mirror of
https://gitlab.com/openconnect/ocserv.git
synced 2026-10-06 22:32:05 +08:00
agents: updated promptkit protocols and requirements
Signed-off-by: Nikos Mavrogiannopoulos <n.mavrogiannopoulos@gmail.com>
This commit is contained in:
@@ -20,7 +20,7 @@ say so and point to where to verify. Do not guess and present guesses as facts.
|
||||
|
||||
---
|
||||
|
||||
## Step 0: Architecture Orientation (Required Before Writing Any Code)
|
||||
## Step 1: Architecture Orientation (Required Before Writing Any Code)
|
||||
|
||||
Before touching a single file, identify which process your change lives in.
|
||||
|
||||
@@ -43,6 +43,37 @@ authentication flow.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Requirements First (Required Before Writing Code)
|
||||
|
||||
`doc/requirements/` is the normative description of what ocserv must do.
|
||||
Before writing any code, find the requirement your change falls under:
|
||||
|
||||
1. Search `doc/requirements/` for the `REQ-*` / `AC-*` entries covering the
|
||||
behavior you're about to add or fix — use `doc/requirements/README.md`'s
|
||||
document map to find the right file by process/subsystem.
|
||||
2. If you are changing behavior an existing requirement describes, update
|
||||
that requirement **first**, before touching implementation code, matching
|
||||
the existing document's ID prefix, category tags, and per-requirement
|
||||
format.
|
||||
3. If no requirement covers the new behavior, add one in the appropriate
|
||||
document before writing the implementation.
|
||||
4. For a bug fix: if the bug violates an existing requirement, cite its ID
|
||||
in your commit/MR. If it reveals a gap, add or extend a requirement
|
||||
describing the *correct* behavior before fixing the code.
|
||||
|
||||
Before opening your MR, self-check the patch against this: load and follow
|
||||
`contrib/ai/protocols/code-compliance-audit.md`, treating your own diff as
|
||||
the code under audit against `doc/requirements/`. The ocserv-specific
|
||||
extensions in that file (document map, ID scheme, the update-before-code
|
||||
ordering check, and the `specification-drift` taxonomy in
|
||||
`contrib/ai/taxonomies/specification-drift.md`) tell you exactly what to
|
||||
check. Any finding classified **D8** (unimplemented requirement) or **D10**
|
||||
(constraint violation in code) means the patch is not ready to submit —
|
||||
resolve it, by fixing the code or updating the requirement, whichever is
|
||||
actually wrong, before opening the MR.
|
||||
|
||||
---
|
||||
|
||||
## Security Disclosure — Stop and Read If This Applies
|
||||
|
||||
**If you believe you have found a security vulnerability:**
|
||||
@@ -156,30 +187,35 @@ If your change touches a subsystem with dedicated documentation, read it first:
|
||||
|
||||
### Submitting a Feature
|
||||
|
||||
1. **Open an issue first.** Describe the motivation, the proposed design, and which
|
||||
1. **Requirements first** — see Step 2 above. Find or add the covering
|
||||
`REQ-*` entry before writing any code.
|
||||
2. **Open an issue first.** Describe the motivation, the proposed design, and which
|
||||
process it lives in. Wait for maintainer feedback before writing code. Features
|
||||
without prior design discussion are often asked to redesign after implementation.
|
||||
2. Implement in the correct process (see Step 0).
|
||||
3. Add configuration if needed: global options go in `src/config.c`; per-module
|
||||
3. Implement in the correct process (see Step 1).
|
||||
4. Add configuration if needed: global options go in `src/config.c`; per-module
|
||||
options go in a struct in `src/common-config.h` and a parser in `src/subconfig.c`.
|
||||
4. Write tests (see checklist below).
|
||||
5. Update relevant documentation (`doc/sample.config`, man pages if applicable).
|
||||
5. Write tests (see checklist below).
|
||||
6. Update relevant documentation (`doc/sample.config`, man pages if applicable).
|
||||
|
||||
### Submitting a Bug Fix
|
||||
|
||||
1. **Characterize the symptom precisely** before touching any code:
|
||||
1. **Requirements first** — see Step 2 above. If the bug is a requirement
|
||||
violation, cite the `REQ-*` ID; if it reveals a gap, extend or add the
|
||||
requirement before fixing the code.
|
||||
2. **Characterize the symptom precisely** before touching any code:
|
||||
- Which process emitted the error (main / sec-mod / worker)?
|
||||
- Is it deterministic or intermittent?
|
||||
- What changed recently that might have introduced it?
|
||||
2. **Generate at least 3 hypotheses** for the root cause before investigating any of them.
|
||||
3. **Generate at least 3 hypotheses** for the root cause before investigating any of them.
|
||||
Include one non-obvious hypothesis (timing, config interaction, allocator mismatch).
|
||||
3. **Distinguish root from proximate cause.**
|
||||
4. **Distinguish root from proximate cause.**
|
||||
Proximate: "null pointer dereference at line X." Root: "the function that
|
||||
initializes the pointer silently fails when Y, leaving the caller with an
|
||||
uninitialized value." Fix the root cause — if you fix only the proximate cause,
|
||||
the root cause will produce a different failure later.
|
||||
4. Write a test that reproduces the bug (it must fail before your fix).
|
||||
5. Apply the fix. Confirm the test passes and no other tests regress.
|
||||
5. Write a test that reproduces the bug (it must fail before your fix).
|
||||
6. Apply the fix. Confirm the test passes and no other tests regress.
|
||||
|
||||
### Writing a Test
|
||||
|
||||
@@ -254,6 +290,12 @@ The canonical checklist is in `CONTRIBUTING.md` → *Before opening a merge requ
|
||||
The following expands it with agent-specific verification steps.
|
||||
|
||||
**Agent-runnable — you must verify these:**
|
||||
- [ ] Relevant `REQ-*`/`AC-*` entry found or added in `doc/requirements/`
|
||||
*before* the code change (Step 2)
|
||||
- [ ] If an existing requirement's described behavior changed: the requirement
|
||||
was updated first, not left contradicting the code
|
||||
- [ ] Self-audit run per Step 2 (`contrib/ai/protocols/code-compliance-audit.md`);
|
||||
no open D8/D10 findings against `doc/requirements/`
|
||||
- [ ] Every changed line is independently justifiable — no drive-by refactoring
|
||||
- [ ] Original types preserved; no unrelated reformatting
|
||||
- [ ] `clang-format --dry-run -Werror` passes on every modified file under `src/` and `tests/`
|
||||
|
||||
@@ -45,48 +45,37 @@ do not continue grading the remaining sections as if the patch were approvable.
|
||||
|
||||
## Protocol: Requirements Compliance
|
||||
|
||||
Run this as step 1 of **Protocol: Contribution Review**. This is the single most
|
||||
important check in the review: a patch that is well-designed but contradicts or
|
||||
ignores `doc/requirements/` is not acceptable, regardless of code quality.
|
||||
Run this as step 1 of **Protocol: Contribution Review**, before Design Review.
|
||||
This is the single most important check in the review: a patch that is
|
||||
well-designed but contradicts or ignores `doc/requirements/` is not acceptable,
|
||||
regardless of code quality.
|
||||
|
||||
For every file, function, or config option touched by the patch:
|
||||
Load and follow `contrib/ai/protocols/code-compliance-audit.md` for the full
|
||||
audit protocol (specification inventory, forward/backward traceability,
|
||||
constraint verification, and classification against the `specification-drift`
|
||||
taxonomy in `contrib/ai/taxonomies/specification-drift.md`). The ocserv-specific
|
||||
extensions — the `doc/requirements/` document map and ID scheme, the
|
||||
update-before-code ordering check from AGENTS.md's Requirements-First Workflow,
|
||||
CFG/SEC/AUTH/IPC-specific traceability evidence, and the mapping from D8–D10
|
||||
findings to the verdicts below — are in the extension section of that file (and
|
||||
of the taxonomy file).
|
||||
|
||||
1. Search `doc/requirements/` for `REQ-*` / `AC-*` / `OC-*` entries that cite it
|
||||
(grep for the file/function/option name, and check the document map in
|
||||
`doc/requirements/README.md` for the right file by process/subsystem).
|
||||
2. Record one of:
|
||||
- **compliant** — an existing requirement covers this behavior and the patch
|
||||
matches it.
|
||||
- **updated** — the patch changes behavior an existing requirement describes;
|
||||
confirm the requirement was updated *first*, in the same or a preceding
|
||||
commit, following the protocol in `contrib/ai/protocols/` that generated
|
||||
that document. If the requirement was not updated, this is a **BLOCK**.
|
||||
- **new requirement added** — the patch introduces behavior with no prior
|
||||
requirement; confirm a new `REQ-*` entry was added in the appropriate
|
||||
document, with the correct ID prefix, category tags, and per-requirement
|
||||
format. If none was added, this is a **BLOCK**.
|
||||
- **gap** — behavior is touched but no requirement covers it and none was
|
||||
added. This is a **BLOCK**, not a note.
|
||||
- **contradicts REQ-X** — the patch's behavior conflicts with an existing
|
||||
requirement that the patch did not update. This is a **BLOCK** unless the
|
||||
requirement itself is independently wrong, in which case say so explicitly
|
||||
and require it be fixed in its own dedicated MR (per AGENTS.md), not
|
||||
silently bundled here.
|
||||
3. Separately, check for collateral damage: search `doc/requirements/` for any
|
||||
`REQ-*`/`AC-*` entries citing the touched files/functions that the patch does
|
||||
*not* intend to change, and confirm each still holds. Flag any that no longer
|
||||
hold as **REVIEW** (requirement vs. code now disagree) — never approve a patch
|
||||
that leaves a `DERIVED` requirement contradicting the code.
|
||||
4. Confirm `doc/ocserv.8.md` / `doc/sample.config` agree with the requirement and
|
||||
the code where applicable (config options, documented behavior).
|
||||
Any finding classified D8 (unimplemented requirement) or D10 (constraint
|
||||
violation in code) is a **BLOCK**: do not approve the patch. A High-severity D9
|
||||
finding (undocumented behavior in a SEC/AUTH/IPC area) is also a **BLOCK**;
|
||||
other D9 findings are a **REVIEW** item to raise with the maintainer, not an
|
||||
automatic rejection.
|
||||
|
||||
Verdict per touched surface: *compliant* | *updated* | *new requirement added* |
|
||||
*gap — BLOCK* | *contradicts REQ-X — BLOCK* | *REVIEW (pre-existing requirement
|
||||
now inconsistent, not caused by this patch)*.
|
||||
Separately, check for collateral damage per the audit protocol's backward
|
||||
traceability phase: search `doc/requirements/` for `REQ-*`/`AC-*` entries
|
||||
citing files/functions the patch touches but does not intend to change, and
|
||||
confirm each still holds. An unrelated requirement now contradicted by the
|
||||
patch is also a **BLOCK** (classify D10) even if the patch's own intended
|
||||
behavior is fully compliant — never approve a patch that leaves a requirement
|
||||
contradicting the code.
|
||||
|
||||
Any `BLOCK` verdict means: do not approve the patch. State which requirement is
|
||||
missing, outdated, or contradicted, and what update (to the requirement, or to
|
||||
the patch) would resolve it.
|
||||
State the verdict per touched requirement, citing the finding's drift label
|
||||
and `REQ-*`/`AC-*`/`OC-*` ID. Do not approve a patch with an open BLOCK.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,269 @@
|
||||
<!-- SPDX-License-Identifier: MIT -->
|
||||
<!-- Copyright (c) PromptKit Contributors -->
|
||||
|
||||
---
|
||||
name: code-compliance-audit
|
||||
type: reasoning
|
||||
description: >
|
||||
Systematic protocol for auditing source code against requirements and
|
||||
design documents. Maps specification claims to code behavior, detects
|
||||
unimplemented requirements, undocumented behavior, and constraint
|
||||
violations. Classifies findings using the specification-drift taxonomy
|
||||
(D8–D10).
|
||||
applicable_to:
|
||||
- audit-code-compliance
|
||||
---
|
||||
|
||||
# Protocol: Code Compliance Audit
|
||||
|
||||
Apply this protocol when auditing source code against requirements and
|
||||
design documents to determine whether the implementation matches the
|
||||
specification. The goal is to find every gap between what was specified
|
||||
and what was built — in both directions.
|
||||
|
||||
## Phase 1: Specification Inventory
|
||||
|
||||
Extract the audit targets from the specification documents.
|
||||
|
||||
1. **Requirements document** — extract:
|
||||
- Every REQ-ID with its summary, acceptance criteria, and category
|
||||
- Every constraint (performance, security, behavioral)
|
||||
- Every assumption that affects implementation
|
||||
- Defined terms and their precise meanings
|
||||
|
||||
2. **Design document** (if provided) — extract:
|
||||
- Components, modules, and interfaces described
|
||||
- API contracts (signatures, pre/postconditions, error handling)
|
||||
- Data models and state management approach
|
||||
- Non-functional strategies (caching, pooling, concurrency model)
|
||||
- Explicit mapping of design elements to REQ-IDs
|
||||
|
||||
3. **Build a requirements checklist**: a flat list of every testable
|
||||
claim from the specification that can be verified against code.
|
||||
Each entry has: REQ-ID, the specific behavior or constraint, and
|
||||
what evidence in code would confirm implementation.
|
||||
|
||||
## Phase 2: Code Inventory
|
||||
|
||||
Survey the source code to understand its structure before tracing.
|
||||
|
||||
1. **Module/component map**: Identify the major code modules, classes,
|
||||
or packages and their responsibilities.
|
||||
2. **API surface**: Catalog public functions, endpoints, interfaces —
|
||||
the externally visible behavior.
|
||||
3. **Configuration and feature flags**: Identify behavior that is
|
||||
conditionally enabled or parameterized.
|
||||
4. **Error handling paths**: Catalog how errors are handled — these
|
||||
often implement (or fail to implement) requirements around
|
||||
reliability and graceful degradation.
|
||||
|
||||
Do NOT attempt to understand every line of code. Focus on the
|
||||
**behavioral surface** — what the code does, not how it does it
|
||||
internally — unless the specification constrains the implementation
|
||||
approach.
|
||||
|
||||
## Phase 3: Forward Traceability (Specification → Code)
|
||||
|
||||
For each requirement in the checklist:
|
||||
|
||||
1. **Search for implementation**: Identify the code module(s),
|
||||
function(s), or path(s) that implement this requirement.
|
||||
- Look for explicit references (comments citing REQ-IDs, function
|
||||
names matching requirement concepts).
|
||||
- Look for behavioral evidence (code that performs the specified
|
||||
action under the specified conditions).
|
||||
- Check configuration and feature flags that may gate the behavior.
|
||||
|
||||
2. **Assess implementation completeness**:
|
||||
- Does the code implement the **full** requirement, including edge
|
||||
cases described in acceptance criteria?
|
||||
- Does the code implement the requirement under all specified
|
||||
conditions, or only the common case?
|
||||
- Are constraints (performance, resource limits, timing) enforced?
|
||||
|
||||
3. **Classify the result**:
|
||||
- **IMPLEMENTED**: Code clearly implements the requirement. Record
|
||||
the code location(s) as evidence.
|
||||
- **PARTIALLY IMPLEMENTED**: Some aspects are present but acceptance
|
||||
criteria are not fully met. Flag as D8_UNIMPLEMENTED_REQUIREMENT
|
||||
with the finding describing what is present and what is missing.
|
||||
Set confidence to Medium.
|
||||
- **NOT IMPLEMENTED**: No code implements this requirement. Flag as
|
||||
D8_UNIMPLEMENTED_REQUIREMENT with confidence High.
|
||||
|
||||
## Phase 4: Backward Traceability (Code → Specification)
|
||||
|
||||
Identify code behavior that is not specified.
|
||||
|
||||
1. **For each significant code module or feature**: determine whether
|
||||
it traces to a requirement or design element.
|
||||
- "Significant" means it implements user-facing behavior, data
|
||||
processing, access control, external communication, or state
|
||||
changes. Infrastructure (logging, metrics, boilerplate) is not
|
||||
significant unless the specification constrains it.
|
||||
|
||||
2. **Flag undocumented behavior**:
|
||||
- Code that implements meaningful behavior with no tracing
|
||||
requirement is a candidate D9_UNDOCUMENTED_BEHAVIOR.
|
||||
- Distinguish between: (a) genuine scope creep, (b) reasonable
|
||||
infrastructure that supports requirements indirectly, and
|
||||
(c) requirements gaps (behavior that should have been specified).
|
||||
Report all three, but note the distinction.
|
||||
|
||||
## Phase 5: Constraint Verification
|
||||
|
||||
Check that specified constraints are respected in the implementation.
|
||||
|
||||
1. **For each constraint in the requirements**:
|
||||
- Identify the code path(s) responsible for satisfying it.
|
||||
- Assess whether the implementation approach **can** satisfy the
|
||||
constraint (algorithmic feasibility, not just correctness).
|
||||
- Check for explicit violations — code that demonstrably contradicts
|
||||
the constraint.
|
||||
|
||||
2. **Common constraint categories to check**:
|
||||
- Performance: response time limits, throughput requirements,
|
||||
resource consumption bounds
|
||||
- Security: encryption requirements, authentication enforcement,
|
||||
input validation, access control
|
||||
- Data integrity: validation rules, consistency guarantees,
|
||||
atomicity requirements
|
||||
- Compatibility: API versioning, backward compatibility,
|
||||
interoperability constraints
|
||||
|
||||
3. **Flag violations** as D10_CONSTRAINT_VIOLATION_IN_CODE with
|
||||
specific evidence (code location, the constraint, and how the
|
||||
code violates it).
|
||||
|
||||
## Phase 6: Classification and Reporting
|
||||
|
||||
Classify every finding using the specification-drift taxonomy
|
||||
(see `taxonomies/specification-drift.md` for full definitions).
|
||||
|
||||
1. Assign exactly one drift label (D8, D9, or D10) to each finding.
|
||||
2. Assign severity using the taxonomy's severity guidance.
|
||||
3. For each finding, provide:
|
||||
- The drift label and short title
|
||||
- The spec location (REQ-ID, section) and code location (file,
|
||||
function, line range). For D9 findings, the spec location is
|
||||
"None — no matching requirement identified" with a description
|
||||
of what was searched.
|
||||
- Evidence: what the spec says and what the code does (or doesn't)
|
||||
- Impact: what could go wrong
|
||||
- Recommended resolution
|
||||
4. Order findings primarily by severity, then by taxonomy ranking
|
||||
within each severity tier.
|
||||
|
||||
## Phase 7: Coverage Summary
|
||||
|
||||
After reporting individual findings, produce aggregate metrics:
|
||||
|
||||
1. **Implementation coverage**: % of REQ-IDs with confirmed
|
||||
implementations in code.
|
||||
2. **Undocumented behavior rate**: count of significant code behaviors
|
||||
with no tracing requirement.
|
||||
3. **Constraint compliance**: count of constraints verified vs.
|
||||
violated vs. unverifiable from code analysis alone.
|
||||
4. **Overall assessment**: a summary judgment of code-to-spec alignment.
|
||||
|
||||
<!-- BEGIN ocserv extensions -->
|
||||
|
||||
## ocserv-Specific Extensions
|
||||
|
||||
The sections below extend the generic protocol with ocserv's requirements
|
||||
tree, document map, and process/privilege model. Apply these alongside the
|
||||
base phases above — they do not replace them. This protocol is the audit
|
||||
engine behind **Protocol: Requirements Compliance** in the core-dev persona,
|
||||
and is also the tool an external contributor should self-apply before
|
||||
opening an MR.
|
||||
|
||||
### Phase 1 — Specification Inventory (ocserv)
|
||||
|
||||
- The "requirements document" is the `doc/requirements/` tree. Consult
|
||||
`doc/requirements/README.md` for the document map (which file covers
|
||||
which process/subsystem) and the ID scheme: `REQ-<AREA>-<NNN>` entries,
|
||||
linked `AC-*` acceptance criteria, and category tags (`AUTH`, `IPC`,
|
||||
`CFG`, `SEC`, `NET`, `COMPAT`, `ACCT`, `LOG`).
|
||||
- ocserv has no separate design document beyond `doc/design.md` (IPC and
|
||||
process architecture), `doc/ocserv.8.md`, and `sample.config` (documented
|
||||
configuration behavior). Treat these three as the "design document" role
|
||||
in Phase 1.
|
||||
- Build the requirements checklist by finding the `REQ-*`/`AC-*` entries
|
||||
that cite the files, functions, or config options actually touched by
|
||||
the diff — this is a patch-scoped audit, not a whole-project audit,
|
||||
unless a full-tree audit is explicitly requested.
|
||||
|
||||
### Phase 1.5 — Update-Before-Code Check (ocserv)
|
||||
|
||||
AGENTS.md's Requirements-First Workflow requires: if the patch changes
|
||||
behavior an existing requirement describes, that requirement must be
|
||||
updated **before** the implementation change, re-applying the protocol in
|
||||
`contrib/ai/protocols/` that generated its document (see
|
||||
`requirements-elicitation.md`, `requirements-from-implementation.md`, or
|
||||
`requirements-reconciliation.md` as applicable).
|
||||
|
||||
Before running Phase 3: for every `REQ-*` whose described behavior the
|
||||
patch changes, confirm the patch includes a corresponding update to that
|
||||
requirement (typically a preceding commit in the same MR). If it does not,
|
||||
this is a hard **BLOCK** — classify as D10_CONSTRAINT_VIOLATION_IN_CODE,
|
||||
since the code now contradicts requirement text that still stands
|
||||
unmodified, regardless of whether the new behavior is otherwise correct.
|
||||
If the requirement is independently wrong for reasons unrelated to this
|
||||
patch, it must be fixed in its own dedicated MR (per AGENTS.md) — do not
|
||||
treat that as satisfying this check.
|
||||
|
||||
### Phase 3 — Forward Traceability (ocserv)
|
||||
|
||||
- "Implemented" evidence for a `CFG` requirement includes the parser
|
||||
itself (`src/config.c` / `src/subconfig.c`) plus the `[scope:]`
|
||||
annotation in `sample.config` and, for reloadable global options, the
|
||||
corresponding `error_on_vhost()` call in `src/config.c`.
|
||||
- For `SEC`/`AUTH`/`IPC`-tagged requirements, a missing **negative**
|
||||
acceptance criterion path (e.g., rejection of a tampered cookie, a
|
||||
replayed SID, a bad password) is at most PARTIALLY IMPLEMENTED — per
|
||||
REQ-GEN-TEST-002's negative-test-first rule, do not classify such a
|
||||
requirement as fully IMPLEMENTED on the strength of the positive path
|
||||
alone.
|
||||
|
||||
### Phase 4 — Backward Traceability (ocserv)
|
||||
|
||||
- Significant, trace-worthy surfaces in ocserv: new or changed config
|
||||
options (`common-config.h` / `subconfig.c` / `config.c`), new or
|
||||
changed IPC fields (`ipc.proto` / `ctl.proto`), new auth/acct module
|
||||
registrations (`sec-mod.c`, `src/auth/`, `src/acct/`), and any code that
|
||||
crosses the main/sec-mod/worker privilege boundary. `mslog()` / `oclog()`
|
||||
/ `seclog()` calls and talloc bookkeeping are infrastructure, not
|
||||
significant, unless a requirement specifically constrains logging.
|
||||
- Classify undocumented behavior touching `SEC`, `AUTH`, or `IPC` category
|
||||
areas as High severity per the taxonomy's own guidance, and separately
|
||||
flag it under AGENTS.md's "Human-judgment required" list (privilege
|
||||
boundary crossing, new auth method) — a maintainer must see this, not
|
||||
just the audit output.
|
||||
|
||||
### Phase 5 — Constraint Verification (ocserv)
|
||||
|
||||
- Treat AGENTS.md's architecture invariant — no credential handling in a
|
||||
worker, no direct filesystem/socket access outside seccomp, no silent
|
||||
collapse of the main/sec-mod/worker boundary — as a standing,
|
||||
project-wide constraint even when no single `REQ-*` states it verbatim.
|
||||
A violation is always D10_CONSTRAINT_VIOLATION_IN_CODE, Critical
|
||||
severity, and requires explicit maintainer acknowledgment per AGENTS.md
|
||||
— a requirements-doc fix alone does not resolve it.
|
||||
- The canonical technology choices (REQ-GEN-TECH-001 through -005:
|
||||
talloc-only allocation, GnuTLS-only cryptography via `tlslib.c`,
|
||||
protobuf-c for IPC with regenerated bindings, INI-only configuration,
|
||||
approval required for new dependencies) are constraints in this sense
|
||||
too. A violation is D10.
|
||||
|
||||
### Reporting (ocserv)
|
||||
|
||||
- Map findings to the BLOCK/REVIEW verdicts used by **Protocol:
|
||||
Requirements Compliance**: D8 and D10 findings are always a BLOCK.
|
||||
A High-severity D9 finding (SEC/AUTH/IPC area) is also a BLOCK; other
|
||||
D9 findings are a REVIEW item to raise with the maintainer rather than
|
||||
an automatic rejection.
|
||||
- Cite `REQ-*` / `AC-*` / `OC-*` IDs exactly as they appear in
|
||||
`doc/requirements/`. Never invent or approximate an ID — if you cannot
|
||||
find the citation, say so and report a gap instead of guessing.
|
||||
|
||||
<!-- END ocserv extensions -->
|
||||
@@ -0,0 +1,354 @@
|
||||
<!-- SPDX-License-Identifier: MIT -->
|
||||
<!-- Copyright (c) PromptKit Contributors -->
|
||||
|
||||
---
|
||||
name: prompt-determinism-analysis
|
||||
type: analysis
|
||||
description: >
|
||||
Systematic analysis of prompt and instruction text for language
|
||||
precision and determinism. Identifies vague quantifiers, subjective
|
||||
adjectives, missing constraints, incomplete conditionals, and
|
||||
ambiguous references that introduce non-deterministic LLM behavior.
|
||||
Classifies each finding as High, Medium, or Low non-determinism
|
||||
potential with concrete rewrite suggestions.
|
||||
applicable_to:
|
||||
- lint-prompt
|
||||
- audit-library-health
|
||||
---
|
||||
|
||||
# Protocol: Prompt Determinism Analysis
|
||||
|
||||
Apply this protocol when analyzing prompt text, instruction files,
|
||||
or prompt library components for language that introduces
|
||||
non-deterministic LLM behavior. Execute all phases in order.
|
||||
|
||||
## Determinism Classification Scale
|
||||
|
||||
| Level | Meaning | Action |
|
||||
|-------|---------|--------|
|
||||
| **High** | Language is vague, subjective, or open-ended. Different LLMs (or the same LLM across runs) will interpret it inconsistently. | Rewrite required — provide concrete rewrite suggestion. |
|
||||
| **Medium** | Language is imprecise but constrained by surrounding context. Interpretation may vary at the margins but the core intent is recoverable. | Rewrite recommended — flag with suggestion. |
|
||||
| **Low** | Language is concrete, specific, and leaves little room for interpretation. Enumerated values, explicit constraints, named artifacts, numbered steps. | No rewrite needed — counted in scorecard only by default. Templates may optionally report Low findings as individual Informational-severity entries when configured for strict analysis. |
|
||||
|
||||
## Phase 1: Lexical Pattern Scan
|
||||
|
||||
Scan the text for specific lexical patterns known to introduce
|
||||
non-determinism. For each occurrence, record the location, the
|
||||
pattern category, and the classification level.
|
||||
|
||||
### 1.1 Vague Quantifiers (High)
|
||||
|
||||
Flag words that leave quantity or degree unspecified:
|
||||
|
||||
- "some", "several", "many", "a few", "a number of", "various",
|
||||
"numerous", "multiple" (when not followed by a specific count)
|
||||
- "often", "usually", "sometimes", "occasionally", "frequently"
|
||||
- "most", "almost all", "nearly"
|
||||
|
||||
**Rewrite pattern**: Replace with a specific count, range, or
|
||||
enumeration. If the exact count is unknowable, state the selection
|
||||
criterion instead (e.g., "at least 3" or "all items matching X").
|
||||
|
||||
### 1.2 Subjective Adjectives (High)
|
||||
|
||||
Flag adjectives that depend on unstated evaluation criteria:
|
||||
|
||||
- "good", "bad", "appropriate", "suitable", "reasonable", "proper",
|
||||
"adequate", "sufficient", "clean", "elegant", "simple",
|
||||
"straightforward", "clear", "obvious", "intuitive"
|
||||
- "important", "significant", "critical", "key", "major", "minor"
|
||||
(when used without a defined severity scale or an explicit
|
||||
enumeration of what qualifies — e.g., "significant" is acceptable
|
||||
if immediately followed by criteria such as "affecting >2 components")
|
||||
|
||||
**Rewrite pattern**: Replace with observable criteria. "Good error
|
||||
handling" → "Error handling that catches all thrown exception types,
|
||||
logs the error with context, and returns a structured error response."
|
||||
|
||||
### 1.3 Open-Ended Enumerations (Medium)
|
||||
|
||||
Flag lists that signal incompleteness without bounding:
|
||||
|
||||
- "etc.", "and so on", "and more", "among others", "for example"
|
||||
(when used as the sole specification, not as illustration before
|
||||
a complete list)
|
||||
- "such as X, Y, …" without a closing exhaustive rule
|
||||
- "including but not limited to"
|
||||
|
||||
**Rewrite pattern**: Either enumerate exhaustively, or state the
|
||||
selection criterion explicitly. "Check for issues such as SQL
|
||||
injection, XSS, etc." → "Check for all OWASP Top 10 vulnerability
|
||||
categories."
|
||||
|
||||
### 1.4 Hedge Words and Weak Modals (Medium)
|
||||
|
||||
Flag words that weaken commitment to an action:
|
||||
|
||||
- "might", "could", "possibly", "perhaps", "consider",
|
||||
"may want to", "it would be nice to", "try to", "attempt to"
|
||||
- "if possible", "if applicable", "when appropriate", "as needed"
|
||||
(without criteria for when it IS applicable/needed)
|
||||
|
||||
**Rewrite pattern**: Replace with a concrete conditional. "Consider
|
||||
checking for null" → "Check for null on every pointer dereference."
|
||||
"If appropriate, add logging" → "Add logging when the function
|
||||
returns an error code."
|
||||
|
||||
### 1.5 Passive Voice Without Actor (Medium)
|
||||
|
||||
Flag passive constructions where the responsible agent is unclear:
|
||||
|
||||
- "should be reviewed", "must be analyzed", "needs to be checked",
|
||||
"is expected to", "will be handled"
|
||||
|
||||
**Rewrite pattern**: Name the actor explicitly. "The output should
|
||||
be reviewed" → "The LLM must review its own output against the
|
||||
checklist in §5 before presenting it as final."
|
||||
|
||||
### 1.6 Unanchored Comparatives and Superlatives (High)
|
||||
|
||||
Flag comparisons without a baseline or reference point:
|
||||
|
||||
- "better", "worse", "more", "less", "improved", "faster",
|
||||
"simpler", "cleaner", "more efficient"
|
||||
- "the best", "the most", "the least", "optimal"
|
||||
|
||||
**Rewrite pattern**: Anchor to a measurable criterion or a specific
|
||||
comparison target. "A better approach" → "An approach that reduces
|
||||
time complexity from O(n²) to O(n log n)." "The most important
|
||||
findings" → "Findings classified as Critical or High severity."
|
||||
|
||||
## Phase 2: Structural Completeness
|
||||
|
||||
Check for structural gaps that leave behavior underspecified.
|
||||
|
||||
### 2.1 Conditionals Without Exhaustive Branches (High)
|
||||
|
||||
For every conditional instruction ("if X, do Y"):
|
||||
|
||||
1. Check whether all branches are specified.
|
||||
2. Flag conditionals that specify the positive case but omit the
|
||||
negative case, the edge case, or the error case.
|
||||
3. Flag "if/else" constructs where the else is vague ("otherwise,
|
||||
use your judgment").
|
||||
|
||||
**Rewrite pattern**: Add explicit else/default branches. "If the
|
||||
file exists, parse it" → "If the file exists, parse it. If the
|
||||
file does not exist, report finding F-NNN with severity High."
|
||||
|
||||
### 2.2 Missing Bounds and Constraints (High)
|
||||
|
||||
Flag instructions that reference quantities, sizes, or durations
|
||||
without concrete limits:
|
||||
|
||||
- "Limit the output" (to what?)
|
||||
- "Keep it concise" (how many words/sections/items?)
|
||||
- "A reasonable number" (what number?)
|
||||
- "Recent" (how recent — last 7 days? last commit?)
|
||||
|
||||
**Rewrite pattern**: Add explicit bounds. "Keep the summary
|
||||
concise" → "The summary must be 2–4 sentences."
|
||||
|
||||
### 2.3 Missing Exit Criteria (Medium)
|
||||
|
||||
Flag loops, iterations, or recursive processes that lack a
|
||||
termination condition:
|
||||
|
||||
- "Repeat until satisfied" (what defines satisfaction?)
|
||||
- "Continue refining" (when does refinement stop?)
|
||||
- "Iterate as needed" (what signals completion?)
|
||||
|
||||
**Rewrite pattern**: Define the exit condition explicitly.
|
||||
"Iterate until the design is complete" → "Iterate until all
|
||||
requirements in the input have a corresponding design section
|
||||
with at least one acceptance criterion addressed."
|
||||
|
||||
### 2.4 Unspecified Ordering or Priority (Medium)
|
||||
|
||||
Flag instructions that present multiple items without specifying
|
||||
execution order or relative priority:
|
||||
|
||||
- "Consider factors A, B, and C" (in what order? equal weight?)
|
||||
- "Review the following areas" (sequentially? in parallel?
|
||||
by priority?)
|
||||
- "Address these concerns" (which first?)
|
||||
|
||||
**Rewrite pattern**: Number the steps or state the priority rule.
|
||||
"Consider security, performance, and readability" → "Evaluate in
|
||||
this priority order: (1) security, (2) correctness,
|
||||
(3) performance, (4) readability."
|
||||
|
||||
### 2.5 Missing Output Specification (High)
|
||||
|
||||
Flag instructions that request output without specifying:
|
||||
|
||||
- The structure (sections, fields, format)
|
||||
- The granularity (per-file, per-function, per-finding)
|
||||
- The artifact type (report, list, table, code block)
|
||||
|
||||
**Rewrite pattern**: Add explicit output structure. "Report your
|
||||
findings" → "Report each finding using the template: Severity,
|
||||
Location, Description, Evidence, Remediation."
|
||||
|
||||
## Phase 3: Semantic Precision
|
||||
|
||||
Assess whether instructions are specific enough to produce
|
||||
consistent behavior across different LLM sessions.
|
||||
|
||||
### 3.1 Abstract Action Verbs (Medium)
|
||||
|
||||
Flag action verbs that describe a goal without specifying the
|
||||
method:
|
||||
|
||||
- "analyze", "evaluate", "assess", "examine", "investigate",
|
||||
"review", "study", "explore"
|
||||
|
||||
These are acceptable ONLY when followed by numbered sub-steps, each
|
||||
naming a concrete action (not another abstract verb) with a
|
||||
measurable completion condition. Flag instances where the verb
|
||||
stands alone as the complete instruction.
|
||||
|
||||
**Rewrite pattern**: Decompose into concrete sub-steps.
|
||||
"Analyze the code for issues" → "For each function: (1) check
|
||||
parameter validation, (2) trace error propagation paths,
|
||||
(3) verify resource cleanup in all exit paths."
|
||||
|
||||
### 3.2 Undefined Domain Terms (Medium)
|
||||
|
||||
Flag terms that have domain-specific meaning but are not defined
|
||||
in the prompt:
|
||||
|
||||
- Technical jargon used without definition or reference
|
||||
- Acronyms not expanded on first use
|
||||
- Terms that have different meanings in different contexts
|
||||
(e.g., "component" in React vs. hardware vs. prompt engineering)
|
||||
|
||||
**Rewrite pattern**: Define the term on first use or reference an
|
||||
external definition. "Check for race conditions" → "Check for race
|
||||
conditions (concurrent access to shared mutable state without
|
||||
synchronization)."
|
||||
|
||||
### 3.3 Implicit Context Dependencies (High)
|
||||
|
||||
Flag instructions that assume context not provided in the prompt:
|
||||
|
||||
- References to "the project", "the codebase", "the system"
|
||||
without specifying what is in scope
|
||||
- Assumed knowledge of conventions, tools, or processes not
|
||||
stated in the prompt
|
||||
- References to "previous" results, "earlier" analysis, or
|
||||
"above" without explicit back-references
|
||||
|
||||
**Rewrite pattern**: Make the context explicit. "Follow the
|
||||
project's conventions" → "Follow the conventions defined in
|
||||
CONTRIBUTING.md, specifically: [list the relevant conventions]."
|
||||
|
||||
### 3.4 Missing Examples (Low–Medium)
|
||||
|
||||
Flag complex or novel instructions that lack illustrative
|
||||
examples:
|
||||
|
||||
- Classification schemes without example classifications
|
||||
- Output formats without a concrete sample
|
||||
- Pattern descriptions without concrete instances
|
||||
|
||||
Classify as Medium when the instruction introduces a concept,
|
||||
schema, category set, output structure, or term that is central
|
||||
to the task and the same document does not provide either (a) an
|
||||
explicit definition, or (b) at least one concrete example.
|
||||
Classify as Low when the instruction lacks an example but the same
|
||||
document already makes the meaning operational through an explicit
|
||||
definition, sample output, glossary entry, or enumerated categories
|
||||
or steps.
|
||||
|
||||
**Rewrite pattern**: Add at least one concrete example for each
|
||||
novel concept. For classification schemes, provide one example
|
||||
per category.
|
||||
|
||||
## Phase 4: Classification and Reporting
|
||||
|
||||
After completing Phases 1–3, produce the determinism assessment.
|
||||
|
||||
### 4.1 Per-Instruction Scoring
|
||||
|
||||
For each flagged instruction or passage:
|
||||
|
||||
1. Record the location (section heading, line, or passage excerpt).
|
||||
2. Assign a determinism level (High / Medium / Low non-determinism).
|
||||
3. Cite the specific pattern from Phase 1, 2, or 3 that triggered
|
||||
the flag.
|
||||
4. For High and Medium findings, provide a concrete rewrite
|
||||
suggestion that would reduce the non-determinism level by at
|
||||
least one step. For Low findings, record "No rewrite needed."
|
||||
|
||||
### 4.2 Per-Section Aggregation
|
||||
|
||||
For each logical section of the analyzed text:
|
||||
|
||||
1. Count findings by level (High / Medium / Low).
|
||||
2. Assign an overall section determinism grade:
|
||||
- **Precise**: 0 High, ≤ 2 Medium
|
||||
- **Acceptable**: 0 High, > 2 Medium; or 1 High with ≤ 2 Medium
|
||||
- **Imprecise**: ≥ 2 High, or 1 High with > 2 Medium
|
||||
3. Sections graded Imprecise should be flagged for priority rewrite.
|
||||
|
||||
### 4.3 Overall Assessment
|
||||
|
||||
Produce an overall determinism summary:
|
||||
|
||||
1. Total findings by level across the entire text.
|
||||
2. Overall grade (Precise / Acceptable / Imprecise) using the
|
||||
same thresholds as section grading, applied to the full text.
|
||||
3. Top 3–5 highest-impact rewrite recommendations, ordered by
|
||||
the degree of non-determinism reduction.
|
||||
4. A per-section scorecard table:
|
||||
|
||||
| Section | High | Medium | Low | Grade |
|
||||
|---------|------|--------|-----|-------|
|
||||
| ... | ... | ... | ... | ... |
|
||||
|
||||
## Output Format
|
||||
|
||||
For each finding, report:
|
||||
|
||||
```
|
||||
[DETERMINISM: High | Medium | Low]
|
||||
Pattern: <pattern category from Phase 1/2/3 — e.g., "1.2 Subjective Adjective">
|
||||
Location: <section heading or line reference>
|
||||
Original: "<exact text flagged>"
|
||||
Issue: <why this introduces non-determinism>
|
||||
Rewrite: "<concrete suggested replacement>"
|
||||
```
|
||||
|
||||
<!-- BEGIN ocserv extensions -->
|
||||
|
||||
## ocserv-Specific Extensions
|
||||
|
||||
Apply this protocol to the **Requirement** and **Acceptance** fields of a
|
||||
drafted `doc/requirements/` entry (see `doc/requirements/README.md`'s
|
||||
per-requirement format), and to any new or edited AGENTS.md /
|
||||
`contrib/ai/personas/` / `contrib/ai/protocols/` prose — instruction text
|
||||
that other agents will follow deserves the same scrutiny as a requirement's
|
||||
acceptance criterion.
|
||||
|
||||
For ocserv's own ambiguous-term list ("secure", "session", "reload", "all
|
||||
clients") and the other requirements-specific ambiguity checks, see
|
||||
`requirements-elicitation.md`'s "Phase 3 — Ambiguity Detection (ocserv)"
|
||||
extension — that list is authoritative for `doc/requirements/` drafting;
|
||||
do not duplicate or fork it here. Use this protocol's Phase 1–3 for
|
||||
everything that list doesn't cover: hedge words, unanchored comparatives,
|
||||
missing exit/bounds/output-format specification, and abstract action verbs
|
||||
— these apply equally to requirement text and to persona/protocol prose.
|
||||
|
||||
### Application to Contribution Review
|
||||
|
||||
When running **Protocol: Contribution Review** on a patch that adds or
|
||||
edits a `doc/requirements/` entry, apply Phase 1 and Phase 3.2 of this
|
||||
protocol to the new/changed **Requirement** and **Acceptance** text before
|
||||
accepting it as the requirement update Requirements Compliance checks for.
|
||||
An ambiguous requirement cannot itself be the basis for a D8/D9/D10 verdict
|
||||
in `code-compliance-audit.md` — fix the wording first (or flag `AMBIGUOUS`
|
||||
per the Status legend in `doc/requirements/README.md`), then re-run the
|
||||
compliance audit against the clarified text.
|
||||
|
||||
<!-- END ocserv extensions -->
|
||||
@@ -165,7 +165,18 @@ corresponding MUST NOT requirement. Examples:
|
||||
|
||||
### Phase 3 — Ambiguity Detection (ocserv)
|
||||
|
||||
Additional ambiguity patterns to check in the ocserv context:
|
||||
The base phase's `prompt-determinism-analysis` reference is now a concrete
|
||||
file: `contrib/ai/protocols/prompt-determinism-analysis.md`. Load and follow
|
||||
it for the full lexical/structural/semantic scan (vague quantifiers, hedge
|
||||
words, unanchored comparatives, missing exit/bounds/output specification,
|
||||
abstract action verbs) and apply it to each drafted requirement's
|
||||
**Requirement** and **Acceptance** text. The ocserv-specific ambiguous-term
|
||||
list below is authoritative for `doc/requirements/` drafting and is not
|
||||
duplicated in that file.
|
||||
|
||||
Additional ambiguity patterns to check in the ocserv context. Canonical
|
||||
definitions for the terms below are in `doc/requirements/README.md`'s
|
||||
Glossary — cite that definition in the requirement, do not restate it:
|
||||
|
||||
- **"Secure"**: Always replace with a concrete property — e.g., "authenticated
|
||||
via TLS client certificate," "protected from replay by SID validation,"
|
||||
|
||||
@@ -0,0 +1,456 @@
|
||||
<!-- SPDX-License-Identifier: MIT -->
|
||||
<!-- Copyright (c) PromptKit Contributors -->
|
||||
|
||||
---
|
||||
name: specification-drift
|
||||
type: taxonomy
|
||||
description: >
|
||||
Classification scheme for specification drift and divergence across
|
||||
requirements, design, and validation artifacts. Use when auditing
|
||||
document sets for traceability gaps, scope creep, assumption drift,
|
||||
and coverage failures.
|
||||
domain: specification-traceability
|
||||
applicable_to:
|
||||
- audit-traceability
|
||||
- audit-code-compliance
|
||||
- audit-test-compliance
|
||||
- audit-integration-compliance
|
||||
---
|
||||
|
||||
# Taxonomy: Specification Drift
|
||||
|
||||
Use these labels to classify findings when auditing requirements, design,
|
||||
and validation documents for consistency and completeness. Every finding
|
||||
MUST use exactly one label from this taxonomy.
|
||||
|
||||
## Label Group Summaries
|
||||
|
||||
When only a subset of labels is applicable to a given audit type, use
|
||||
these summaries for cross-reference context. The primary label group
|
||||
(determined by the audit template) uses full definitions below; non-primary
|
||||
groups use these summaries to preserve semantic context without consuming
|
||||
prompt tokens.
|
||||
|
||||
- **D1–D7 (Traceability)**: Document-level drift — untraced requirements,
|
||||
untested requirements, orphaned design decisions, orphaned test cases,
|
||||
assumption drift between documents, constraint violations in design,
|
||||
and acceptance criteria mismatches between test plans and requirements.
|
||||
- **D8–D10 (Code Compliance)**: Code-to-spec drift — unimplemented
|
||||
requirements, undocumented behavior in source code, and constraint
|
||||
violations in implementation.
|
||||
- **D11–D13 (Test Compliance)**: Test-to-plan drift — unimplemented test
|
||||
cases, untested acceptance criteria, and assertion mismatches between
|
||||
test code and validation plans.
|
||||
- **D14–D16 (Integration)**: Cross-component drift — unspecified
|
||||
integration flows, interface contract mismatches between components,
|
||||
and untested integration paths.
|
||||
|
||||
## Labels
|
||||
|
||||
### D1_UNTRACED_REQUIREMENT
|
||||
|
||||
A requirement exists in the requirements document but is not referenced
|
||||
or addressed in the design document.
|
||||
|
||||
**Pattern**: REQ-ID appears in the requirements document. No section of
|
||||
the design document references this REQ-ID or addresses its specified
|
||||
behavior.
|
||||
|
||||
**Risk**: The requirement may be silently dropped during implementation.
|
||||
Without a design realization, there is no plan to deliver this capability.
|
||||
|
||||
**Severity guidance**: High when the requirement is functional or
|
||||
safety-critical. Medium when it is a non-functional or low-priority
|
||||
constraint.
|
||||
|
||||
### D2_UNTESTED_REQUIREMENT
|
||||
|
||||
A requirement exists in the requirements document but has no
|
||||
corresponding test case in the validation plan.
|
||||
|
||||
**Pattern**: REQ-ID appears in the requirements document and may appear
|
||||
in the traceability matrix, but no test case (TC-NNN) is linked to it —
|
||||
or the traceability matrix entry is missing entirely.
|
||||
|
||||
**Risk**: The requirement will not be verified. Defects against this
|
||||
requirement will not be caught by the validation process.
|
||||
|
||||
**Severity guidance**: Critical when the requirement is safety-critical
|
||||
or security-related. High for functional requirements. Medium for
|
||||
non-functional requirements with measurable criteria.
|
||||
|
||||
### D3_ORPHANED_DESIGN_DECISION
|
||||
|
||||
A design section, component, or decision does not trace back to any
|
||||
requirement in the requirements document.
|
||||
|
||||
**Pattern**: A design section describes a component, interface, or
|
||||
architectural decision. No REQ-ID from the requirements document is
|
||||
referenced or addressed by this section.
|
||||
|
||||
**Risk**: Scope creep — the design introduces capabilities or complexity
|
||||
not justified by the requirements. Alternatively, the requirements
|
||||
document is incomplete and the design is addressing an unstated need.
|
||||
|
||||
**Severity guidance**: Medium. Requires human judgment — the finding may
|
||||
indicate scope creep (remove from design) or a requirements gap (add a
|
||||
requirement).
|
||||
|
||||
### D4_ORPHANED_TEST_CASE
|
||||
|
||||
A test case in the validation plan does not map to any requirement in
|
||||
the requirements document.
|
||||
|
||||
**Pattern**: TC-NNN exists in the validation plan but references no
|
||||
REQ-ID, or references a REQ-ID that does not exist in the requirements
|
||||
document.
|
||||
|
||||
**Risk**: Test effort is spent on behavior that is not required.
|
||||
Alternatively, the requirements document is incomplete and the test
|
||||
covers an unstated need.
|
||||
|
||||
**Severity guidance**: Low to Medium. The test may still be valuable
|
||||
(e.g., regression or exploratory), but it is not contributing to
|
||||
requirements coverage.
|
||||
|
||||
### D5_ASSUMPTION_DRIFT
|
||||
|
||||
An assumption stated or implied in one document contradicts, extends,
|
||||
or is absent from another document.
|
||||
|
||||
**Pattern**: The design document states an assumption (e.g., "the system
|
||||
will have at most 1000 concurrent users") that is not present in the
|
||||
requirements document's assumptions section — or contradicts a stated
|
||||
constraint. Similarly, the validation plan may assume environmental
|
||||
conditions not specified in requirements.
|
||||
|
||||
**Risk**: Documents are based on incompatible premises. Implementation
|
||||
may satisfy the design's assumptions while violating the requirements'
|
||||
constraints, or vice versa.
|
||||
|
||||
**Severity guidance**: High when the assumption affects architectural
|
||||
decisions or test validity. Medium when it affects non-critical behavior.
|
||||
|
||||
### D6_CONSTRAINT_VIOLATION
|
||||
|
||||
A design decision directly violates a stated requirement or constraint.
|
||||
|
||||
**Pattern**: The requirements document states a constraint (e.g.,
|
||||
"the system MUST respond within 200ms") and the design document
|
||||
describes an approach that cannot satisfy it (e.g., a synchronous
|
||||
multi-service call chain with no caching), or explicitly contradicts
|
||||
it (e.g., "response times up to 2 seconds are acceptable").
|
||||
|
||||
**Risk**: The implementation will not meet requirements by design.
|
||||
This is not a gap but an active conflict.
|
||||
|
||||
**Severity guidance**: Critical when the violated constraint is
|
||||
safety-critical, regulatory, or a hard performance requirement. High
|
||||
for functional constraints.
|
||||
|
||||
### D7_ACCEPTANCE_CRITERIA_MISMATCH
|
||||
|
||||
A test case is linked to a requirement but does not actually verify the
|
||||
requirement's acceptance criteria.
|
||||
|
||||
**Pattern**: TC-NNN is mapped to REQ-XXX-NNN in the traceability matrix,
|
||||
but the test case's steps, inputs, or expected results do not correspond
|
||||
to the acceptance criteria defined for that requirement. The test may
|
||||
verify related but different behavior, or may be too coarse to confirm
|
||||
the specific criterion.
|
||||
|
||||
**Risk**: The traceability matrix shows coverage, but the coverage is
|
||||
illusory. The requirement appears tested but its actual acceptance
|
||||
criteria are not verified.
|
||||
|
||||
**Severity guidance**: High. This is more dangerous than D2 (untested
|
||||
requirement) because it creates a false sense of coverage.
|
||||
|
||||
## Code Compliance Labels
|
||||
|
||||
### D8_UNIMPLEMENTED_REQUIREMENT
|
||||
|
||||
A requirement exists in the requirements document but has no
|
||||
corresponding implementation in the source code.
|
||||
|
||||
**Pattern**: REQ-ID specifies a behavior, constraint, or capability.
|
||||
No function, module, class, or code path in the source implements
|
||||
or enforces this requirement.
|
||||
|
||||
**Risk**: The requirement was specified but never built. The system
|
||||
does not deliver this capability despite it being in the spec.
|
||||
|
||||
**Severity guidance**: Critical when the requirement is safety-critical
|
||||
or security-related. High for functional requirements. Medium for
|
||||
non-functional requirements that affect quality attributes.
|
||||
|
||||
### D9_UNDOCUMENTED_BEHAVIOR
|
||||
|
||||
The source code implements behavior that is not specified in any
|
||||
requirement or design document.
|
||||
|
||||
**Pattern**: A function, module, or code path implements meaningful
|
||||
behavior (not just infrastructure like logging or error handling)
|
||||
that does not trace to any REQ-ID in the requirements document or
|
||||
any section in the design document.
|
||||
|
||||
**Risk**: Scope creep in implementation — the code does more than
|
||||
was specified. The undocumented behavior may be intentional (a missing
|
||||
requirement) or accidental (a developer's assumption). Either way,
|
||||
it is untested against any specification.
|
||||
|
||||
**Severity guidance**: Medium when the behavior is benign feature
|
||||
logic. High when the behavior involves security, access control,
|
||||
data mutation, or external communication — undocumented behavior
|
||||
in these areas is a security concern.
|
||||
|
||||
### D10_CONSTRAINT_VIOLATION_IN_CODE
|
||||
|
||||
The source code violates a constraint stated in the requirements or
|
||||
design document.
|
||||
|
||||
**Pattern**: The requirements document states a constraint (e.g.,
|
||||
"MUST respond within 200ms", "MUST NOT store passwords in plaintext",
|
||||
"MUST use TLS 1.3 or later") and the source code demonstrably violates
|
||||
it — through algorithmic choice, missing implementation, or explicit
|
||||
contradiction.
|
||||
|
||||
**Risk**: The implementation will not meet requirements. Unlike D6
|
||||
(constraint violation in design), this is a concrete defect in code,
|
||||
not a planning gap.
|
||||
|
||||
**Severity guidance**: Critical when the violated constraint is
|
||||
safety-critical, security-related, or regulatory. High for performance
|
||||
or functional constraints. Assess based on the constraint itself,
|
||||
not the code's complexity.
|
||||
|
||||
## Test Compliance Labels
|
||||
|
||||
### D11_UNIMPLEMENTED_TEST_CASE
|
||||
|
||||
A test case is defined in the validation plan but has no corresponding
|
||||
automated test in the test code.
|
||||
|
||||
**Pattern**: TC-NNN is specified in the validation plan with steps,
|
||||
inputs, and expected results. No test function, test class, or test
|
||||
file in the test code implements this test case — either by name
|
||||
reference, by TC-NNN identifier, or by behavioral equivalence.
|
||||
|
||||
**Risk**: The validation plan claims coverage that does not exist in
|
||||
the automated test suite. The requirement linked to this test case
|
||||
is effectively untested in CI, even though the validation plan says
|
||||
it is covered.
|
||||
|
||||
**Severity guidance**: High when the linked requirement is
|
||||
safety-critical or security-related. Medium for functional
|
||||
requirements. Note: test cases classified as manual-only or deferred
|
||||
in the validation plan are excluded from D11 findings and reported
|
||||
only in the coverage summary.
|
||||
|
||||
### D12_UNTESTED_ACCEPTANCE_CRITERION
|
||||
|
||||
A test implementation exists for a test case, but it does not assert
|
||||
one or more acceptance criteria specified for the linked requirement.
|
||||
|
||||
**Pattern**: TC-NNN is implemented as an automated test. The linked
|
||||
requirement (REQ-XXX-NNN) has multiple acceptance criteria. The test
|
||||
implementation asserts some criteria but omits others — for example,
|
||||
it checks the happy-path output but does not verify error handling,
|
||||
boundary conditions, or timing constraints specified in the acceptance
|
||||
criteria.
|
||||
|
||||
**Risk**: The test passes but does not verify the full requirement.
|
||||
Defects in the untested acceptance criteria will not be caught by CI.
|
||||
This is the test-code equivalent of D7 (acceptance criteria mismatch
|
||||
in the validation plan) but at the implementation level.
|
||||
|
||||
**Severity guidance**: High when the missing criterion is a security
|
||||
or safety property. Medium for functional criteria. Assess based on
|
||||
what the missing criterion protects, not on the test's overall
|
||||
coverage.
|
||||
|
||||
### D13_ASSERTION_MISMATCH
|
||||
|
||||
A test implementation exists for a test case, but its assertions do
|
||||
not match the expected behavior specified in the validation plan.
|
||||
|
||||
**Pattern**: TC-NNN is implemented as an automated test. The test
|
||||
asserts different conditions, thresholds, or outcomes than what the
|
||||
validation plan specifies — for example, the plan says "verify
|
||||
response within 200ms" but the test asserts "response is not null",
|
||||
or the plan says "verify error code 403" but the test asserts "status
|
||||
is not 200".
|
||||
|
||||
**Risk**: The test passes but does not verify what the validation plan
|
||||
says it should. This creates illusory coverage — the traceability
|
||||
matrix shows the requirement as tested, but the actual test checks
|
||||
something different. More dangerous than D11 (missing test) because
|
||||
it is invisible without comparing test code to the validation plan.
|
||||
|
||||
**Severity guidance**: High. This is the most dangerous test
|
||||
compliance drift type because it creates false confidence. Severity
|
||||
should be assessed based on the gap between what is asserted and what
|
||||
should be asserted.
|
||||
|
||||
## Integration Compliance Labels
|
||||
|
||||
### D14_UNSPECIFIED_INTEGRATION_FLOW
|
||||
|
||||
A cross-component integration flow is described in the integration
|
||||
specification but is not reflected in one or more component specs.
|
||||
|
||||
**Pattern**: The integration spec describes an end-to-end flow that
|
||||
traverses components A → B → C. Component B's specification does not
|
||||
mention its role in this flow, does not describe receiving input from
|
||||
A, or does not describe producing output for C. The flow exists at
|
||||
the system level but has a gap at the component level.
|
||||
|
||||
**Risk**: The flow may be implemented by convention or tribal knowledge
|
||||
but is not contractually specified. Changes to component B may break
|
||||
the flow without any specification-level signal. Per-component audits
|
||||
will not detect this because no component's spec claims responsibility
|
||||
for the missing step.
|
||||
|
||||
**Severity guidance**: High when the flow is safety-critical, involves
|
||||
data integrity, or is a core user-facing workflow. Medium for
|
||||
operational or diagnostic flows. Assess based on what breaks if the
|
||||
gap causes a runtime failure.
|
||||
|
||||
### D15_INTERFACE_CONTRACT_MISMATCH
|
||||
|
||||
Two components describe the same interface differently in their
|
||||
respective specifications.
|
||||
|
||||
**Pattern**: Component A's spec says it produces output in format X
|
||||
with error codes {E1, E2}. Component B's spec says it consumes input
|
||||
in format Y with error codes {E2, E3}. The interface exists on both
|
||||
sides but the descriptions are incompatible — different data formats,
|
||||
different error sets, different sequencing assumptions, or different
|
||||
timing constraints.
|
||||
|
||||
**Risk**: Runtime failures at the integration boundary — data
|
||||
corruption, unhandled errors, deadlocks, or silent degradation.
|
||||
Per-component audits see each side as internally consistent; the
|
||||
mismatch is only visible when comparing both sides.
|
||||
|
||||
**Severity guidance**: Critical when the mismatch involves data
|
||||
integrity, security properties, or will cause deterministic runtime
|
||||
failure. High when it involves error handling or sequencing that may
|
||||
cause intermittent failures. Medium for cosmetic or logging
|
||||
differences that do not affect correctness.
|
||||
|
||||
### D16_UNTESTED_INTEGRATION_PATH
|
||||
|
||||
A cross-component integration flow or interface contract is specified
|
||||
but has no corresponding integration or end-to-end test.
|
||||
|
||||
**Pattern**: The integration spec describes flow F-NNN traversing
|
||||
components A → B → C. No integration test exercises this flow
|
||||
end-to-end. Individual component tests may test A's output and B's
|
||||
input separately, but no test verifies the handoff between them under
|
||||
realistic conditions.
|
||||
|
||||
**Risk**: Defects at integration boundaries will not be caught until
|
||||
production. Per-component test-compliance audits will show full
|
||||
coverage within each component, masking the integration gap. This is
|
||||
the integration-level equivalent of D11 (unimplemented test case).
|
||||
|
||||
**Severity guidance**: High when the flow is safety-critical or
|
||||
involves data that crosses trust boundaries. Medium for well-understood
|
||||
interfaces with stable contracts. Note: flows explicitly marked as
|
||||
"manual integration test" or "deferred" in the integration spec are
|
||||
excluded from D16 findings and reported only in the coverage summary.
|
||||
|
||||
## Ranking Criteria
|
||||
|
||||
Within a given severity level, order findings by impact on specification
|
||||
integrity:
|
||||
|
||||
1. **Highest risk**: D6 (constraint violation in design), D7 (illusory
|
||||
test coverage), D10 (constraint violation in code), D13
|
||||
(assertion mismatch), and D15 (interface contract mismatch) —
|
||||
these indicate active conflicts between artifacts.
|
||||
2. **High risk**: D2 (untested requirement), D5 (assumption drift),
|
||||
D8 (unimplemented requirement), D12 (untested acceptance
|
||||
criterion), and D14 (unspecified integration flow) — these
|
||||
indicate silent gaps that will surface late.
|
||||
3. **Medium risk**: D1 (untraced requirement), D3 (orphaned design),
|
||||
D9 (undocumented behavior), D11 (unimplemented test case), and
|
||||
D16 (untested integration path) — these indicate incomplete
|
||||
traceability that needs human resolution.
|
||||
4. **Lowest risk**: D4 (orphaned test case) — effort misdirection but
|
||||
no safety or correctness impact.
|
||||
|
||||
## Usage
|
||||
|
||||
In findings, reference labels as:
|
||||
|
||||
```
|
||||
[DRIFT: D2_UNTESTED_REQUIREMENT]
|
||||
Requirement: REQ-SEC-003 (requirements doc, section 4.2)
|
||||
Evidence: REQ-SEC-003 does not appear in the traceability matrix
|
||||
(validation plan, section 4). No test case references this REQ-ID.
|
||||
Impact: The encryption-at-rest requirement will not be verified.
|
||||
```
|
||||
|
||||
<!-- END PromptKit base -->
|
||||
|
||||
---
|
||||
|
||||
<!-- BEGIN ocserv extensions -->
|
||||
|
||||
## ocserv-Specific Extensions
|
||||
|
||||
ocserv's normative specification is `doc/requirements/` (REQ-*/AC-*/OC-*
|
||||
entries; see `doc/requirements/README.md` for the ID scheme and document
|
||||
map). It has no separate design document or validation plan in the sense
|
||||
this taxonomy assumes — `doc/design.md`, `doc/ocserv.8.md`, and
|
||||
`sample.config` play the design-document role, and `tests/` (registered in
|
||||
`tests/meson.build`) plays the validation-plan role. The primary label
|
||||
group used in ocserv reviews is **D8–D10 (Code Compliance)**, via
|
||||
`protocols/code-compliance-audit.md`. D1–D7 apply if auditing
|
||||
`doc/requirements/` documents against each other or against `doc/design.md`
|
||||
(e.g. with `requirements-reconciliation.md`); D11–D13 apply if auditing
|
||||
`tests/` against the acceptance criteria in `doc/requirements/`.
|
||||
|
||||
### D8_UNIMPLEMENTED_REQUIREMENT — ocserv examples
|
||||
|
||||
- A `doc/sample.config` option is documented with an acceptance criterion
|
||||
(e.g., a bad value must be rejected at vhost scope) but `src/config.c` /
|
||||
`src/subconfig.c` only enforces it globally, or not at all.
|
||||
- A `SEC`/`AUTH`-tagged negative acceptance criterion ("MUST reject a
|
||||
replayed SID", "MUST reject a tampered cookie") has no corresponding
|
||||
rejection path in `sec-mod.c` or the worker's cookie-handling code.
|
||||
|
||||
### D9_UNDOCUMENTED_BEHAVIOR — ocserv examples
|
||||
|
||||
- A new config option added to `common-config.h` / `subconfig.c` with no
|
||||
`REQ-CFG-*` entry and no mention in `sample.config` / `doc/ocserv.8.md`.
|
||||
- A new field added to `src/ipc.proto` or `src/ctl.proto` with no
|
||||
acceptance criteria in `doc/requirements/` describing what a receiving
|
||||
process should do with it, or a validation rule the receiver silently
|
||||
omits.
|
||||
- Classify High severity whenever the undocumented behavior touches
|
||||
`SEC`, `AUTH`, or `IPC` category areas — per this taxonomy's own
|
||||
guidance, undocumented behavior in those areas is a security concern,
|
||||
not just a documentation gap.
|
||||
|
||||
### D10_CONSTRAINT_VIOLATION_IN_CODE — ocserv examples
|
||||
|
||||
- Code that stores, logs, or transmits a credential or session secret in
|
||||
a way a `SEC`/`AUTH` requirement's acceptance criteria forbid.
|
||||
- Code that silently collapses the main/sec-mod/worker privilege boundary
|
||||
(AGENTS.md, "Architecture — The Invariant You Must Not Violate") —
|
||||
always Critical, and requires explicit maintainer acknowledgment, not
|
||||
just a fix.
|
||||
- A patch that changes behavior an existing `REQ-*` describes without
|
||||
updating that requirement first (AGENTS.md's Requirements-First
|
||||
Workflow) — the code now contradicts requirement text that still
|
||||
stands unmodified. Classify as D10 with the specific REQ-ID as
|
||||
evidence, per `protocols/code-compliance-audit.md` Phase 1.5.
|
||||
- A violation of a canonical technology choice (REQ-GEN-TECH-001..005):
|
||||
non-talloc allocation, direct OpenSSL use, hand-edited `*.pb-c.c`/`.h`,
|
||||
a new structured config format outside INI, or an unapproved new
|
||||
external dependency.
|
||||
|
||||
<!-- END ocserv extensions -->
|
||||
@@ -140,3 +140,60 @@ sources:
|
||||
incidental** — they are `SEC` requirements, always essential.
|
||||
- **IPC acceptance criteria must cite protobuf field names** from
|
||||
`src/ipc.proto` / `src/ctl.proto`, not vague descriptions.
|
||||
- **Normative language uses RFC 2119 keywords only** — MUST, MUST NOT,
|
||||
SHALL, SHALL NOT, SHOULD, SHOULD NOT, MAY, REQUIRED, RECOMMENDED,
|
||||
OPTIONAL. Informal equivalents ("needs to," "has to," "can," "will")
|
||||
MUST NOT be used to express a normative obligation in requirement prose.
|
||||
|
||||
## Glossary
|
||||
|
||||
Canonical definitions for terms that are ambiguous across ocserv's own
|
||||
documentation and code. A requirement MUST NOT redefine a term listed here;
|
||||
if a requirement needs a meaning not covered below, add it here first, then
|
||||
cite it. New entries are added the first time a term is flagged during the
|
||||
Ambiguity Detection phase of `requirements-elicitation.md` or by
|
||||
`contrib/ai/protocols/prompt-determinism-analysis.md`.
|
||||
|
||||
| Term | Definition |
|
||||
|------|------------|
|
||||
| session | Disambiguate per use: a **TLS/DTLS session** (GnuTLS session object, resumable via session tickets); a **VPN session** (the authenticated user's SID and lease, spanning reconnects/roaming); or a **PAM session** (`pam_open_session`/`pam_close_session`). A requirement using "session" unqualified MUST specify which. |
|
||||
| connection | One TCP/UDP socket-level attachment to a single worker process, bounded by that socket's lifetime. Distinct from a **session** (above): a VPN session can span multiple connections via cookie resumption (`doc/design.md`, "IPC Communication for SID assignment": "client/worker may disconnect and reconnect, using SID cookie to resume the authenticated session"). |
|
||||
| secure | Not a standalone property. Always state the concrete guarantee meant: encrypted transport, authenticated peer, integrity-protected, or a specific cipher/version floor (e.g. "TLS 1.2 or later"). |
|
||||
| reload | A `SIGHUP`-triggered live config reload (main/sec-mod only, requires procfs — REQ-GEN-COMPAT-001), distinct from a full process restart. Check `doc/sample.config`'s `[reload]`/`[not-reloadable]` annotation for the specific option before using this term. |
|
||||
| worker / client | "Worker" is always the ocserv worker process; "client" is always the remote OpenConnect/AnyConnect endpoint. Never use one to mean the other, even informally — they sit on opposite sides of the privilege/trust boundary. |
|
||||
| security module (sec-mod) | The `sec-mod` process. `doc/design.md` uses "security module" and "sec-mod" interchangeably for the same root process that holds private keys, PAM/RADIUS state, and session/SID state (`src/sec-mod*.c`). Not an external HSM or a generic security concept. |
|
||||
| cookie | The SID-bound authentication ticket issued by sec-mod on successful auth, forwarded via `AUTH_COOKIE_REQ`/`AUTH_COOKIE_REP`, valid for `cookie-timeout`, and used to resume a session across reconnects. `doc/design.md` also calls this a "ticket." Not a generic HTTP `Set-Cookie` value, even though it travels over HTTPS. |
|
||||
| SID vs. safe_id | The **SID** is the internal session identifier assigned by sec-mod on `SEC_AUTH_INIT` and used directly in IPC (`src/ipc.proto`, `src/ctl.proto`) and as cookie material; treat it as sensitive. **safe_id** is `base64(SHA1(SID))` (`calc_safe_id()`, `src/common/common.c`), a one-way, non-reversible derivation used wherever a session must be referenced externally without exposing the SID: `occtl` session listing/termination, logs, and RADIUS accounting (`PW_ACCT_SESSION_ID`, `src/acct/radius.c`). A requirement or acceptance criterion MUST say which one it means — they are not interchangeable, and safe_id cannot be reversed to recover the SID. |
|
||||
| accounting | Post-authentication usage/session data reporting (RADIUS accounting, `src/acct/`) forwarded by sec-mod. Distinct from PAM **account management** (`pam_acct_mgmt`), which is an authorization check performed during login, not usage reporting — `doc/design.md`'s "Gatekeeper for accounting information keeping and reporting" refers to the former. |
|
||||
|
||||
## Dependencies
|
||||
|
||||
External/optional build dependencies that make some requirements
|
||||
conditionally inapplicable. Format:
|
||||
|
||||
```markdown
|
||||
### DEP-<NNN>
|
||||
**Dependency:** <build option / library>
|
||||
**Required by:** <REQ-ID(s) or document>
|
||||
**Impact if unavailable:** <what becomes inapplicable or degraded>
|
||||
```
|
||||
|
||||
### DEP-001
|
||||
**Dependency:** seccomp (`-Dseccomp`, auto-detected)
|
||||
**Required by:** REQ-GEN-SEC-002(d)
|
||||
**Impact if unavailable:** the worker runs without syscall confinement;
|
||||
the privilege-boundary requirement's process-separation intent still
|
||||
holds, but this specific enforcement mechanism is absent and MUST be
|
||||
called out in deployment documentation as reduced defense-in-depth.
|
||||
|
||||
### DEP-002
|
||||
**Dependency:** PAM (`-Dpam`, auto-detected)
|
||||
**Required by:** PAM-backed entries in `internal/authentication.md`
|
||||
**Impact if unavailable:** those entries are not applicable to the build;
|
||||
authentication falls back to other configured modules (plain, RADIUS,
|
||||
GSSAPI, OIDC).
|
||||
|
||||
### DEP-003
|
||||
**Dependency:** RADIUS (`-Dradius`, auto-detected)
|
||||
**Required by:** RADIUS-backed entries in `internal/authentication.md`
|
||||
**Impact if unavailable:** those entries are not applicable to the build.
|
||||
|
||||
Reference in New Issue
Block a user