mirror of
https://gitlab.com/openconnect/ocserv.git
synced 2026-08-08 09:21:48 +08:00
ocserv-core-dev: added root-cause-analysis protocol from promptkit
Signed-off-by: Nikos Mavrogiannopoulos <n.mavrogiannopoulos@gmail.com>
This commit is contained in:
@@ -194,42 +194,13 @@ New auth method ↔ `auth_mod_st` vtable in `src/sec-mod-auth.h` ↔ registratio
|
||||
|
||||
## Protocol: Root Cause Analysis
|
||||
|
||||
Apply when investigating a defect or unexpected behavior. The goal is the
|
||||
**fundamental cause**, not the proximate trigger.
|
||||
|
||||
**Phase 1 — Characterize the symptom precisely.**
|
||||
- What is the observed behavior vs. the expected behavior?
|
||||
- What does `doc/ocserv.8.md` or `doc/sample.config` document as the expected
|
||||
behavior for this option or feature? If the observed behavior contradicts the
|
||||
documentation, that divergence is itself part of the bug characterization.
|
||||
- Which process emitted the error (main / sec-mod / worker)? Use `VERBOSE=1` to
|
||||
identify the source.
|
||||
- Is it deterministic or intermittent? If intermittent, does it correlate with
|
||||
load, timing, or specific client behavior?
|
||||
- What changed recently? (code, config, dependencies)
|
||||
|
||||
**Phase 2 — Generate hypotheses (at least 3 before investigating any).**
|
||||
For each: state the hypothesis, what evidence would confirm it, what would refute
|
||||
it, and a plausibility rating (High / Medium / Low). Include at least one
|
||||
non-obvious hypothesis (IPC race, config reload timing, seccomp filter gap,
|
||||
allocator mismatch across process boundary).
|
||||
|
||||
**Phase 3 — Eliminate.**
|
||||
For each hypothesis, identify the minimal investigation needed (specific file,
|
||||
log line, or code path). Classify: CONFIRMED / ELIMINATED / INCONCLUSIVE.
|
||||
Do not anchor on the first plausible hypothesis.
|
||||
|
||||
**Phase 4 — Distinguish root from proximate cause.**
|
||||
- Proximate: "null pointer dereference in worker at `worker-http.c:312`."
|
||||
- Root: "sec-mod returned a zero-length SID when config reload raced with
|
||||
`SEC_AUTH_INIT`, leaving the worker's session pointer uninitialized."
|
||||
- Ask: if we fix only the proximate cause, will the root cause produce other
|
||||
failures? If yes, the fix is incomplete.
|
||||
|
||||
**Phase 5 — Remediation.**
|
||||
Propose a fix for the root cause. Identify secondary fixes (assertions, improved
|
||||
error messages, tests that would have caught this). Assess the risk of the fix:
|
||||
could it introduce new failures in adjacent code paths?
|
||||
Load and follow `contrib/ai/protocols/root-cause-analysis.md` for the full
|
||||
analysis protocol. The ocserv-specific guidance — process identification
|
||||
(main / sec-mod / worker), documentation sources (`doc/ocserv.8.md`,
|
||||
`doc/sample.config`, `doc/design.md`), canonical ocserv hypotheses (IPC race,
|
||||
config-reload timing, seccomp filter gap, allocator mismatch, SID/cookie
|
||||
lifecycle), IPC call-chain tracing, and remediation conventions — is in the
|
||||
extension section of that file.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,270 @@
|
||||
<!-- SPDX-License-Identifier: MIT -->
|
||||
<!-- Copyright (c) PromptKit Contributors -->
|
||||
|
||||
---
|
||||
name: root-cause-analysis
|
||||
type: reasoning
|
||||
description: >
|
||||
Structured reasoning protocol for tracing symptoms to root causes.
|
||||
Applies systematic hypothesis generation, evidence evaluation,
|
||||
and elimination. Language-agnostic.
|
||||
applicable_to:
|
||||
- investigate-bug
|
||||
- investigate-trace
|
||||
- root-cause-ci-failure
|
||||
---
|
||||
|
||||
# Protocol: Root Cause Analysis
|
||||
|
||||
Apply this protocol when investigating a defect, failure, or unexpected behavior.
|
||||
The goal is to trace from **observed symptoms** to the **fundamental cause** —
|
||||
not just the proximate trigger.
|
||||
|
||||
## Phase 1: Symptom Characterization
|
||||
|
||||
1. **Describe the symptom precisely**:
|
||||
- What is the observed behavior?
|
||||
- What is the expected behavior?
|
||||
- Under what conditions does it occur? (inputs, timing, environment, load)
|
||||
- Is it deterministic or intermittent?
|
||||
2. **Establish the timeline**:
|
||||
- When was it first observed?
|
||||
- What changed recently? (code, configuration, dependencies, infrastructure)
|
||||
- Has it ever worked correctly? When did it stop?
|
||||
3. **Determine the blast radius**:
|
||||
- What is affected? (single user, all users, specific configurations)
|
||||
- What is NOT affected? (this constrains the hypothesis space)
|
||||
|
||||
## Phase 2: Hypothesis Generation
|
||||
|
||||
Generate a **ranked list of hypotheses**, ordered by likelihood. For each:
|
||||
|
||||
1. State the hypothesis clearly: "The root cause is X because Y."
|
||||
2. State what **evidence would confirm** the hypothesis.
|
||||
3. State what **evidence would refute** the hypothesis.
|
||||
4. Rate plausibility: High / Medium / Low — with reasoning.
|
||||
|
||||
Rules:
|
||||
- Generate at least 3 hypotheses before investigating any of them.
|
||||
- Include at least one "non-obvious" hypothesis (environment, timing, config,
|
||||
upstream dependency, data corruption).
|
||||
- Do NOT anchor on the first plausible hypothesis.
|
||||
|
||||
## Phase 3: Evidence Gathering and Elimination
|
||||
|
||||
For each hypothesis, starting with the most plausible:
|
||||
|
||||
1. Identify the **minimal investigation** needed to confirm or eliminate it.
|
||||
2. Examine the evidence:
|
||||
- Does the code/config/log support or contradict the hypothesis?
|
||||
- Are there alternative explanations for the same evidence?
|
||||
3. Classify the hypothesis:
|
||||
- **CONFIRMED**: Strong evidence supports it; no contradicting evidence.
|
||||
- **ELIMINATED**: Evidence directly contradicts it.
|
||||
- **INCONCLUSIVE**: Evidence is insufficient; state what is needed.
|
||||
|
||||
## Phase 3a: Iterative Deepening
|
||||
|
||||
Investigation MUST proceed in layers of increasing resolution. Each layer
|
||||
informs the next — do NOT skip layers or jump directly to deep analysis.
|
||||
|
||||
1. **Broad survey**: Identify top contributors at the coarsest granularity
|
||||
(e.g., by process, module, subsystem, or component). Rank by impact.
|
||||
2. **Attribution**: For the top 5–10 contributors, break down by the next
|
||||
level of detail (e.g., by module within a process, by function within
|
||||
a module, by allocation site within a function).
|
||||
3. **Deep analysis**: For the top contributors at the attribution level,
|
||||
obtain the most detailed evidence available (e.g., call stacks, data
|
||||
flow traces, lock contention chains, allocation histories). Call stacks
|
||||
and execution traces reveal *why* something is happening — module-level
|
||||
data only reveals *where*.
|
||||
4. **Cross-component tracing**: Identify causal chains that span component
|
||||
or process boundaries (see Phase 4a).
|
||||
|
||||
Do NOT write the final report until layer 3 is complete for the top
|
||||
contributors, up to 5, using the most detailed evidence available.
|
||||
If fewer than 5 contributors exist, analyze all of them. If available
|
||||
evidence does not support layer-3 completion for some contributors, you
|
||||
MAY proceed to the final report only if you explicitly document the
|
||||
limitation, identify which contributors remain inconclusive, and state
|
||||
what additional evidence would be needed. Premature reporting without
|
||||
this disclosure produces surface-level findings that miss the actual
|
||||
root cause.
|
||||
|
||||
## Phase 4: Root Cause Identification
|
||||
|
||||
1. Distinguish between the **root cause** (fundamental defect) and the
|
||||
**proximate cause** (immediate trigger).
|
||||
- Example: The proximate cause is "null pointer dereference on line 42."
|
||||
The root cause is "the initialization function silently fails when
|
||||
the config file is missing, leaving the pointer uninitialized."
|
||||
2. Trace the **causal chain** from root cause to symptom.
|
||||
3. Ask: "If we fix only the proximate cause, will the root cause
|
||||
produce other failures?" If yes, the fix is incomplete.
|
||||
|
||||
## Phase 4a: Cross-Component Causal Chains
|
||||
|
||||
When the investigation involves multiple components, processes, or
|
||||
subsystems, trace causal chains across boundaries:
|
||||
|
||||
1. **Identify trigger-response pairs**: Does activity in component A
|
||||
cause work in component B? For example, a file write by one process
|
||||
may trigger scanning by an antivirus service, which triggers hashing
|
||||
by an EDR agent, which triggers network inspection by another service.
|
||||
2. **Map the amplification cascade**: A single action may fan out into
|
||||
disproportionate downstream work. Document the full chain:
|
||||
`Trigger → Reactor₁ → Reactor₂ → ... → Observed symptom`.
|
||||
3. **Quantify amplification**: For each link in the chain, estimate the
|
||||
cost ratio (e.g., "1 file write triggers 3 scan operations, each
|
||||
consuming 50ms of CPU"). The amplification factor often explains why
|
||||
a seemingly minor activity produces outsized impact.
|
||||
4. **Identify the leverage point**: The most effective fix targets the
|
||||
link in the chain with the highest amplification factor, not
|
||||
necessarily the initial trigger or the final symptom.
|
||||
|
||||
Skip this phase when the investigation is confined to a single component
|
||||
with no cross-boundary interactions.
|
||||
|
||||
## Phase 5: Remediation
|
||||
|
||||
1. Propose a fix for the **root cause**, not just the symptom.
|
||||
2. Identify **secondary fixes** needed to prevent recurrence:
|
||||
- Assertions or precondition checks
|
||||
- Improved error handling
|
||||
- Logging or monitoring
|
||||
- Tests that would have caught this
|
||||
3. Assess the **risk of the fix**: could it introduce new issues?
|
||||
|
||||
<!-- BEGIN ocserv extensions -->
|
||||
|
||||
## ocserv-Specific Extensions
|
||||
|
||||
The sections below extend the generic protocol with ocserv's three-process
|
||||
architecture, documentation sources, and diagnostic conventions. Apply these
|
||||
alongside the base phases above — they do not replace them.
|
||||
|
||||
### Phase 1 — Symptom Characterization (ocserv)
|
||||
|
||||
In addition to the generic checklist:
|
||||
|
||||
- **Consult documentation first.** Check `doc/ocserv.8.md` and
|
||||
`doc/sample.config` for the documented behavior of the option or feature
|
||||
under investigation. If the observed behavior contradicts the documentation,
|
||||
that divergence is itself part of the bug characterization — note it
|
||||
explicitly. The code must match the documentation, or both must be updated
|
||||
together.
|
||||
- **Identify the originating process.** ocserv has three processes with
|
||||
distinct responsibilities:
|
||||
- **main** (`main.c`, `main-*.c`) — TCP/UDP listeners, TUN devices, IP
|
||||
allocation, process lifecycle.
|
||||
- **sec-mod** (`sec-mod.c`, `sec-mod-*.c`) — authentication, private keys,
|
||||
session state, PAM, accounting.
|
||||
- **worker** (`worker.c`, `worker-*.c`) — per-client TLS/DTLS, VPN traffic.
|
||||
|
||||
Run `VERBOSE=1 ./tests/<test>` (or start the server with verbose logging)
|
||||
and identify which process emitted the error. The originating process
|
||||
directly constrains the hypothesis space.
|
||||
- **Intermittent failures.** If the failure is intermittent, check whether
|
||||
it correlates with:
|
||||
- Concurrent client connections (IPC contention).
|
||||
- Config reload timing (`SIGHUP` during active sessions).
|
||||
- Specific authentication backends (PAM, RADIUS, OIDC).
|
||||
- Client reconnection / cookie resumption flows.
|
||||
|
||||
### Phase 2 — Hypothesis Generation (ocserv)
|
||||
|
||||
In addition to the generic "non-obvious hypothesis" requirement, always
|
||||
consider at least one hypothesis from each of these ocserv-specific
|
||||
categories before investigating any hypothesis:
|
||||
|
||||
- **IPC race**: A message arrived at main or sec-mod while another request
|
||||
was in-flight; the response was matched to the wrong caller.
|
||||
- **Config reload timing**: `SIGHUP` triggered a config reload that raced
|
||||
with an active authentication or session establishment.
|
||||
- **seccomp filter gap**: A syscall needed by a new code path is missing
|
||||
from the worker's seccomp whitelist, causing the worker to be killed by
|
||||
the kernel.
|
||||
- **Allocator mismatch across process boundary**: A talloc pointer was
|
||||
embedded in an IPC message and dereferenced in a different process where
|
||||
it is meaningless; or `gnutls_malloc` / `gnutls_free` was used for memory
|
||||
that outlives the GnuTLS context.
|
||||
- **SID / cookie lifecycle**: The session ID was invalidated (e.g., by a
|
||||
sec-mod restart or cookie expiry) before the worker consumed it, leaving
|
||||
a dangling reference.
|
||||
|
||||
### Phase 3a — Iterative Deepening (ocserv)
|
||||
|
||||
Apply the generic layers with these ocserv-specific granularities:
|
||||
|
||||
| Layer | ocserv granularity |
|
||||
|-------|--------------------|
|
||||
| Broad survey | Which process (main / sec-mod / worker)? |
|
||||
| Attribution | Which module / source file within that process? |
|
||||
| Deep analysis | Which function, IPC message field, or auth backend call? |
|
||||
| Cross-component | Trace the IPC call chain (e.g., `SEC_AUTH_INIT` → `SEC_AUTH_REP` → `AUTH_COOKIE_REQ` → ...) |
|
||||
|
||||
For IPC-level bugs, read `doc/design.md` for the authoritative message
|
||||
sequence before tracing code paths.
|
||||
|
||||
### Phase 4 — Root vs. Proximate Cause (ocserv)
|
||||
|
||||
ocserv-specific examples of the root/proximate distinction:
|
||||
|
||||
- **Proximate**: "null pointer dereference in worker at `worker-http.c:312`."
|
||||
**Root**: "sec-mod returned a zero-length SID when a config reload raced
|
||||
with `SEC_AUTH_INIT`, leaving the worker's session pointer uninitialized."
|
||||
|
||||
- **Proximate**: "worker killed with SIGKILL during a new feature's first
|
||||
system call."
|
||||
**Root**: "the new syscall is not in the seccomp whitelist; the kernel
|
||||
kills workers that attempt it."
|
||||
|
||||
- **Proximate**: "`talloc_free` crashes on a session struct."
|
||||
**Root**: "ownership of the struct was transferred across an IPC boundary
|
||||
without re-parenting under the receiving process's talloc context."
|
||||
|
||||
After identifying the root cause, ask: if we fix only the proximate cause,
|
||||
will the root cause produce failures elsewhere in the IPC call chain?
|
||||
|
||||
### Phase 4a — Cross-Component Causal Chains (ocserv)
|
||||
|
||||
ocserv's IPC topology makes cross-component tracing essential. The canonical
|
||||
chain is:
|
||||
|
||||
```
|
||||
worker → sec-mod: SEC_AUTH_INIT
|
||||
sec-mod → worker: SEC_AUTH_REP (SID assigned)
|
||||
worker → sec-mod: SEC_AUTH_CONT (SID + credential)
|
||||
sec-mod → worker: SEC_AUTH_REP (OK / FAILED)
|
||||
worker → main: AUTH_COOKIE_REQ (SID)
|
||||
main → sec-mod: SECM_SESSION_OPEN (SID)
|
||||
sec-mod → main: SECM_SESSION_REPLY (user config)
|
||||
main → worker: AUTH_COOKIE_REP (TUN device, config)
|
||||
```
|
||||
|
||||
When tracing a causal chain, map each IPC message hop and identify at
|
||||
which hop the invariant breaks. The leverage point is usually the first
|
||||
hop where a precondition is violated, not the hop where the crash or
|
||||
error is ultimately observed.
|
||||
|
||||
### Phase 5 — Remediation (ocserv)
|
||||
|
||||
Secondary fixes to consider for every ocserv bug:
|
||||
|
||||
- **Assertions**: Add `assert()` or a log-and-return guard at the point
|
||||
where the violated precondition should have been caught.
|
||||
- **Error messages**: ocserv's `mslog()` / `oclog()` / `seclog()` messages
|
||||
should identify the process, the IPC message type, and the session ID (SID)
|
||||
so that field-reported bugs can be diagnosed from logs alone.
|
||||
- **Tests**: Write a regression test in `tests/` that reproduces the bug.
|
||||
Register it in `tests/meson.build`. For IPC-race bugs, a test that
|
||||
exercises concurrent connections under load is preferred.
|
||||
- **Documentation**: If the bug arose because the documented behavior was
|
||||
ambiguous or incorrect, update `doc/ocserv.8.md` and `doc/sample.config`
|
||||
alongside the code fix.
|
||||
|
||||
Assess fix risk with ocserv's privilege model in mind: a change in the
|
||||
worker that looks local may require a corresponding IPC or seccomp change;
|
||||
a change in sec-mod's session handling may affect all auth backends.
|
||||
|
||||
<!-- END ocserv extensions -->
|
||||
Reference in New Issue
Block a user