mirror of
https://gitlab.com/openconnect/ocserv.git
synced 2026-08-08 09:21:48 +08:00
ocserv-core-dev: added root-cause-analysis protocol from promptkit
Signed-off-by: Nikos Mavrogiannopoulos <n.mavrogiannopoulos@gmail.com>
This commit is contained in:
@@ -194,42 +194,13 @@ New auth method ↔ `auth_mod_st` vtable in `src/sec-mod-auth.h` ↔ registratio
|
|||||||
|
|
||||||
## Protocol: Root Cause Analysis
|
## Protocol: Root Cause Analysis
|
||||||
|
|
||||||
Apply when investigating a defect or unexpected behavior. The goal is the
|
Load and follow `contrib/ai/protocols/root-cause-analysis.md` for the full
|
||||||
**fundamental cause**, not the proximate trigger.
|
analysis protocol. The ocserv-specific guidance — process identification
|
||||||
|
(main / sec-mod / worker), documentation sources (`doc/ocserv.8.md`,
|
||||||
**Phase 1 — Characterize the symptom precisely.**
|
`doc/sample.config`, `doc/design.md`), canonical ocserv hypotheses (IPC race,
|
||||||
- What is the observed behavior vs. the expected behavior?
|
config-reload timing, seccomp filter gap, allocator mismatch, SID/cookie
|
||||||
- What does `doc/ocserv.8.md` or `doc/sample.config` document as the expected
|
lifecycle), IPC call-chain tracing, and remediation conventions — is in the
|
||||||
behavior for this option or feature? If the observed behavior contradicts the
|
extension section of that file.
|
||||||
documentation, that divergence is itself part of the bug characterization.
|
|
||||||
- Which process emitted the error (main / sec-mod / worker)? Use `VERBOSE=1` to
|
|
||||||
identify the source.
|
|
||||||
- Is it deterministic or intermittent? If intermittent, does it correlate with
|
|
||||||
load, timing, or specific client behavior?
|
|
||||||
- What changed recently? (code, config, dependencies)
|
|
||||||
|
|
||||||
**Phase 2 — Generate hypotheses (at least 3 before investigating any).**
|
|
||||||
For each: state the hypothesis, what evidence would confirm it, what would refute
|
|
||||||
it, and a plausibility rating (High / Medium / Low). Include at least one
|
|
||||||
non-obvious hypothesis (IPC race, config reload timing, seccomp filter gap,
|
|
||||||
allocator mismatch across process boundary).
|
|
||||||
|
|
||||||
**Phase 3 — Eliminate.**
|
|
||||||
For each hypothesis, identify the minimal investigation needed (specific file,
|
|
||||||
log line, or code path). Classify: CONFIRMED / ELIMINATED / INCONCLUSIVE.
|
|
||||||
Do not anchor on the first plausible hypothesis.
|
|
||||||
|
|
||||||
**Phase 4 — Distinguish root from proximate cause.**
|
|
||||||
- Proximate: "null pointer dereference in worker at `worker-http.c:312`."
|
|
||||||
- Root: "sec-mod returned a zero-length SID when config reload raced with
|
|
||||||
`SEC_AUTH_INIT`, leaving the worker's session pointer uninitialized."
|
|
||||||
- Ask: if we fix only the proximate cause, will the root cause produce other
|
|
||||||
failures? If yes, the fix is incomplete.
|
|
||||||
|
|
||||||
**Phase 5 — Remediation.**
|
|
||||||
Propose a fix for the root cause. Identify secondary fixes (assertions, improved
|
|
||||||
error messages, tests that would have caught this). Assess the risk of the fix:
|
|
||||||
could it introduce new failures in adjacent code paths?
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,270 @@
|
|||||||
|
<!-- SPDX-License-Identifier: MIT -->
|
||||||
|
<!-- Copyright (c) PromptKit Contributors -->
|
||||||
|
|
||||||
|
---
|
||||||
|
name: root-cause-analysis
|
||||||
|
type: reasoning
|
||||||
|
description: >
|
||||||
|
Structured reasoning protocol for tracing symptoms to root causes.
|
||||||
|
Applies systematic hypothesis generation, evidence evaluation,
|
||||||
|
and elimination. Language-agnostic.
|
||||||
|
applicable_to:
|
||||||
|
- investigate-bug
|
||||||
|
- investigate-trace
|
||||||
|
- root-cause-ci-failure
|
||||||
|
---
|
||||||
|
|
||||||
|
# Protocol: Root Cause Analysis
|
||||||
|
|
||||||
|
Apply this protocol when investigating a defect, failure, or unexpected behavior.
|
||||||
|
The goal is to trace from **observed symptoms** to the **fundamental cause** —
|
||||||
|
not just the proximate trigger.
|
||||||
|
|
||||||
|
## Phase 1: Symptom Characterization
|
||||||
|
|
||||||
|
1. **Describe the symptom precisely**:
|
||||||
|
- What is the observed behavior?
|
||||||
|
- What is the expected behavior?
|
||||||
|
- Under what conditions does it occur? (inputs, timing, environment, load)
|
||||||
|
- Is it deterministic or intermittent?
|
||||||
|
2. **Establish the timeline**:
|
||||||
|
- When was it first observed?
|
||||||
|
- What changed recently? (code, configuration, dependencies, infrastructure)
|
||||||
|
- Has it ever worked correctly? When did it stop?
|
||||||
|
3. **Determine the blast radius**:
|
||||||
|
- What is affected? (single user, all users, specific configurations)
|
||||||
|
- What is NOT affected? (this constrains the hypothesis space)
|
||||||
|
|
||||||
|
## Phase 2: Hypothesis Generation
|
||||||
|
|
||||||
|
Generate a **ranked list of hypotheses**, ordered by likelihood. For each:
|
||||||
|
|
||||||
|
1. State the hypothesis clearly: "The root cause is X because Y."
|
||||||
|
2. State what **evidence would confirm** the hypothesis.
|
||||||
|
3. State what **evidence would refute** the hypothesis.
|
||||||
|
4. Rate plausibility: High / Medium / Low — with reasoning.
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- Generate at least 3 hypotheses before investigating any of them.
|
||||||
|
- Include at least one "non-obvious" hypothesis (environment, timing, config,
|
||||||
|
upstream dependency, data corruption).
|
||||||
|
- Do NOT anchor on the first plausible hypothesis.
|
||||||
|
|
||||||
|
## Phase 3: Evidence Gathering and Elimination
|
||||||
|
|
||||||
|
For each hypothesis, starting with the most plausible:
|
||||||
|
|
||||||
|
1. Identify the **minimal investigation** needed to confirm or eliminate it.
|
||||||
|
2. Examine the evidence:
|
||||||
|
- Does the code/config/log support or contradict the hypothesis?
|
||||||
|
- Are there alternative explanations for the same evidence?
|
||||||
|
3. Classify the hypothesis:
|
||||||
|
- **CONFIRMED**: Strong evidence supports it; no contradicting evidence.
|
||||||
|
- **ELIMINATED**: Evidence directly contradicts it.
|
||||||
|
- **INCONCLUSIVE**: Evidence is insufficient; state what is needed.
|
||||||
|
|
||||||
|
## Phase 3a: Iterative Deepening
|
||||||
|
|
||||||
|
Investigation MUST proceed in layers of increasing resolution. Each layer
|
||||||
|
informs the next — do NOT skip layers or jump directly to deep analysis.
|
||||||
|
|
||||||
|
1. **Broad survey**: Identify top contributors at the coarsest granularity
|
||||||
|
(e.g., by process, module, subsystem, or component). Rank by impact.
|
||||||
|
2. **Attribution**: For the top 5–10 contributors, break down by the next
|
||||||
|
level of detail (e.g., by module within a process, by function within
|
||||||
|
a module, by allocation site within a function).
|
||||||
|
3. **Deep analysis**: For the top contributors at the attribution level,
|
||||||
|
obtain the most detailed evidence available (e.g., call stacks, data
|
||||||
|
flow traces, lock contention chains, allocation histories). Call stacks
|
||||||
|
and execution traces reveal *why* something is happening — module-level
|
||||||
|
data only reveals *where*.
|
||||||
|
4. **Cross-component tracing**: Identify causal chains that span component
|
||||||
|
or process boundaries (see Phase 4a).
|
||||||
|
|
||||||
|
Do NOT write the final report until layer 3 is complete for the top
|
||||||
|
contributors, up to 5, using the most detailed evidence available.
|
||||||
|
If fewer than 5 contributors exist, analyze all of them. If available
|
||||||
|
evidence does not support layer-3 completion for some contributors, you
|
||||||
|
MAY proceed to the final report only if you explicitly document the
|
||||||
|
limitation, identify which contributors remain inconclusive, and state
|
||||||
|
what additional evidence would be needed. Premature reporting without
|
||||||
|
this disclosure produces surface-level findings that miss the actual
|
||||||
|
root cause.
|
||||||
|
|
||||||
|
## Phase 4: Root Cause Identification
|
||||||
|
|
||||||
|
1. Distinguish between the **root cause** (fundamental defect) and the
|
||||||
|
**proximate cause** (immediate trigger).
|
||||||
|
- Example: The proximate cause is "null pointer dereference on line 42."
|
||||||
|
The root cause is "the initialization function silently fails when
|
||||||
|
the config file is missing, leaving the pointer uninitialized."
|
||||||
|
2. Trace the **causal chain** from root cause to symptom.
|
||||||
|
3. Ask: "If we fix only the proximate cause, will the root cause
|
||||||
|
produce other failures?" If yes, the fix is incomplete.
|
||||||
|
|
||||||
|
## Phase 4a: Cross-Component Causal Chains
|
||||||
|
|
||||||
|
When the investigation involves multiple components, processes, or
|
||||||
|
subsystems, trace causal chains across boundaries:
|
||||||
|
|
||||||
|
1. **Identify trigger-response pairs**: Does activity in component A
|
||||||
|
cause work in component B? For example, a file write by one process
|
||||||
|
may trigger scanning by an antivirus service, which triggers hashing
|
||||||
|
by an EDR agent, which triggers network inspection by another service.
|
||||||
|
2. **Map the amplification cascade**: A single action may fan out into
|
||||||
|
disproportionate downstream work. Document the full chain:
|
||||||
|
`Trigger → Reactor₁ → Reactor₂ → ... → Observed symptom`.
|
||||||
|
3. **Quantify amplification**: For each link in the chain, estimate the
|
||||||
|
cost ratio (e.g., "1 file write triggers 3 scan operations, each
|
||||||
|
consuming 50ms of CPU"). The amplification factor often explains why
|
||||||
|
a seemingly minor activity produces outsized impact.
|
||||||
|
4. **Identify the leverage point**: The most effective fix targets the
|
||||||
|
link in the chain with the highest amplification factor, not
|
||||||
|
necessarily the initial trigger or the final symptom.
|
||||||
|
|
||||||
|
Skip this phase when the investigation is confined to a single component
|
||||||
|
with no cross-boundary interactions.
|
||||||
|
|
||||||
|
## Phase 5: Remediation
|
||||||
|
|
||||||
|
1. Propose a fix for the **root cause**, not just the symptom.
|
||||||
|
2. Identify **secondary fixes** needed to prevent recurrence:
|
||||||
|
- Assertions or precondition checks
|
||||||
|
- Improved error handling
|
||||||
|
- Logging or monitoring
|
||||||
|
- Tests that would have caught this
|
||||||
|
3. Assess the **risk of the fix**: could it introduce new issues?
|
||||||
|
|
||||||
|
<!-- BEGIN ocserv extensions -->
|
||||||
|
|
||||||
|
## ocserv-Specific Extensions
|
||||||
|
|
||||||
|
The sections below extend the generic protocol with ocserv's three-process
|
||||||
|
architecture, documentation sources, and diagnostic conventions. Apply these
|
||||||
|
alongside the base phases above — they do not replace them.
|
||||||
|
|
||||||
|
### Phase 1 — Symptom Characterization (ocserv)
|
||||||
|
|
||||||
|
In addition to the generic checklist:
|
||||||
|
|
||||||
|
- **Consult documentation first.** Check `doc/ocserv.8.md` and
|
||||||
|
`doc/sample.config` for the documented behavior of the option or feature
|
||||||
|
under investigation. If the observed behavior contradicts the documentation,
|
||||||
|
that divergence is itself part of the bug characterization — note it
|
||||||
|
explicitly. The code must match the documentation, or both must be updated
|
||||||
|
together.
|
||||||
|
- **Identify the originating process.** ocserv has three processes with
|
||||||
|
distinct responsibilities:
|
||||||
|
- **main** (`main.c`, `main-*.c`) — TCP/UDP listeners, TUN devices, IP
|
||||||
|
allocation, process lifecycle.
|
||||||
|
- **sec-mod** (`sec-mod.c`, `sec-mod-*.c`) — authentication, private keys,
|
||||||
|
session state, PAM, accounting.
|
||||||
|
- **worker** (`worker.c`, `worker-*.c`) — per-client TLS/DTLS, VPN traffic.
|
||||||
|
|
||||||
|
Run `VERBOSE=1 ./tests/<test>` (or start the server with verbose logging)
|
||||||
|
and identify which process emitted the error. The originating process
|
||||||
|
directly constrains the hypothesis space.
|
||||||
|
- **Intermittent failures.** If the failure is intermittent, check whether
|
||||||
|
it correlates with:
|
||||||
|
- Concurrent client connections (IPC contention).
|
||||||
|
- Config reload timing (`SIGHUP` during active sessions).
|
||||||
|
- Specific authentication backends (PAM, RADIUS, OIDC).
|
||||||
|
- Client reconnection / cookie resumption flows.
|
||||||
|
|
||||||
|
### Phase 2 — Hypothesis Generation (ocserv)
|
||||||
|
|
||||||
|
In addition to the generic "non-obvious hypothesis" requirement, always
|
||||||
|
consider at least one hypothesis from each of these ocserv-specific
|
||||||
|
categories before investigating any hypothesis:
|
||||||
|
|
||||||
|
- **IPC race**: A message arrived at main or sec-mod while another request
|
||||||
|
was in-flight; the response was matched to the wrong caller.
|
||||||
|
- **Config reload timing**: `SIGHUP` triggered a config reload that raced
|
||||||
|
with an active authentication or session establishment.
|
||||||
|
- **seccomp filter gap**: A syscall needed by a new code path is missing
|
||||||
|
from the worker's seccomp whitelist, causing the worker to be killed by
|
||||||
|
the kernel.
|
||||||
|
- **Allocator mismatch across process boundary**: A talloc pointer was
|
||||||
|
embedded in an IPC message and dereferenced in a different process where
|
||||||
|
it is meaningless; or `gnutls_malloc` / `gnutls_free` was used for memory
|
||||||
|
that outlives the GnuTLS context.
|
||||||
|
- **SID / cookie lifecycle**: The session ID was invalidated (e.g., by a
|
||||||
|
sec-mod restart or cookie expiry) before the worker consumed it, leaving
|
||||||
|
a dangling reference.
|
||||||
|
|
||||||
|
### Phase 3a — Iterative Deepening (ocserv)
|
||||||
|
|
||||||
|
Apply the generic layers with these ocserv-specific granularities:
|
||||||
|
|
||||||
|
| Layer | ocserv granularity |
|
||||||
|
|-------|--------------------|
|
||||||
|
| Broad survey | Which process (main / sec-mod / worker)? |
|
||||||
|
| Attribution | Which module / source file within that process? |
|
||||||
|
| Deep analysis | Which function, IPC message field, or auth backend call? |
|
||||||
|
| Cross-component | Trace the IPC call chain (e.g., `SEC_AUTH_INIT` → `SEC_AUTH_REP` → `AUTH_COOKIE_REQ` → ...) |
|
||||||
|
|
||||||
|
For IPC-level bugs, read `doc/design.md` for the authoritative message
|
||||||
|
sequence before tracing code paths.
|
||||||
|
|
||||||
|
### Phase 4 — Root vs. Proximate Cause (ocserv)
|
||||||
|
|
||||||
|
ocserv-specific examples of the root/proximate distinction:
|
||||||
|
|
||||||
|
- **Proximate**: "null pointer dereference in worker at `worker-http.c:312`."
|
||||||
|
**Root**: "sec-mod returned a zero-length SID when a config reload raced
|
||||||
|
with `SEC_AUTH_INIT`, leaving the worker's session pointer uninitialized."
|
||||||
|
|
||||||
|
- **Proximate**: "worker killed with SIGKILL during a new feature's first
|
||||||
|
system call."
|
||||||
|
**Root**: "the new syscall is not in the seccomp whitelist; the kernel
|
||||||
|
kills workers that attempt it."
|
||||||
|
|
||||||
|
- **Proximate**: "`talloc_free` crashes on a session struct."
|
||||||
|
**Root**: "ownership of the struct was transferred across an IPC boundary
|
||||||
|
without re-parenting under the receiving process's talloc context."
|
||||||
|
|
||||||
|
After identifying the root cause, ask: if we fix only the proximate cause,
|
||||||
|
will the root cause produce failures elsewhere in the IPC call chain?
|
||||||
|
|
||||||
|
### Phase 4a — Cross-Component Causal Chains (ocserv)
|
||||||
|
|
||||||
|
ocserv's IPC topology makes cross-component tracing essential. The canonical
|
||||||
|
chain is:
|
||||||
|
|
||||||
|
```
|
||||||
|
worker → sec-mod: SEC_AUTH_INIT
|
||||||
|
sec-mod → worker: SEC_AUTH_REP (SID assigned)
|
||||||
|
worker → sec-mod: SEC_AUTH_CONT (SID + credential)
|
||||||
|
sec-mod → worker: SEC_AUTH_REP (OK / FAILED)
|
||||||
|
worker → main: AUTH_COOKIE_REQ (SID)
|
||||||
|
main → sec-mod: SECM_SESSION_OPEN (SID)
|
||||||
|
sec-mod → main: SECM_SESSION_REPLY (user config)
|
||||||
|
main → worker: AUTH_COOKIE_REP (TUN device, config)
|
||||||
|
```
|
||||||
|
|
||||||
|
When tracing a causal chain, map each IPC message hop and identify at
|
||||||
|
which hop the invariant breaks. The leverage point is usually the first
|
||||||
|
hop where a precondition is violated, not the hop where the crash or
|
||||||
|
error is ultimately observed.
|
||||||
|
|
||||||
|
### Phase 5 — Remediation (ocserv)
|
||||||
|
|
||||||
|
Secondary fixes to consider for every ocserv bug:
|
||||||
|
|
||||||
|
- **Assertions**: Add `assert()` or a log-and-return guard at the point
|
||||||
|
where the violated precondition should have been caught.
|
||||||
|
- **Error messages**: ocserv's `mslog()` / `oclog()` / `seclog()` messages
|
||||||
|
should identify the process, the IPC message type, and the session ID (SID)
|
||||||
|
so that field-reported bugs can be diagnosed from logs alone.
|
||||||
|
- **Tests**: Write a regression test in `tests/` that reproduces the bug.
|
||||||
|
Register it in `tests/meson.build`. For IPC-race bugs, a test that
|
||||||
|
exercises concurrent connections under load is preferred.
|
||||||
|
- **Documentation**: If the bug arose because the documented behavior was
|
||||||
|
ambiguous or incorrect, update `doc/ocserv.8.md` and `doc/sample.config`
|
||||||
|
alongside the code fix.
|
||||||
|
|
||||||
|
Assess fix risk with ocserv's privilege model in mind: a change in the
|
||||||
|
worker that looks local may require a corresponding IPC or seccomp change;
|
||||||
|
a change in sec-mod's session handling may affect all auth backends.
|
||||||
|
|
||||||
|
<!-- END ocserv extensions -->
|
||||||
Reference in New Issue
Block a user