secskills
secskills / core / auditing-code-for-vulnerabilities

auditing-code-for-vulnerabilities

core verified 2026-07-28

Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis. Use when reviewing a codebase or diff for security bugs, performing a security audit, hunting for vulnerabilities in a target's source, or validating whether a suspected finding is real.

$ /plugin install secskills-core

Finding real bugs in source code is a different job from running a scanner. A scanner matches patterns; an auditor builds a model of what the code is supposed to guarantee and then hunts for the paths where that guarantee breaks. This skill is the methodology for the second job.

When to Use

When NOT to Use

builder/critic separation and iteration — use orchestrating-vulnerability-research, which dispatches this skill as the per-slice hunter

The Loop

Auditing is four passes, not one. Do not skip to pass 3.

1. Context      → what does this system protect, and from whom?
2. Attack surface → where does untrusted input enter, and what does it reach?
3. Hunt          → trace specific bug classes along those paths
4. Verify        → prove exploitability before you write a word

Pass 1: Build context before reading code

Do not open files at random. Spend the first block of effort answering:

QuestionWhere to look
What are the security-relevant assets?README, docs, data models, DB schema
Who are the actors and trust tiers?Auth middleware, role enums, tenant models
What are the stated invariants?Tests, assertions, comments containing "must", "never", "invariant"
What has already been fixed here?git log --grep for security keywords, CVE files, SECURITY.md
What is out of scope?Engagement brief, vendor code, generated files
# Prior security work is the cheapest source of bug leads
git log --oneline --grep='security\|CVE\|vuln\|injection\|auth bypass\|overflow' -i | head -40

# Stated invariants often mark where the author was nervous
rg -n --stats 'MUST NOT|must never|SECURITY|XXX|HACK|TODO.*(auth|secur|valid)' -i

# Where does privilege actually get checked?
rg -n 'is_admin|require_role|authorize|has_permission|@login_required|checkAccess'

Write a short target model before hunting: assets, actors, trust boundaries, and the three invariants whose violation would matter most. Everything after this is a search for counterexamples to those invariants.

Pass 2: Map the attack surface

Enumerate entry points, then rank them. An entry point matters in proportion to how far it reaches before it is validated.

# HTTP/RPC routes
rg -n '@(app|router)\.(get|post|put|delete|patch)|app\.(get|post)\(|@RequestMapping|http\.HandleFunc'

# Deserialization, template rendering, and dynamic execution sinks
rg -n 'pickle\.loads|yaml\.load\(|Marshal|unserialize|ObjectInputStream|eval\(|new Function|exec\(|Runtime\.getRuntime'

# Command, SQL, and path sinks
rg -n 'os\.system|subprocess.*shell\s*=\s*True|child_process\.exec\(|execSync|Statement\.execute|\.raw\(|fmt\.Sprintf.*SELECT'

# Where authentication is decided rather than enforced
rg -n 'verify=False|InsecureSkipVerify|jwt\.decode\(.*verify.*False|algorithms=\[.*none'

Rank entry points by: reachable without authentication > reachable by a low privilege tier > reachable only by an admin. Then follow the highest-ranked ones inward. Depth beats breadth — one fully traced path is worth twenty grep hits.

Grep hits are candidates, not findings. The commands above are a cheap wide net; each match is an unresolved lead until you have traced it. Persist the candidate set — a worklist of (file:line, bug class, entry point) — and drive every entry to an explicit verdict: traced-safe, confirmed, or needs-PoC. Widen the net cheaply, then spend expensive reasoning per candidate — never the reverse. The failure mode is not a missing grep pattern; it is enumerating fifty candidates, eyeballing five, and calling the tree clean. On a large codebase, fold the project's own conventions into the net — its ORM's raw-query escape hatch, its auth decorator's name, its templating call — because the highest-yield sinks are the ones generic patterns miss.

Pass 3: Hunt bug classes along the traced paths

For each promising path, trace taint from source to sink and ask what the code assumes. The high-yield classes, in rough order of how often they survive to production:

Authorization, not authentication. Most real breaches are missing object level checks, not broken login. For every handler that takes an ID, ask: is the object scoped to the caller's tenant/user, or only looked up by ID? Check the query, not the decorator.

Trust-boundary confusion. Data validated at one layer and re-parsed at another. Look for values that cross a serialization boundary — a validated string re-parsed as a URL, a path, a template, or a query.

State and concurrency. Check-then-use gaps, non-atomic balance updates, idempotency keys that are not actually unique, retry paths that replay side effects. Search for reads followed by writes with no lock or transaction.

Injection into a secondary interpreter. SQL, shell, LDAP, XPath, template engines, log formats, and regex. The question is never "is there a filter" but "does the filter and the interpreter agree on the grammar."

Memory safety (C/C++/unsafe Rust/CGo). Length arithmetic before bounds checks, memcpy with an attacker-influenced size, off-by-one in loop bounds, signed/unsigned conversions, use-after-free on error paths.

Error and cleanup paths. The happy path is usually reviewed; the except, catch, defer, and goto fail branches are not. Audit them specifically.

Secrets and cryptographic misuse. Hardcoded keys, non-constant-time comparison of tokens, predictable IDs from Math.random/rand(), missing signature verification. Deep crypto review belongs in reviewing-cryptography.

Pass 4: Variant analysis

A bug is a template, not an incident. When you confirm one, immediately search for its siblings — the same mistake made by the same author, the same copied block, the same missing check on a neighbouring route.

# You found one unscoped lookup. Find every other one.
rg -n 'find_by_id|findOne\(\{ *_id|get_object_or_404' -A3 | rg -v 'tenant|owner|user_id'

Variant analysis is where audits produce disproportionate value. Budget time for it explicitly — roughly one unit of variant search per confirmed finding.

Verification Before Reporting

A finding you cannot demonstrate is a hypothesis. Before it goes in the report, answer all four:

  1. Reachability — name the concrete entry point and the caller privilege

required. "An attacker who can reach POST /api/export unauthenticated."

  1. Control — show which part of the dangerous value the attacker controls.
  2. Impact — state what breaks: which invariant from Pass 1, and what an

attacker gains.

  1. No mitigating control — check for a WAF rule, a framework default, a

middleware, a DB constraint, or a caller that already sanitizes.

If a proof of concept is in scope, write the smallest one that proves control of the sink — not a weaponized exploit.

Revalidate to prune false positives

The four checks above confirm a finding; this pass tries to kill it. Run it on every confirmed candidate before it reaches the report — a report's credibility is set by its worst false positive, not its best true finding.

may sit on a branch you have not pulled. Confirm the vulnerable code is what actually ships before you file it.

``bash git log -S'<dangerous token>' --oneline -- <file> # when this line changed, and toward what git log --oneline <checkout>..origin/main -- <file> # a fix on main you are not reading git blame -L <line>,<line> <file> # the commit that introduced it, for context ``

safe and go find the control that makes it so — the middleware, the DB constraint, the caller that already sanitizes. A finding that survives a genuine attempt to disprove it is one you can defend.

refactor often slips a validator or an encoder between source and sink that a first read glides past.

Drop what dies here, and say so in your coverage notes. A candidate you cannot revalidate is a note to yourself, not a finding.

Rationalizations to Reject

These are the thoughts that turn an audit into a formality. Each one is wrong.

assumptions about a caller are the single most common source of missed bugs.

the version, check the config, check whether this call uses the safe API.

queue consumer or admin import is usually reachable from outside.

not a reason to drop the finding.

the finding.

your own first read. Revalidate it against what actually ships and against a genuine attempt to disprove it before it goes in the report.

surface you mapped in Pass 2, not against file count.

Tool Assist, Not Tool Substitute

Static analysis is for coverage and for variant search after you know the pattern. Write a rule once you have a confirmed bug, and let it find the rest.

# Semgrep: broad pass, then a rule you write for your specific finding
semgrep --config=auto --severity=ERROR --json -o semgrep.json .
semgrep --config=./rules/my-variant-rule.yaml .

# CodeQL for dataflow questions grep cannot answer
codeql database create db --language=<lang> && codeql database analyze db --format=sarif-latest -o out.sarif

# Language-specific
bandit -r . -f json            # Python
gosec -fmt=json ./...          # Go
cargo audit && cargo geiger    # Rust deps + unsafe surface
npm audit --json               # JS deps

Triage every tool finding through the four verification questions above. A report of unverified scanner output is worse than no report — it burns the reader's trust and buries the real bugs.

Audit Deliverable

Track coverage as you go, and state it honestly:

## Coverage
| Component | Files | Depth      | Notes                          |
|-----------|-------|------------|--------------------------------|
| auth/     | 12    | Full trace | All routes traced to sinks     |
| billing/  | 30    | Partial    | Webhook handlers only          |
| vendor/   | -     | Excluded   | Out of scope per brief         |

## Findings
F1. [High] Tenant isolation bypass in GET /api/reports/:id — <impact> — <repro>

Say what you did not cover. An audit that claims full coverage it did not achieve is the most damaging artifact you can produce.

References