Security Lab Integration Guide

Technology: security · Category: tooling · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/security

Insight:

A security lab is an isolated, offline set of deliberately vulnerable targets you own — Juice Shop, DVWA, WebGoat, the MASTG crackmes — where you break something on purpose so you understand the defence that stops it. The trade-off is discipline: every offensive concept only earns its place when it ships back a countermeasure and a detection signal, and the lab must never have a route to a system you are not authorised to touch. For codeAmani this is where the CLAUDE.md security rules stop being a checklist — you feel why parameterised queries, webhook signature verification, RLS, and boundary sanitisation are non-negotiable because you watched each one fail.

███████╗███████╗ ██████╗██╗   ██╗██████╗ ██╗████████╗██╗   ██╗
██╔════╝██╔════╝██╔════╝██║   ██║██╔══██╗██║╚══██╔══╝╚██╗ ██╔╝
███████╗█████╗  ██║     ██║   ██║██████╔╝██║   ██║    ╚████╔╝
╚════██║██╔══╝  ██║     ██║   ██║██╔══██╗██║   ██║     ╚██╔╝
███████║███████╗╚██████╗╚██████╔╝██║  ██║██║   ██║      ██║
╚══════╝╚══════╝ ╚═════╝ ╚═════╝ ╚═╝  ╚═╝╚═╝   ╚═╝      ╚═╝

Security Lab Integration Guide

Focus: a defensive, offline lab of deliberately vulnerable targets you own, used to learn each attack class alongside the countermeasure that kills it and the signal that detects it — never as a playbook against real systems.

Read this before you install anything. Testing a computer system you do not own and are not authorised to test is a crime in essentially every jurisdiction codeAmani operates in — in the US under the Computer Fraud and Abuse Act (18 U.S.C. § 1030) and state equivalents, in Kenya under the Computer Misuse and Cybercrimes Act, 2018. "I was only learning" is not a defence, and neither is "the system was already broken." Intent does not create authorisation; a signed document does.

The rules this guide operates under, without exception:

  1. Own the target or hold written authorisation. Every host, app, container, and phone image in this lab is either something you installed on your own hardware or a purpose-built training platform whose terms of service explicitly invite testing (PortSwigger Web Security Academy, TryHackMe, Hack The Box, VulnHub images running locally). Nothing else. A public bug bounty programme's policy page is a form of authorisation — read it, and stay inside it.
  2. Written scope before any activity. A real engagement names the exact hosts, domains, IP ranges, and app versions in scope, and explicitly lists what is out of scope. Anything not named is out of scope. In your own lab, write the scope down anyway — it builds the habit and it stops "I'll just check whether the router does that too."
  3. Rules of engagement. Agree the test window, the rate limits, a named emergency contact on both sides, what happens if you find live customer data (stop, do not exfiltrate, report immediately), and the fact that you will not test availability. Data you encounter is handled under the client's data-protection obligations, not yours.
  4. No spillover. The lab has no default route to the internet and no route to production. Bind every target to 127.0.0.1 or a host-only network. A misconfigured docker run -p 3000:3000 publishes a knowingly-vulnerable app to your whole LAN — and, on a laptop with a public IP or an open Wi-Fi network, to strangers.
  5. Defence is the deliverable. A finding is not finished when you reproduce it. It is finished when you can state the root cause, the code-level fix, and the log line or rule that would have caught it. That is the entire point of this folder.

Out of scope for this guide, permanently: anything aimed at real or third-party systems, malware / ransomware / C2 development, denial-of-service techniques, detection or EDR evasion, and credential attacks at scale. Where a topic is dual-use, only the lab-isolated defensive framing appears here.

Overview

The lab is three things: targets (apps built to be broken), a proxy to watch traffic, and a notebook where each finding becomes a rule. Nothing exotic — the whole thing runs in Docker on a laptop.

The value is not the exploit. It is the loop: you see input reach a place it should never have reached, you find the line of code that let it, and you write the version that does not. After the fourth time you watch ' change the meaning of a SQL statement, you stop writing string-concatenated queries — permanently, without needing to be told.

Pick targets by what you want to learn:

Platform Runs Best for Cost
PortSwigger Web Security Academy hosted, per-lab the canonical, current explanation of every web class — free, no signup wall on the content free
OWASP Juice Shop Docker / npm a full modern JS SPA + API; realistic, gamified, ~100 challenges free, self-hosted
DVWA Docker Compose classic PHP app with a security level dial (low → impossible) — the single best vulnerable-vs-safe code diff free, self-hosted
OWASP WebGoat Docker lesson-by-lesson teaching with explanation built in; ships WebWolf as the "attacker's own server" free, self-hosted
TryHackMe hosted VMs guided rooms, gentle ramp, structured learning paths freemium
Hack The Box hosted VMs unguided machines; closest to a real engagement's ambiguity freemium
VulnHub download → local VM offline boot-to-root images you run on your own hypervisor free
OWASP MASTG crackmes Android / iOS mobile reverse-engineering and MASVS-RESILIENCE practice free
flowchart TB
  subgraph HOST["Your workstation"]
    PX["Intercepting proxy<br/>Burp Community · OWASP ZAP"]
  end
  subgraph LAB["Isolated lab — host-only, no default route"]
    J["OWASP Juice Shop<br/>127.0.0.1:3000"]
    D["DVWA<br/>127.0.0.1:4280"]
    W["WebGoat + WebWolf<br/>127.0.0.1:8080 / 9090"]
    M["Android emulator<br/>MASTG crackmes"]
  end
  subgraph OUT["Never a target"]
    P["Production · client systems"]
    T["Any third-party host"]
  end
  PX --> J
  PX --> D
  PX --> W
  PX --> M
  PX -.->|"no written authorisation"| P
  PX -.->|"illegal"| T
  J --> R["Root cause →<br/>countermeasure + detection"]
  D --> R
  W --> R
  M --> R
  R --> C["codeAmani CLAUDE.md rule<br/>parameterise · verify webhooks<br/>RLS · sanitise at boundaries"]

Official Documentation

Resource URL
PortSwigger Web Security Academy https://portswigger.net/web-security
Academy learning paths https://portswigger.net/web-security/learning-paths
OWASP Top 10 (current: 2025) https://owasp.org/Top10/
OWASP Cheat Sheet Series https://cheatsheetseries.owasp.org/
SQL Injection Prevention Cheat Sheet https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection_Prevention_Cheat_Sheet.html
OWASP Web Security Testing Guide (WSTG) https://owasp.org/www-project-web-security-testing-guide/
OWASP Juice Shop https://owasp.org/www-project-juice-shop/
DVWA (source) https://github.com/digininja/DVWA
OWASP WebGoat https://owasp.org/www-project-webgoat/
OWASP MAS (MASVS + MASTG) https://mas.owasp.org/
MASTG crackmes https://mas.owasp.org/crackmes/
MITRE ATT&CK (defensive mapping) https://attack.mitre.org/
NIST SP 800-115 (technical testing methodology) https://csrc.nist.gov/pubs/sp/800/115/final

Setup — build the lab

1. Start the targets

All three publish to 127.0.0.1 only. That prefix is not decoration — drop it and you are hosting a known-vulnerable app on your network.

# OWASP Juice Shop — modern SPA + REST API
docker run --rm -p 127.0.0.1:3000:3000 bkimminich/juice-shop

# OWASP WebGoat + WebWolf — guided lessons
docker run -it -p 127.0.0.1:8080:8080 -p 127.0.0.1:9090:9090 webgoat/webgoat

# DVWA — classic PHP, with the security-level dial
git clone https://github.com/digininja/DVWA.git
cd DVWA
docker compose up -d
# → http://localhost:4280   (default creds admin / password, then "Create / Reset Database")

Juice Shop lands on http://localhost:3000, WebGoat on http://localhost:8080/WebGoat/, WebWolf on http://localhost:9090/WebWolf/, DVWA on http://localhost:4280.

2. Add an intercepting proxy

An intercepting proxy is a debugging tool first and a testing tool second — the same thing you already reach for when a webhook body does not look like you expected.

# Burp Suite Community  — https://portswigger.net/burp/communitydownload
# OWASP ZAP (free, open source, scriptable, runs in CI)
docker run --rm -u zap -p 127.0.0.1:8080:8080 \
  ghcr.io/zaproxy/zaproxy:stable zap-webswing.sh

Point the browser at the proxy, install its CA certificate in a throwaway browser profile only, and never in your daily driver.

3. Isolate the network

Isolation is the load-bearing control, and it has its own guide — see sandbox/CLAUDE_CODE_INTEGRATION.md for VM/container isolation and networking/CLAUDE_CODE_INTEGRATION.md for the addressing and routing model. Minimum bar:

# A Docker network with no route off the host
docker network create --internal lab-net

# Verify a container on it genuinely cannot reach the internet
docker run --rm --network lab-net alpine sh -c "wget -qO- -T3 https://example.com || echo 'no egress — correct'"

For VM-based targets (VulnHub images especially — they are untrusted third-party disk images), use a host-only adapter, snapshot before first boot, and revert after. Never bridge a VulnHub image to your LAN.

4. Environment variables

The lab itself needs no secrets, which is the point. If you script anything against it, keep the values local and never point them at anything real:

# .env.local — lab only, never committed, never a production host
LAB_TARGET_URL=http://127.0.0.1:3000
LAB_PROXY_URL=http://127.0.0.1:8080

Never place a production URL, API key, or database connection string in a lab script. If a tool asks for a target and you have to think about whether the answer is allowed, the answer is no.


The attack classes — and what kills them

The current standard is the OWASP Top 10:2025. Two things changed that matter for how you plan a lab: SSRF was folded into A01 Broken Access Control, and Software Supply Chain Failures (A03) is now its own category — which is exactly why codeAmani's SLSA provenance policy exists (supply-chain/CLAUDE_CODE_INTEGRATION.md).

# (2025) Class What goes wrong Countermeasure How you detect it
A01 Broken Access Control (now includes SSRF) The server trusts a client-supplied identifier — ?orderId=1002, a JWT claim, a hidden field, a URL it is asked to fetch Deny by default; authorise server-side on every request against the session subject, never a request parameter. For SSRF: allowlist destination hosts, resolve-then-validate the IP, block link-local 169.254.169.254 and RFC1918 Log (subject, object, decision) on every access check; alert on a spike of denies from one session, or on outbound requests to internal ranges
A02 Security Misconfiguration Debug mode in prod, default creds, permissive CORS, verbose stack traces, a storage bucket left public Hardened build baseline in CI; explicit CORS origins; generic error responses; config diffed against a known-good template Config drift scanning; alert on Access-Control-Allow-Origin: * or a 500 that leaks a file path
A03 Software Supply Chain Failures A dependency, build system, or distribution channel is compromised — not just "old library" Pin and lock; npm audit signatures; SLSA Build L3 provenance on anything shipped; pin GitHub Actions to SHAs Dependency review in CI; alert on a lockfile change in a PR that touches no source
A04 Cryptographic Failures Secrets at rest in plaintext, home-rolled crypto, weak hashing, TLS not enforced Platform primitives only (argon2/bcrypt, AEAD ciphers); HSTS; secrets in a manager (Hazina), never in the repo Secret scanning (gitleaks) as a pre-push gate; TLS posture monitoring
A05 Injection (SQLi, command, XSS, template, LDAP) User input crosses from data into a grammar — SQL, a shell command, HTML, a template Parameterise. Prepared statements, execFileSync(cmd, [args]), contextual output encoding, DOMPurify for rendered HTML. Validate at the boundary as defence-in-depth, never as the primary control WAF / CRS rules; DB error-rate spikes; alert on queries whose shape changes (normalised statement fingerprint)
A06 Insecure Design The feature is unsafe as specified — no rate limit on OTP, no re-auth before an email change Threat-model at plan time; write abuse cases next to user stories Business-logic anomaly detection: N password resets/hour, refunds exceeding charges
A07 Authentication Failures Credential stuffing survivable, no MFA, weak session lifecycle, tokens that never expire MFA; rate limits + lockout on credential routes; rotate session ID on privilege change; short-lived tokens Alert on auth failure rate per account and per IP; impossible-travel; new-device sign-in
A08 Software or Data Integrity Failures Unsigned updates, insecure deserialisation, CI that trusts an unpinned action Signed artefacts + verification; never deserialise untrusted input into live objects; pin the pipeline Verify signatures at install; alert on unexpected artefact digests
A09 Security Logging and Alerting Failures The attack happened and nothing recorded it — or it recorded and nobody was paged Structured security events (authn, authz denials, admin actions, payment state changes) with a real alerting path This is the detection layer. Test it: run a lab attack and confirm something fires
A10 Mishandling of Exceptional Conditions Errors leak internals, or a failure path silently falls open (catch {} → grant access) Fail closed; generic client errors + detailed server-side logs; never swallow an exception on a security path Alert on error-rate spikes on auth/payment paths and on any authorisation code path reached via a catch block

Practise each row against a target: A01 on Juice Shop's basket and order endpoints, A05 on DVWA (flip the security level to read the fix diff), A07 and A09 on WebGoat's lesson sequence, and the whole set on the Academy's per-topic labs.


SQL injection in depth

SQL injection is the reference case because the mechanism generalises to every other injection class — and because the fix is one line.

The mechanism

A SQL statement has two things in it: grammar (SELECT, WHERE, ', --) and data (the phone number a user typed). String concatenation destroys that boundary — the user's text is parsed as grammar. A single ' closes the string literal early, and everything after it is executed as SQL.

// ❌ VULNERABLE — the input becomes part of the query's grammar
const phone = req.query.phone as string;
const rows = await db.query(
  `SELECT id, name, phone FROM riders WHERE phone = '${phone}'`
);

With phone = 254712000000 the database parses ... WHERE phone = '254712000000'. With phone = ' OR '1'='1 it parses ... WHERE phone = '' OR '1'='1' — a tautology, and the endpoint returns every rider. The attacker did not "guess a password"; they rewrote the query.

The fix: parameterise

A parameterised (prepared) statement sends the query text and the values to the database as separate things. The parser sees the statement first and finalises the grammar; values are then bound to placeholders. A ' inside a bound value is a ' character in a string — it can no longer become punctuation.

// ✅ SAFE — node-postgres / Neon: $1 placeholder, values in an array
const rows = await db.query(
  "SELECT id, name, phone FROM riders WHERE phone = $1",
  [phone]
);
// ✅ SAFE — postgres.js tagged template: interpolations are parameterised, not concatenated
const rows = await sql`SELECT id, name, phone FROM riders WHERE phone = ${phone}`;
// ✅ SAFE — Supabase / PostgREST: the filter builder parameterises for you
const { data, error } = await supabase
  .from("riders")
  .select("id, name, phone")
  .eq("phone", phone);

Escaping is not the fix. Hand-written quote-doubling, blocklists of the word UNION, and "strip the apostrophes" all fail against numeric contexts, second-order injection (safe on write, concatenated on read), encoding tricks, and the next database driver you swap in. Parameterise.

The parameterisation footguns

Placeholders bind values, never identifiers or keywords. Anywhere the query shape is dynamic, you cannot parameterise — so you must map through an allowlist:

// ❌ VULNERABLE — sort column concatenated straight in
const rows = await db.query(`SELECT * FROM deliveries ORDER BY ${req.query.sort}`);

// ✅ SAFE — allowlist maps an opaque token to a literal you wrote
const SORTS = { newest: "created_at DESC", fare: "fare_kes DESC" } as const;
const orderBy = SORTS[req.query.sort as keyof typeof SORTS] ?? SORTS.newest;
const rows = await db.query(`SELECT * FROM deliveries ORDER BY ${orderBy}`);

Every ORM keeps a raw escape hatch, and every one of them is the place SQLi comes back: Prisma's $queryRawUnsafe, Drizzle's sql.raw(), Knex's knex.raw() with template interpolation, TypeORM's query(). Grep for them in review. The safe raw form always takes values separately:

// ❌  prisma.$queryRawUnsafe(`SELECT * FROM riders WHERE phone = '${phone}'`)
// ✅  prisma.$queryRaw`SELECT * FROM riders WHERE phone = ${phone}`   // tagged template = parameterised

The layers behind it

Parameterisation is the control. These reduce the blast radius when something else slips through:

Least-privilege database user. The app's role should not be able to read tables it never touches, and should not be able to change schema. A read-only reporting path gets its own role.

-- The application role: exactly the verbs it needs, on exactly the tables it needs
REVOKE ALL ON SCHEMA public FROM PUBLIC;
CREATE ROLE app_rw LOGIN PASSWORD :'app_password';
GRANT USAGE ON SCHEMA public TO app_rw;
GRANT SELECT, INSERT, UPDATE ON deliveries, riders TO app_rw;
GRANT SELECT ON fare_rates TO app_rw;
-- deliberately absent: DROP, CREATE, TRUNCATE, and any grant on audit_log or webhook_jobs

Row Level Security, on Supabase and on plain Postgres, turns "the query returned rows it should not have" into "the database refused." It is the last line that holds when application-layer authorisation has a bug — which is exactly the A01 failure mode.

ALTER TABLE deliveries ENABLE ROW LEVEL SECURITY;
CREATE POLICY rider_reads_own ON deliveries
  FOR SELECT USING (rider_id = auth.uid());

Input validation at the boundary. Validate shape, type, and range with a schema at the edge of the system — every form, API route, and webhook handler. It catches a large class of nonsense early and it makes the intended domain explicit. It is defence-in-depth, never the primary control.

import { z } from "zod";

const RiderLookup = z.object({
  // 254XXXXXXXXX — the M-Pesa phone format from CLAUDE.md
  phone: z.string().regex(/^254\d{9}$/),
});

const parsed = RiderLookup.safeParse(await req.json());
if (!parsed.success) return Response.json({ error: "invalid request" }, { status: 400 });

A WAF (Cloudflare's managed rules, or the OWASP Core Rule Set in front of your own origin) buys time against automated scanning and mass exploitation of a freshly-published CVE. It is a speed bump on a determined, targeted attempt. Never let its presence justify a concatenated query.

Detecting it


Mobile — OWASP MASVS and MASTG

The mobile equivalent of the Top 10 is the OWASP Mobile Application Security project: MASVS (the requirements standard) and MASTG (the testing guide, plus the crackmes to practise on). The controls are organised into eight categories, each of which maps to a decision you make while writing an app:

MASVS category The defensive question
MASVS-STORAGE Is anything sensitive on disk, in a backup, in a log, or hardcoded in the package?
MASVS-CRYPTO Are keys in the platform keystore, and is the crypto the platform's, not yours?
MASVS-AUTH Is authorisation enforced server-side, with the device only presenting a credential?
MASVS-NETWORK Is TLS enforced, with no user-added CA trust, and validation never disabled "for debugging"?
MASVS-PLATFORM Are IPC surfaces (exported activities, intents, deep links, WebViews, pasteboard) locked down?
MASVS-CODE Are dependencies current, is untrusted input validated, is debug tooling stripped from release?
MASVS-RESILIENCE Does tampering/reverse engineering raise the cost — knowing it never makes the app safe?
MASVS-PRIVACY Is data minimised, disclosed, and does the user actually have control?

Current MASTG is v2, which introduced MASWE weakness IDs (MASWE-0001 …) alongside MASTG-TEST-*, MASTG-DEMO-*, and MASTG-TOOL-* identifiers — a demo per test, so you can see the finding reproduced and then fixed.

Practise on the official crackmes (https://mas.owasp.org/crackmes/) for Android and iOS, on an emulator or a dedicated wiped device. Never on a phone that holds real accounts.

The single most useful lesson the mobile lab teaches: the client is not a trust boundary. Anything the app can compute, an attacker with the binary can compute. Root/jailbreak detection, certificate pinning, and obfuscation (MASVS-RESILIENCE) raise cost — they do not create security. Every authorisation decision belongs on the server. This is the same rule as "middleware is a redirect optimisation, not an authorisation boundary" from better-auth/CLAUDE_CODE_INTEGRATION.md, wearing different clothes.


Desktop and OS hardening

The lab teaches the app layer; the host layer is where a compromise becomes permanent. A defensible baseline, per platform:

Platform Baseline that matters Official reference
Windows BitLocker on, Secure Boot + TPM, virtualisation-based security / Credential Guard, SmartScreen, Defender with tamper protection, standard (non-admin) daily account, Attack Surface Reduction rules https://learn.microsoft.com/en-us/windows/security/
macOS FileVault, System Integrity Protection left on, Gatekeeper + notarisation enforced, firewall on, standard account for daily use, Lockdown Mode where the threat model warrants https://support.apple.com/guide/security/welcome/web
Linux Full-disk encryption, automatic security updates, no password SSH (keys only) with root login disabled, host firewall default-deny inbound, SELinux/AppArmor enforcing, minimal installed surface https://ubuntu.com/security
iOS Current iOS, strong passcode + biometric, automatic updates, minimal profile trust (never install an unknown MDM/CA profile), Lockdown Mode for high-risk users https://support.apple.com/guide/security/welcome/web
Android Current Android + Play Protect, verified boot, sideloading off, per-app permission review, work profile to separate contexts, no root on a device holding real accounts https://source.android.com/docs/security

Cross-cutting, and worth more than any individual toggle: patch fast, run as a standard user, use a password manager plus phishing-resistant MFA (passkeys / hardware keys), full-disk encryption everywhere, and have a restore-tested backup. Ransomware defence is backup and recovery, not a product.

For developer machines specifically: keep the lab off the machine that holds production credentials. Lab VMs get snapshotted and reverted; VulnHub images are untrusted binaries from strangers and are treated as such. See sandbox/CLAUDE_CODE_INTEGRATION.md.


The loop: finding → fix → detection

Every lab session produces three artefacts, and stops being useful if it produces fewer:

  1. Root cause — the specific line or design decision, not the symptom. "The order endpoint reads orderId from the query string and never checks it against the session subject."
  2. Countermeasure — the diff. Parameterised query, server-side authorisation check, output encoding, signature verification. Written as code, not as advice.
  3. Detection — the log event and the alert that would have fired. If you cannot name it, you have found A09:2025 in your own stack.

Map the finding to MITRE ATT&CK (https://attack.mitre.org/) when it helps you talk to a security team, and score severity with CVSS (https://www.first.org/cvss/) when you need a shared vocabulary for prioritisation. Sigma (https://github.com/SigmaHQ/sigma) is a portable way to write the detection rule once.


codeAmani notes

This lab is where the CLAUDE.md security rules come from

Every rule in the workspace CLAUDE.md is the scar tissue of one of the classes above. The lab is how an engineer gets the instinct instead of the checklist:

Secrets and the lab

Kenya-targeted projects: the M-Pesa callback is hostile input

For boda-dispatch, duka-order-bot, and the rest of the Kenya-targeted builds, the highest-value application of this lab is the Daraja callback path, because it combines three classes at once:

Two more that are genuinely regional rather than forced: WhatsApp inbound messages are the same category of untrusted input as a form field — validate before they reach a query or a template. And on low-bandwidth Android, a tight Content-Security-Policy plus a small JS bundle is both a performance win and a real XSS mitigation, since CSP is what stops an injected script from executing at all.

Pair this guide with

The one-line summary

You break a target you own so that you never ship the bug — and the session is only finished when the finding has a diff and an alert attached to it.

Official docs: