Usage
Usage
Day-to-day operation of the lab assumes Installation is complete: the aisec-lab NixOS machine converged, nix develop active, and Jupyter serving on Mac localhost:8888. Two ways to run it — non-interactive and interactive — plus the rule-iteration workflow after the scripted tests pass.
Non-interactive: nbdev_test, the acceptance suite
From a VM shell inside the repository:
nix develop
nbdev_test --n_workers 0 # serial — one shared clusternbdev_test exits 0 only when, in notebook order:
- cluster reachable — one Ready node,
aisec-lab; - all four
.novrule files parse, and each evaluator tier (keywords → semantics → LLM judge) catches its target while a benign prompt passes; - the gate’s policy configuration catches the same tiers in a live NOVA engine;
- the victim’s traps assert armed: canary key, poisoned support note, CRM sink;
- OpenTofu validates the generated configuration;
- workloads pass their rollouts and the edge answers its first probe;
- benign prompts return 200 with real model completions (transparency);
- every attack category returns 403 at the edge and the CRM log gains no entries (the gate holds);
- the gate audit log attributes every block (
engine+ tier/label) and every pass.
Two cells are # slow-flagged and therefore skipped by nbdev_test: the tofu apply (04_deploy) and the direct-path leak proof (05_attacks, a real LLM turn + tool loop). Run those cells interactively on first setup — click through 04_deploy and 05_attacks in order — then nbdev_test re-verifies the rest against the standing lab. The free-play scratchpad is #| notest and is not part of the suite.
Interactive: click through the notebooks
Open notebooks/index.ipynb first — it carries a linked contents table for the lab — then follow its table in order. Each markdown cell states the next cell’s function, expected output, and failure modes; each notebook ends with the run-order entry for the next one.
| Run order | Notebook | Covers |
|---|---|---|
| 1 | 00_core | shared constants + the fail-closed shell helper + cluster connectivity test |
| 2 | 01_nova_rules | the four .nov rule files, one per evaluator tier, live-scan tested |
| 3 | 02_gate_app | the FastAPI gate (NOVA + Laya) and its container image |
| 4 | 03_victim | Atlas — the vulnerable support app and its three weaknesses |
| 5 | 04_deploy | OpenTofu: validate → apply → rollouts → first edge probe |
| 6 | 05_attacks | the two-path test: leak (direct path), block (edge path) |
Rule iteration: operate the gate, don’t only test it
The scripted corpus covers the committed rules. The deployment also supports ad-hoc iteration with zero image rebuilds:
The rules are a mounted ConfigMap — edit the live policy and restart the gate:
kubectl edit configmap nova-rules -n ai-sec kubectl rollout restart deploy/nova-gate -n ai-sec kubectl rollout status deploy/nova-gate -n ai-sec --timeout=180sEvery 403 body names the engine and tier that fired (
{"engine":"nova","tier":…}or{"engine":"laya","label":…}) — the feedback signal for which rule matched, or what passed.Prompts that pass reach the real engine with its tool loop — multi-turn interaction through
try_promptin 05_attacks’#| notestscratchpad.The CRM sink (
crm_log()) reports whether anything leaked.
The iteration workflow: attack → read the verdict metadata → tighten the matching .nov rule or the Laya threshold → restart the gate → re-run until 403 → then commit the learned rule to nova-rules/ so the scripted tests cover it. Findings promoted from scratch cell to tested rule are how this lab grows.
Publish the documentation site
nbdev_docs # renders all seven notebooks (code, prose, tests) as a static siteTear-down / re-run
cd terraform && tofu destroy -auto-approve # removes everything it created
docker rmi ai-sec-lab/laya-gate:1.0.0 ai-sec-lab/atlas-victim:1.0.0
nix store gc # reclaim the Nix store
# full teardown of the lab machine itself (from the Mac):
# orb delete aisec-lab # the VM is disposable; the flake rebuilds itSubsequent runs re-apply cheaply: tofu apply is a no-op when state matches, and nbdev_test re-verifies the acceptance criteria in order.