Umbra scores any AI-built repo 0–100. Security. Slop. Does it actually run. Is the agent lying about its tests. One command, fully local, evidence for every finding.
40–60% of AI-generated code ships with exploitable vulnerabilities, and agents report outcomes they never measured. The tooling for writing code with AI is a year ahead of the tooling for trusting it. Umbra measures four things no other tool checks together.
UMBRA TRUST SCORE: 49/100 ● SAFE ✅ 100/100 — 0 findings CLEAN ✅ 100/100 — 0 findings RUNS — not measured — no detectable run path HONEST ⚠️ 50/100 — 2 claims failed, 2 verified Score capped below passing: a documented claim was verified false. Claim receipts: CLAIM FAILED: "14 tests pass" — README.md:7 — actually 3 tests pass, 0 fail CLAIM FAILED: "build passes" — README.md:9 — actually build exits 1 CLAIM VERIFIED: "All tests pass" — CLAUDE.md:3 — 3 tests pass
Scanners ask "is this pattern dangerous?" Umbra asks the question vibe coding actually raises: the AI wrote this — can I trust it? Pattern matching for what it wrote. A sandbox for what it does. Receipts for what it said.
Deterministic rules. A locked-down sandbox. One score, versioned forever.
Real numbers from public, actively-maintained AI-built repositories — with per-class deep dives, the full per-repo table, and an honest methodology including the false positives we found in our own rules and fixed.