Vybel Standard v1.1
A method for grading the health of a software project, designed for projects built with AI and vibe coding.
Principles
- Deterministic. The grade involves no AI: the same code analyzed with the same version of the standard always gets the same grade and the same report ID.
- Reproducible. Time-based measures (how old a document is, recent activity) are calculated relative to the date of the analyzed commit, not the date of the analysis.
- Static. None of the project’s code is executed. Uncertain heuristics are marked “to verify” and only count for half.
- Comparable. Quality rules are scaled to the size of the project (share of files, occurrences per 1,000 lines), so that small and large projects are compared fairly.
- Verifiable. A label is verifiable when it covers a git commit with no local changes: any new analysis of that commit, under the same version of the standard, yields the same grade.
How the grade is calculated
- Each area starts at 100. Each issue subtracts a penalty, applied multiplicatively:
score = 100 × Π(1 − penalty/100). - The overall score is the average of the areas, weighted by their weight.
- Requirements for each grade, regardless of the average. When a requirement isn’t met, the score is brought below the limit gradually (
limit − 0.3 × (100 − average)), so that every improvement stays visible:- A requires automated tests, no important or blocking issues, and at least 70 in every area (otherwise the score stays below 85);
- B requires no important issues in security, structure, or duplication (otherwise the score stays below 70);
- one critical issue caps the score at 49 (D at best), three or more critical issues at 39 (E).
| Area | Weight | Question |
|---|---|---|
| Security | 1.5 | Are your data and keys protected? |
| Production readiness | 1.2 | Can you ship and maintain it with confidence? |
| Structure | 1 | Is the code split up in a way that’s easy to follow? |
| Duplication | 1 | Is the same code written more than once? |
| Dead code & debt | 0.8 | Is the project weighed down by abandoned code? |
| Stack | 0.8 | Are the tools consistent and up to date? |
| Docs & AI context | 0.7 | Can a human or an AI agent understand the project? |
Scale
| Grade | Score | Meaning |
|---|---|---|
| A — Excellent | 85–100 | Healthy project: you’re in control of your code. |
| B — Good | 70–84 | Project in good shape, with a few knots to untangle. |
| C — Fair | 55–69 | The project is starting to get tangled: now’s the time to clean up. |
| D — Fragile | 40–54 | The project is turning into a tangled mess: every change risks breaking something else. |
| E — Critical | 0–39 | Project at risk: deal with the blocking issues before going any further. |
Rules
Security (weight 1.5)
| Rule | What is measured | Penalty |
|---|---|---|
security/secrets — Secrets in the code | Known key formats (OpenAI, Anthropic, Stripe, AWS, GitHub, Supabase service_role, private keys, database URLs with a password…) in committed files. | Critical. −60 points, −10 per additional secret (max −80). |
security/client-exposed-secrets — Secrets exposed to the browser | Public environment variables (NEXT_PUBLIC_, VITE_…) with secret-like names, dangerouslyAllowBrowser, direct calls to AI APIs, or a service_role key in client code. | Critical. −50 points. |
security/env-committed — Committed .env file with secrets | .env file tracked by git that contains secret values. | Critical. −45 points. |
security/env-not-ignored — .env not ignored by git | .env file present but missing from .gitignore. | Important. −25 points. |
security/supabase-rls — Supabase tables without RLS | Tables created in SQL migrations without enable row level security. | Critical. −50 points. |
security/permissive-policies — RLS policies open to everyone | Write policies (insert, update, delete, all) with using (true) or with check (true). | Important. −25 points. |
security/rls-unverifiable — RLS not verifiable | Supabase used without SQL migrations in the repository. | To verify. −10 points (counts for half). |
security/firebase-open-rules — Open Firebase rules | allow read, write: if true, time-limited test mode, .read/.write: true. | Critical. −50 points. |
security/unverified-webhooks — Webhooks without signature verification | Webhook routes (Stripe…) that don’t verify the signature. | Important. −20 points. |
security/unprotected-api-routes — API routes with no apparent authentication | API routes that modify data with no visible access control. | To verify. −4 points + 2 per route (max −20), counts for half. |
security/dangerous-code — Dynamically injected code | dangerouslySetInnerHTML without a sanitization library, eval, new Function. | Medium. −8 points. |
Production readiness (weight 1.2)
| Rule | What is measured | Penalty |
|---|---|---|
prod/no-tests — No automated tests | No test files in a project with more than 10 source files. | Important. −30 points. |
prod/few-tests — Very few tests | Fewer than one test file per 20 source files. | Medium. −12 points. |
prod/ts-not-strict — TypeScript in permissive mode | strict disabled, or strictNullChecks / noImplicitAny set to false. | Medium. −10 points. |
prod/no-ci — No continuous integration | No CI pipeline (GitHub Actions, GitLab CI…). | Minor. −6 points. |
prod/no-linter — No linter | No ESLint, Biome, oxlint, or Ruff configuration. | Minor. −4 points. |
prod/missing-scripts — Missing check scripts | No build, lint, test, or typecheck script in package.json. | Minor. −4 points. |
prod/env-undocumented — Undocumented environment variables | Variables used in the code but missing from .env.example. | Medium. −8 points without .env.example, −4 if incomplete. |
prod/console-logs — Leftover debug logs | console.log/debug/info in the source code, per 1,000 lines. | −1.5 points per log per 1,000 lines (max −10). |
prod/committed-artifacts — Committed generated files | node_modules, venv, dist, build, .next… tracked by git. | −20 points for dependencies, −8 for builds. |
prod/hardcoded-localhost — Hardcoded localhost URLs | http://localhost URLs in the application code. | Medium. −5 points + 2 per occurrence (max −15). |
prod/no-error-monitoring — No error tracking | No error tracking tool in production (Sentry…). | Minor. −3 points. |
Structure (weight 1)
| Rule | What is measured | Penalty |
|---|---|---|
structure/giant-files — Giant files | Share of source files over 400 lines of code (weighted: ×2 above 800, ×3 above 1,500). | −120 × weighted share, +5 if a file exceeds 1,500 lines (max −40). Important if a file exceeds 1,500 lines or if the weighted share exceeds 15%. |
structure/long-functions — Overly long functions | Functions over 80 lines and components over 200 lines, relative to the number of files. | −60 × weighted share (max −30). Important if a function exceeds 3 times the threshold or if the weighted share exceeds 20%. |
structure/circular-deps — Circular dependencies | Groups of files that import each other (runtime imports). | −4 points + 100 × share of affected files (max −25). |
structure/broken-imports — Broken imports | Relative imports, and declared Rust modules, pointing to files that don’t exist. | Important. −12 points + 3 per import (max −30). |
structure/crowded-folders — Catch-all folders | Folders directly containing more than 35 source files. | Minor. −3 points per folder (max −9). |
Duplication (weight 1)
| Rule | What is measured | Penalty |
|---|---|---|
duplication/clones — Copy-pasted code | Share of duplicated lines of code (identical blocks of at least 50 tokens and 5 lines). | −2.5 points per % of duplication (max −60), starting at 1.5%. Important above 12%. |
duplication/same-name — Functions defined multiple times | Same function or component defined in several files. | −3 points per near-identical group, −1.5 otherwise (max −25). |
duplication/identical-files — Identical files | Source files with identical content in several places. | −4 points per group (max −20). |
duplication/multiple-clients — Service clients created multiple times | Supabase, Prisma, Stripe, OpenAI… clients instantiated in several files. | −6 points per service, −10 for Prisma (max −20). |
Dead code & debt (weight 0.8)
| Rule | What is measured | Penalty |
|---|---|---|
legacy/orphans — Unused files | Source files that nothing imports and that aren’t entry points. | −100 × share of affected files (max −30). |
legacy/unused-ui — Unused UI components | shadcn/ui components copied into the project but never imported. | −0.25 points per component (max −6). |
legacy/backup-files — Old versions and copies | Files named .old, .bak, copy, v2, backup… | −3 points per file (max −15). |
legacy/commented-code — Commented-out code | Lines of code commented out, per 1,000 lines. | −0.5 points per line per 1,000 (max −15). |
legacy/todos — TODOs and FIXMEs | TODO, FIXME, HACK, XXX markers in comments, per 1,000 lines. | −1 point per marker per 1,000 lines (max −10). |
legacy/suppressions — Disabled checks | @ts-ignore, @ts-nocheck, eslint-disable per 1,000 lines. | −2 points per occurrence per 1,000 lines, +2 per @ts-nocheck (max −12). |
legacy/any — Type checking bypassed (any) | Uses of any per 1,000 lines of TypeScript. | −1 point per use per 1,000 lines (max −15). |
Stack (weight 0.8)
| Rule | What is measured | Penalty |
|---|---|---|
stack/redundant-libraries — Redundant libraries | Several libraries for the same need (dates, state, UI, auth, routing…). | −6 points per category (−3 if low impact), max −30. |
stack/unused-dependencies — Unused dependencies | Declared dependencies that are never imported. | To verify. −1.5 points per dependency (max −12), counts for half. |
stack/multiple-lockfiles — Multiple package managers | Several lockfiles (npm, yarn, pnpm, bun) in the same folder. | Medium. −10 points. |
stack/no-lockfile — No lockfile | Dependencies without a lockfile: installs are not reproducible. | Medium. −8 points. |
stack/deprecated — Abandoned dependencies | Officially deprecated or archived packages. | −5 points per package (−2 if minor), max −20. |
Docs & AI context (weight 0.7)
| Rule | What is measured | Penalty |
|---|---|---|
docs/ai-working-docs — AI-generated working documents | Markdown report files (SUMMARY, PLAN, FIXES, PHASE…) left in the project. | −3 points per document (max −30). |
docs/stale-agent-instructions — Stale agent instructions | CLAUDE.md, AGENTS.md, .cursorrules… that reference files or commands that no longer exist. | −8 points + 2 per broken reference (max −25). |
docs/broken-references — Outdated documentation | Files and commands mentioned in the documentation that no longer exist. | −2 points per reference (max −15). |
docs/multiple-agent-files — Scattered AI instructions | Several agent instruction formats side by side. | −5 points (−8 from 3 tools). |
docs/no-agent-instructions — No instructions for AI agents | No AGENTS.md, CLAUDE.md, .cursorrules… file in a non-trivial project. | Minor. −6 points. |
docs/readme — Missing or generic README | README missing or still the one from the starter template. | −12 points if missing, −10 if generic. |
Limitations
- The analysis is static: it doesn’t prove that the app works; it measures risk and maintainability.
- Detection of dead code and unprotected routes relies on heuristics; these results are marked “to verify.”
- Functions (length, duplicates) are analyzed in JavaScript, TypeScript, Python, Rust, Go, Java, Kotlin, Swift, C#, C, C++, PHP and Scala; the import graph (dead code, cycles, broken imports) in JavaScript, TypeScript, Python and Rust. Ruby, Dart, Lua and Objective-C are covered by the generic rules (size, duplication, secrets, documentation).
- Outside JavaScript and TypeScript, comments and docstrings don’t count as lines of code, and only near-identical functions sharing a name are reported: each language has its own naming conventions.