No description
  • Rust 64%
  • TypeScript 32.7%
  • CSS 2.7%
  • Dockerfile 0.5%
  • HTML 0.1%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Arzaroth 5dd6cc650f [master] chore(release): 0.4.0
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 03:42:54 +02:00
.claude/skills [feature/todo-sweep] docs: branches, watch, report, the alarm, the drill-down and the new module rule; TODO emptied 2026-09-29 03:02:19 +02:00
crates [feature/todo-sweep] refactor(web): one percent formatter over the share derive already computes, file names made relative in derive, and the dead fallbacks gone 2026-09-29 03:34:09 +02:00
web [feature/todo-sweep] refactor(web): one percent formatter over the share derive already computes, file names made relative in derive, and the dead fallbacks gone 2026-09-29 03:34:09 +02:00
web-dist [master] feat(server): the api the dashboard reads, and the dashboard itself 2026-09-06 22:26:37 +02:00
.dockerignore [master] feat(docker): a static binary in a sixteen megabyte image 2026-09-06 22:58:58 +02:00
.gitignore [master] fix(coverage): a new file section closes an unclosed one 2026-09-07 00:40:27 +02:00
Cargo.lock [master] chore(release): 0.4.0 2026-09-29 03:42:54 +02:00
Cargo.toml [master] chore(release): 0.4.0 2026-09-29 03:42:54 +02:00
compose.yaml [master] feat(docker): repositories at the paths the store already knows 2026-09-06 22:58:58 +02:00
Dockerfile [master] feat(docker): a static binary in a sixteen megabyte image 2026-09-06 22:58:58 +02:00
mise.toml [master] fix(mise): declare cargo-llvm-cov, and find LLVM without rustup 2026-09-08 16:40:34 +02:00
README.md [feature/todo-sweep] docs: what review changed - numbered branch slugs, unambiguous commit prefixes, in-file tests clearing the alarm, and what older coverage cannot follow 2026-09-29 03:34:55 +02:00
TODO.md [master] docs(release): 0.4.0 sweep 2026-09-29 03:41:51 +02:00

Codegauge

A codebase observatory. It walks a repository's git history, counts every line of every tracked file at one snapshot per UTC day, and serves a dashboard that charts how production code and test code moved against each other over time.

One static binary with the frontend baked in. Nothing to install on the machine that runs it, and no git on the PATH: history comes from gitoxide, in process.

Use

mise run build                   # dashboard first, then the binary
mise run scan ~/repos/my-project
mise run serve

mise run check type-checks the dashboard and lints the Rust; mise run test runs the suite; mise run install symlinks the binary into ~/.local/bin.

codegauge scan <repo-path...>   --branch <name>  --rescan  --name <label>
codegauge serve                 --port <n>  --host <addr>  --open
codegauge cover <slug>          --report <file>    read a coverage report
codegauge report <slug>         --at <day|commit>  --since <day|commit>  --metric <m>  --json
codegauge watch                 --every <30s|15m|1h>  --once
codegauge list                  tracked repositories
codegauge forget <slug>         drop a repository and its snapshots
codegauge where                 print the data directory

Data lives in $XDG_DATA_HOME/codegauge/codegauge.db, overridable with CODEGAUGE_HOME or --db.

Scanning is incremental: stored days are skipped, and line counts are cached per git blob, so a file that never changes is counted once for the whole history no matter how many snapshots it appears in. Eleven repositories and 4,908 daily snapshots take 1.6 seconds cold on sixteen cores.

Each branch of a repository is tracked on its own. The first branch scanned keeps the repository's slug; scanning another, with --branch or from a checkout sitting on it, adds a row beside it slugged <slug>@<branch> (numbered when two branch names slugify alike, as feature/login and feature-login do), and the dashboard offers a branch picker once a repository has more than one. Checking out a feature branch for an afternoon therefore leaves the main history where it was. Every branch shares the repository's configuration, keyed by the first slug.

Keeping snapshots fresh

A rescan costs about as much as the day's diff, so the only reason a dashboard goes stale is that nothing asks it not to. codegauge watch rescans every tracked repository every fifteen minutes, printing a line only for the ones that moved; a repository that cannot be read is reported and skipped. --once runs a single pass for a timer to call instead:

# ~/.config/systemd/user/codegauge.service
[Service]
Type=oneshot
ExecStart=%h/.local/bin/codegauge watch --once

# ~/.config/systemd/user/codegauge.timer
[Timer]
OnBootSec=5min
OnUnitActiveSec=15min

[Install]
WantedBy=timers.target

systemctl --user enable --now codegauge.timer starts it.

For CI

codegauge report <slug> describes one stored day - named by the day or by a commit it ended on, a prefix of at least four characters that names only one stored commit - with its role totals, test ratio, per-module rows and the coverage report in force. --since adds what moved against a baseline, per role and per module, including modules that appeared or vanished; --json prints the same thing as data for a pull-request comment to be built from. It carries no verdict and there is no --fail-under, for the reason given under On the test ratio.

Stack

Layer Choice Why
Git access gix Tree traversal and object reads in process; no subprocess per day
Counting rayon The hot loop is byte scanning over thousands of blobs, and it is embarrassingly parallel
Diffs imara-diff git diff --numstat semantics without shelling out
Storage rusqlite One file, one writer, real transactions
API axum on tokio Server is async; the store is only ever locked on spawn_blocking, since a scan holds it for its whole length, and the scan runs there with rayon because it is CPU-bound
Scan progress server-sent events, a tokio watch One way, over the request that starts the scan; a reader slower than the scan skips to the newest state instead of working through a backlog of days
Assets rust-embed The dashboard ships inside the binary
Frontend Solid, Vite Signals own state and shell; no virtual DOM to sit in the hot path
Charts hand-written SVG See below

The charts are not components. A crosshair tracking the pointer and a brush that re-slices six charts per drag frame want to write to the DOM directly, so web/src/lib/charts.ts is framework-free and imperative and ChartCanvas.tsx is the seam, calling update() so a pinned point survives a re-render.

Nothing above a chart is a singleton. createDashboard() builds the store, and the tooltip with it, so a test gets state nobody else has touched; DashboardProvider hands it down. The arithmetic behind every panel lives in web/src/lib/derive.ts as plain functions over plain data, because a number worth checking should not need a mounted chart to reach it.

What it measures

Every file is assigned a role, and the production/test split is the story the dashboard tells:

Role What lands here
production authored source that ships
test test/, spec/, e2e/, __tests__/, plus *.test.*, *_test.go, test_*.py, *Test.java …
docs Markdown, reStructuredText, AsciiDoc, plain text
config JSON, YAML, TOML, INI, Dockerfiles, Terraform …
other generated output: *.g.dart, *.pb.go, migrations/, drizzle/, snapshots, minified bundles

Lockfiles, node_modules, dist, vendor, build directories and binaries are not counted at all. Four measures are available: nonblank lines (default), code lines with comments stripped, all lines, and bytes. Comment detection is per language, and a line counts as a comment only when it holds no code.

A file can hold both roles. Rust writes unit tests in #[cfg(test)] modules and Zig in test blocks, so for those languages the region is counted as test and the rest of the file as whatever the file is. Without it, a Rust project reads as having almost no tests: TokenGauge measured 0.02 tests per production line and is really 0.35. Only languages with an actual convention are read this way - guessing at one that has none would be worse than the undercount. The file is still counted once, against the role it is filed as.

A .d.ts in the source tree is filed as production: plenty are written by hand, and emitted ones land in dist/ or a generated directory, which the directory rules already catch.

Alongside the role, each file carries a zone (top-level area, with monorepo containers like apps/ expanded one level: apps/web-nuxt), a module (one level deeper: apps/web-nuxt/server), and a language. Every chart slices on those.

A module is named by what it is, not by where its code or its tests sit. A leading src/, lib/, source/ or sources/ is looked through, and so is a test directory or a main/ below it, so crates/core/src and crates/core/tests are both crates/core, web/src/lib and web/tests/lib are both web/lib, and Maven's src/main/java meets src/test/java. Without it, Rust's convention of integration tests beside the crate reads as a module with no tests next to a module with nothing else.

The rules above are versioned in the store. A release that files the same path differently drops the stored facts on first open and the next scan rebuilds them from the blob cache, which the change leaves alone.

In a container

mise run docker:build                    # image user takes your uid and gid
mise run docker:serve                    # dashboard on 127.0.0.1:7777
mise run docker:scan ~/repos/my-project

The runtime image is Alpine plus a static musl binary, 16 MB, and it contains no git: history comes from gitoxide, in process.

Two things about the mounts are deliberate. Repositories go in read-only, because scanning never writes to one. And they are mounted at the same path they have on the host, because the snapshot store records each repository's absolute path - identical paths are what let a container and a codegauge on the host share one database, and what lets "Update snapshots" in the dashboard find the repository it is talking about.

CG_REPOS, CG_DATA and CG_ARGS override the repository root, the snapshot store and the command.

Tests

mise run test       # 332 Rust, 197 dashboard
mise run coverage   # both halves, merged into coverage/all.lcov
mise run cover:self # and then read by codegauge

529 tests. The dashboard covers 92.0% of its 1,079 lines, measured with vitest --coverage. The Rust half covered 95.5% at v0.3.0 with cargo llvm-cov, 94.2% of 3,844 lines merged - a number codegauge reads back with its own cover command - and has not been measured since.

The scanner is tested against real repositories, built commit by commit with gitoxide rather than by shelling out, so the suite keeps the property the binary has: no git needed on the machine. Every fixture commit takes an explicit timestamp, because daily snapshots, gap filling and "the last commit of the day wins" are all arithmetic on committer time and none of it is testable against commits stamped now.

Writing them turned up six real defects, which is the argument for having written them:

count a line-comment token shadowed a longer block opener, so Lua block comments never opened and their contents counted as code
classify a generated-file suffix that could never match, because the language lookup runs first
server unknown /api/ routes fell through to the single-page fallback and answered HTML with a 200
server creating a repository returned 200 where the previous implementation returned 201
web the headline figures said "lines" whatever the measure was: 4.2 MB lines of production code
coverage concatenated reports lost a whole file when the first omitted its final newline - found by pointing codegauge at itself

Telling it about your repository

Every repository has an idiosyncrasy, and editing the classifier and rebuilding is absurd ceremony for it. A .codegauge.toml at the root says what the built-in rules cannot work out:

name = "Ouroboros"
exclude = ["generated/**"]        # not authored source, on top of the built-in list
containers = ["tier"]             # directories whose children are the unit, like apps/

[roles]
"src/api/*.d.ts" = "other"        # emitted by codegen, not written

# Where the tests are, when the language will not say.
[[tests]]
paths = ["ouroboros.py"]
functions = ["_selftest", "fuzz"]                       # a definition and its body
spans = [["# Layer 17: differential fuzzing", "# ="]]   # from a line to the next marker

functions matches a definition whose name starts with the pattern - so fuzz takes fuzz, fuzz_limits and fuzz_refusals - and takes its body, by braces or by indentation depending on the language. spans runs from a line to the next line matching the end pattern, skipping any that immediately follow the start, so a banner whose closing rule looks like its opening one does not end the section it introduces.

The file describes the repository as it is now, so it is read from the working tree rather than from each commit and applies to the whole history. Changing it re-scans the repository, since the stored days were measured under different rules. Coverage follows without cover being run again: the report is kept per file, and every scan rolls it through the rules it just used.

Without it, ouroboros - a self-hosting compiler in one 22k-line Python file, with a differential fuzzer and a self-test inside it - reads as 0.00 tests per production line. With it, 0.05.

For repositories that are not yours

A clone you do not own cannot carry that file, and a rule that holds for every repository has nowhere to live inside any one of them. The same settings are therefore readable from $XDG_CONFIG_HOME/codegauge/config.toml, under [default] for every repository and under [repo.<slug>] for one, named by the slug that codegauge list prints:

[default]
exclude = ["**/mise-tasks/**"]

[repo.axeo-suite]
exclude = ["apps/processing-pipeline/modelconverter-cpp/include/**"]

[repo.otel-demo]
name = "OpenTelemetry Demo"

One key is global-only. Coverage reports are build output, so they are gitignored in most repositories and there is nothing to discover until something has produced one:

[repo.codegauge.coverage]
command = "mise run coverage"      # run this first
report  = "coverage/all.lcov"      # then read this

codegauge cover <slug> runs the command in the repository root, streaming its output, and reads the report it leaves behind. --report skips the command.

coverage.command is refused in a repository's own .codegauge.toml, and that is the point of the split: that file is read from a working tree you do not always own, and codegauge otherwise spawns no subprocess at all - not even for git. Naming a path describes the repository; naming a command borrows your shell. A repository may still set coverage.report.

Three layers apply in order: the global [default], the repository's own .codegauge.toml, then the global [repo.<slug>]. Lists concatenate and the last matching role wins, so the layer you can always edit is the layer that wins. A name set globally renames a repository without changing its slug, since the slug is what the section is keyed by.

The fingerprint that decides whether a repository must be measured again is taken over the merged result, so adding a section for one repository leaves every other repository's stored days alone - and moving a rule between layers without changing what it means is not a rescan.

Coverage

Coverage cannot be reconstructed from history the way line counts can. Line counts come from the tree; coverage only exists where a suite actually ran. Codegauge does not pretend otherwise: it reads a report a run already produced and stores one point per commit, rather than inventing a daily series.

codegauge cover my-project                       # finds a well-known report
codegauge cover my-project --report build/lcov.info

Three formats are read directly:

  • LCOV (lcov.info) — the closest thing to a lingua franca. Vitest and Jest via c8, flutter test --coverage, cargo llvm-cov --lcov, gcov, coverage.py lcov, Elixir and PHP all emit it.
  • Istanbul JSON summary (coverage-summary.json).
  • Go cover profiles (coverage.out). Go counts statements, not lines; those are summed into the line columns, which is the closest honest mapping.

JaCoCo and Cobertura are not read directly; convert to LCOV first.

LCOV being the lingua franca is not only a claim about tooling. ouroboros, a self-hosting compiler in a single dependency-free Python file, writes its own tracefile from its own sys.settrace tracer with no coverage library at all, taking executable lines from co_lines() rather than the text because a third of that file is string literals holding other languages. Codegauge reads it without knowing any of that.

The part that makes this uniform across tools is path matching. Reports disagree about what a path is relative to — absolute build paths, a package root, the repository root. Codegauge matches each report path onto the longest suffix that exists in the tree at that commit, which handles all three without per-tool configuration, and tells you when paths match nothing so a stale report is obvious:

nostragoalus: istanbul format, 0 files matched, no measurable lines
  208 report paths matched nothing in the tree (stale report?)

Coverage rolls up through the same (role, language, zone, module) cube as the line counts, so both slice identically. The breakdown draws each row's covered share beside its bars, and a row opened onto its files shows it per file. The report itself is kept per file too, so a change of rules rolls it again instead of asking for it a second time.

Reports read before per-file keeping (v0.3.0 and earlier) have only the cube, and a cube cannot be sliced again. They keep their totals, so the coverage chart is unchanged, but they stay sliced by the rules they were read under: after the module rule changed, their per-module figures no longer line up with the breakdown until cover reads a report again.

On the test ratio

The dashboard leads with test lines per production line because the number is cheap and legible, not because a particular value is correct. Treat it as a lens:

  • The slope matters, the level does not. A repository holding 0.5 while production grows forty-fold is healthier than one sitting at 2.0 and flat.
  • Per module beats per repository. A single number hides the module with four thousand lines and no tests. The breakdown panel is where the signal is.
  • There is no target line and no red/green verdict, on purpose. The cheapest way to move a ratio is verbose setup and snapshot dumps, and a threshold invites exactly that.

The one alarm it does raise is an event, not a level: a day where production moved by fifty lines or more, additions and removals together, and tests did not move at all. It is marked above that day in the movement chart and counted beside it, and /series?floor= moves the fifty. Padding the test side cannot make a past day go away, which is what a ratio threshold gets wrong. Churn is split by the role a file is filed as, so tests are judged twice: no test file changed, and the test lines in the day's facts, which count Rust and Zig in-file tests, did not change either.

How a snapshot is built

For each UTC day, the last commit of that day is chosen; days with no commits carry the previous tree forward, so the series is continuous. A merge can be that commit, since its tree is the real state of the branch, but it is not counted as a commit or credited to an author. gix lists the tree, each blob is classified by path, and anything not already in the blob cache is counted in parallel. The result rolls into a (day, role, language, zone, module) cube, and consecutive snapshots are diffed for the daily movement figures. While a day's files are counted, the rayon workers bump a counter the scan reads every tenth of a second, so the progress stream moves through a large day instead of sitting at its start.

The cube stops at module. Per-path facts for every day would be orders of magnitude larger and almost never read, so the files inside a row are computed when asked for, for one day, from that day's tree and the blob cache.

Fact totals were verified against the previous Node implementation across 49,494 rows and eleven repositories: identical everywhere except empty files, which the old implementation counted as one line and this one counts as zero, matching wc -l. Churn totals track git diff --numstat to within 0.003%; the residue is inherent, since imara-diff and git's xdiff pick different but equally minimal edit scripts on a handful of files.

Layout

crates/core/     classify, count, git, scan, coverage, report, store
crates/server/   axum routes, embedded assets
crates/cli/      the codegauge binary
web/             Solid dashboard
web-dist/        vite output, embedded at compile time

API: /api/repos (each row carries its family, shared by a repository's branches), /api/repo/:slug, /api/repo/:slug/series (with an untested flag per day, ?floor= to move it), /api/repo/:slug/breakdown?by=zone|module|lang, /api/repo/:slug/files?by=…&key=…&day=… (the files inside one breakdown row, 404 for a day with no snapshot), /api/repo/:slug/authors, /api/repo/:slug/coverage, /api/repo/:slug/export (CSV), and POST /api/repo/:slug/rescan. Sent with Accept: text/event-stream, the rescan answers with server-sent events instead: progress while it runs, then done with the report or failed with the error. A second rescan of a repository already being scanned answers 409 rather than queueing behind the first.

A request that panics while holding the store poisons its lock. The server takes the lock back instead of failing every later request, because SQLite has already rolled back whatever transaction the panic abandoned.

The dashboard loads Big Shoulders Display and Karla from Google Fonts, falling back to system faces offline.

Node and pnpm are pinned in mise.toml. Rust is not, so it comes from whatever toolchain is on the machine; rust-version in Cargo.toml records the floor.