- Rust 64%
- TypeScript 32.7%
- CSS 2.7%
- Dockerfile 0.5%
- HTML 0.1%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| .claude/skills | ||
| crates | ||
| web | ||
| web-dist | ||
| .dockerignore | ||
| .gitignore | ||
| Cargo.lock | ||
| Cargo.toml | ||
| compose.yaml | ||
| Dockerfile | ||
| mise.toml | ||
| README.md | ||
| TODO.md | ||
Codegauge
A codebase observatory. It walks a repository's git history, counts every line of every tracked file at one snapshot per UTC day, and serves a dashboard that charts how production code and test code moved against each other over time.
One static binary with the frontend baked in. Nothing to install on the machine that
runs it, and no git on the PATH: history comes from gitoxide, in process.
Use
mise run build # dashboard first, then the binary
mise run scan ~/repos/my-project
mise run serve
mise run check type-checks the dashboard and lints the Rust; mise run test runs
the suite; mise run install symlinks the binary into ~/.local/bin.
codegauge scan <repo-path...> --branch <name> --rescan --name <label>
codegauge serve --port <n> --host <addr> --open
codegauge cover <slug> --report <file> read a coverage report
codegauge report <slug> --at <day|commit> --since <day|commit> --metric <m> --json
codegauge watch --every <30s|15m|1h> --once
codegauge list tracked repositories
codegauge forget <slug> drop a repository and its snapshots
codegauge where print the data directory
Data lives in $XDG_DATA_HOME/codegauge/codegauge.db, overridable with
CODEGAUGE_HOME or --db.
Scanning is incremental: stored days are skipped, and line counts are cached per git blob, so a file that never changes is counted once for the whole history no matter how many snapshots it appears in. Eleven repositories and 4,908 daily snapshots take 1.6 seconds cold on sixteen cores.
Each branch of a repository is tracked on its own. The first branch scanned keeps
the repository's slug; scanning another, with --branch or from a checkout sitting
on it, adds a row beside it slugged <slug>@<branch> (numbered when two branch
names slugify alike, as feature/login and feature-login do), and the dashboard offers a
branch picker once a repository has more than one. Checking out a feature branch
for an afternoon therefore leaves the main history where it was. Every branch shares
the repository's configuration, keyed by the first slug.
Keeping snapshots fresh
A rescan costs about as much as the day's diff, so the only reason a dashboard goes
stale is that nothing asks it not to. codegauge watch rescans every tracked
repository every fifteen minutes, printing a line only for the ones that moved; a
repository that cannot be read is reported and skipped. --once runs a single pass
for a timer to call instead:
# ~/.config/systemd/user/codegauge.service
[Service]
Type=oneshot
ExecStart=%h/.local/bin/codegauge watch --once
# ~/.config/systemd/user/codegauge.timer
[Timer]
OnBootSec=5min
OnUnitActiveSec=15min
[Install]
WantedBy=timers.target
systemctl --user enable --now codegauge.timer starts it.
For CI
codegauge report <slug> describes one stored day - named by the day or by a
commit it ended on, a prefix of at least four characters that names only one stored
commit - with its role totals, test ratio, per-module rows and the
coverage report in force. --since adds what moved against a baseline, per role
and per module, including modules that appeared or vanished; --json prints the
same thing as data for a pull-request comment to be built from. It carries no
verdict and there is no --fail-under, for the reason given under
On the test ratio.
Stack
| Layer | Choice | Why |
|---|---|---|
| Git access | gix |
Tree traversal and object reads in process; no subprocess per day |
| Counting | rayon |
The hot loop is byte scanning over thousands of blobs, and it is embarrassingly parallel |
| Diffs | imara-diff |
git diff --numstat semantics without shelling out |
| Storage | rusqlite |
One file, one writer, real transactions |
| API | axum on tokio |
Server is async; the store is only ever locked on spawn_blocking, since a scan holds it for its whole length, and the scan runs there with rayon because it is CPU-bound |
| Scan progress | server-sent events, a tokio watch |
One way, over the request that starts the scan; a reader slower than the scan skips to the newest state instead of working through a backlog of days |
| Assets | rust-embed |
The dashboard ships inside the binary |
| Frontend | Solid, Vite | Signals own state and shell; no virtual DOM to sit in the hot path |
| Charts | hand-written SVG | See below |
The charts are not components. A crosshair tracking the pointer and a brush that
re-slices six charts per drag frame want to write to the DOM directly, so
web/src/lib/charts.ts is framework-free and imperative and ChartCanvas.tsx is the
seam, calling update() so a pinned point survives a re-render.
Nothing above a chart is a singleton. createDashboard() builds the store, and the
tooltip with it, so a test gets state nobody else has touched; DashboardProvider
hands it down. The arithmetic behind every panel lives in web/src/lib/derive.ts as
plain functions over plain data, because a number worth checking should not need a
mounted chart to reach it.
What it measures
Every file is assigned a role, and the production/test split is the story the dashboard tells:
| Role | What lands here |
|---|---|
production |
authored source that ships |
test |
test/, spec/, e2e/, __tests__/, plus *.test.*, *_test.go, test_*.py, *Test.java … |
docs |
Markdown, reStructuredText, AsciiDoc, plain text |
config |
JSON, YAML, TOML, INI, Dockerfiles, Terraform … |
other |
generated output: *.g.dart, *.pb.go, migrations/, drizzle/, snapshots, minified bundles |
Lockfiles, node_modules, dist, vendor, build directories and binaries are not
counted at all. Four measures are available: nonblank lines (default), code lines with
comments stripped, all lines, and bytes. Comment detection is per language, and a line
counts as a comment only when it holds no code.
A file can hold both roles. Rust writes unit tests in #[cfg(test)] modules and
Zig in test blocks, so for those languages the region is counted as test and the
rest of the file as whatever the file is. Without it, a Rust project reads as
having almost no tests: TokenGauge measured 0.02 tests per production line and is
really 0.35. Only languages with an actual convention are read this way - guessing
at one that has none would be worse than the undercount. The file is still counted
once, against the role it is filed as.
A .d.ts in the source tree is filed as production: plenty are written by hand,
and emitted ones land in dist/ or a generated directory, which the directory rules
already catch.
Alongside the role, each file carries a zone (top-level area, with monorepo
containers like apps/ expanded one level: apps/web-nuxt), a module (one level
deeper: apps/web-nuxt/server), and a language. Every chart slices on those.
A module is named by what it is, not by where its code or its tests sit. A leading
src/, lib/, source/ or sources/ is looked through, and so is a test
directory or a main/ below it, so crates/core/src and crates/core/tests are
both crates/core, web/src/lib and web/tests/lib are both web/lib, and
Maven's src/main/java meets src/test/java. Without it, Rust's convention of
integration tests beside the crate reads as a module with no tests next to a
module with nothing else.
The rules above are versioned in the store. A release that files the same path differently drops the stored facts on first open and the next scan rebuilds them from the blob cache, which the change leaves alone.
In a container
mise run docker:build # image user takes your uid and gid
mise run docker:serve # dashboard on 127.0.0.1:7777
mise run docker:scan ~/repos/my-project
The runtime image is Alpine plus a static musl binary, 16 MB, and it contains no git: history comes from gitoxide, in process.
Two things about the mounts are deliberate. Repositories go in read-only,
because scanning never writes to one. And they are mounted at the same path they
have on the host, because the snapshot store records each repository's absolute
path - identical paths are what let a container and a codegauge on the host share
one database, and what lets "Update snapshots" in the dashboard find the repository
it is talking about.
CG_REPOS, CG_DATA and CG_ARGS override the repository root, the snapshot
store and the command.
Tests
mise run test # 332 Rust, 197 dashboard
mise run coverage # both halves, merged into coverage/all.lcov
mise run cover:self # and then read by codegauge
529 tests. The dashboard covers 92.0% of its 1,079 lines, measured with
vitest --coverage. The Rust half covered 95.5% at v0.3.0 with cargo llvm-cov,
94.2% of 3,844 lines merged - a number codegauge reads back with its own cover
command - and has not been measured since.
The scanner is tested against real repositories, built commit by commit with
gitoxide rather than by shelling out, so the suite keeps the property the binary
has: no git needed on the machine. Every fixture commit takes an explicit
timestamp, because daily snapshots, gap filling and "the last commit of the day
wins" are all arithmetic on committer time and none of it is testable against
commits stamped now.
Writing them turned up six real defects, which is the argument for having written them:
count |
a line-comment token shadowed a longer block opener, so Lua block comments never opened and their contents counted as code |
classify |
a generated-file suffix that could never match, because the language lookup runs first |
server |
unknown /api/ routes fell through to the single-page fallback and answered HTML with a 200 |
server |
creating a repository returned 200 where the previous implementation returned 201 |
web |
the headline figures said "lines" whatever the measure was: 4.2 MB lines of production code |
coverage |
concatenated reports lost a whole file when the first omitted its final newline - found by pointing codegauge at itself |
Telling it about your repository
Every repository has an idiosyncrasy, and editing the classifier and rebuilding is
absurd ceremony for it. A .codegauge.toml at the root says what the built-in
rules cannot work out:
name = "Ouroboros"
exclude = ["generated/**"] # not authored source, on top of the built-in list
containers = ["tier"] # directories whose children are the unit, like apps/
[roles]
"src/api/*.d.ts" = "other" # emitted by codegen, not written
# Where the tests are, when the language will not say.
[[tests]]
paths = ["ouroboros.py"]
functions = ["_selftest", "fuzz"] # a definition and its body
spans = [["# Layer 17: differential fuzzing", "# ="]] # from a line to the next marker
functions matches a definition whose name starts with the pattern - so fuzz
takes fuzz, fuzz_limits and fuzz_refusals - and takes its body, by braces or
by indentation depending on the language. spans runs from a line to the next line
matching the end pattern, skipping any that immediately follow the start, so a
banner whose closing rule looks like its opening one does not end the section it
introduces.
The file describes the repository as it is now, so it is read from the working tree
rather than from each commit and applies to the whole history. Changing it re-scans
the repository, since the stored days were measured under different rules. Coverage
follows without cover being run again: the report is kept per file, and every scan
rolls it through the rules it just used.
Without it, ouroboros - a self-hosting compiler in one 22k-line Python file, with
a differential fuzzer and a self-test inside it - reads as 0.00 tests per production
line. With it, 0.05.
For repositories that are not yours
A clone you do not own cannot carry that file, and a rule that holds for every
repository has nowhere to live inside any one of them. The same settings are
therefore readable from $XDG_CONFIG_HOME/codegauge/config.toml, under [default]
for every repository and under [repo.<slug>] for one, named by the slug that
codegauge list prints:
[default]
exclude = ["**/mise-tasks/**"]
[repo.axeo-suite]
exclude = ["apps/processing-pipeline/modelconverter-cpp/include/**"]
[repo.otel-demo]
name = "OpenTelemetry Demo"
One key is global-only. Coverage reports are build output, so they are gitignored in most repositories and there is nothing to discover until something has produced one:
[repo.codegauge.coverage]
command = "mise run coverage" # run this first
report = "coverage/all.lcov" # then read this
codegauge cover <slug> runs the command in the repository root, streaming its
output, and reads the report it leaves behind. --report skips the command.
coverage.command is refused in a repository's own .codegauge.toml, and that
is the point of the split: that file is read from a working tree you do not
always own, and codegauge otherwise spawns no subprocess at all - not even for
git. Naming a path describes the repository; naming a command borrows your
shell. A repository may still set coverage.report.
Three layers apply in order: the global [default], the repository's own
.codegauge.toml, then the global [repo.<slug>]. Lists concatenate and the last
matching role wins, so the layer you can always edit is the layer that wins. A name
set globally renames a repository without changing its slug, since the slug is what
the section is keyed by.
The fingerprint that decides whether a repository must be measured again is taken over the merged result, so adding a section for one repository leaves every other repository's stored days alone - and moving a rule between layers without changing what it means is not a rescan.
Coverage
Coverage cannot be reconstructed from history the way line counts can. Line counts come from the tree; coverage only exists where a suite actually ran. Codegauge does not pretend otherwise: it reads a report a run already produced and stores one point per commit, rather than inventing a daily series.
codegauge cover my-project # finds a well-known report
codegauge cover my-project --report build/lcov.info
Three formats are read directly:
- LCOV (
lcov.info) — the closest thing to a lingua franca. Vitest and Jest via c8,flutter test --coverage,cargo llvm-cov --lcov, gcov,coverage.py lcov, Elixir and PHP all emit it. - Istanbul JSON summary (
coverage-summary.json). - Go cover profiles (
coverage.out). Go counts statements, not lines; those are summed into the line columns, which is the closest honest mapping.
JaCoCo and Cobertura are not read directly; convert to LCOV first.
LCOV being the lingua franca is not only a claim about tooling. ouroboros, a
self-hosting compiler in a single dependency-free Python file, writes its own
tracefile from its own sys.settrace tracer with no coverage library at all,
taking executable lines from co_lines() rather than the text because a third of
that file is string literals holding other languages. Codegauge reads it without
knowing any of that.
The part that makes this uniform across tools is path matching. Reports disagree about what a path is relative to — absolute build paths, a package root, the repository root. Codegauge matches each report path onto the longest suffix that exists in the tree at that commit, which handles all three without per-tool configuration, and tells you when paths match nothing so a stale report is obvious:
nostragoalus: istanbul format, 0 files matched, no measurable lines
208 report paths matched nothing in the tree (stale report?)
Coverage rolls up through the same (role, language, zone, module) cube as the line counts, so both slice identically. The breakdown draws each row's covered share beside its bars, and a row opened onto its files shows it per file. The report itself is kept per file too, so a change of rules rolls it again instead of asking for it a second time.
Reports read before per-file keeping (v0.3.0 and earlier) have only the cube, and a
cube cannot be sliced again. They keep their totals, so the coverage chart is
unchanged, but they stay sliced by the rules they were read under: after the module
rule changed, their per-module figures no longer line up with the breakdown until
cover reads a report again.
On the test ratio
The dashboard leads with test lines per production line because the number is cheap and legible, not because a particular value is correct. Treat it as a lens:
- The slope matters, the level does not. A repository holding 0.5 while production grows forty-fold is healthier than one sitting at 2.0 and flat.
- Per module beats per repository. A single number hides the module with four thousand lines and no tests. The breakdown panel is where the signal is.
- There is no target line and no red/green verdict, on purpose. The cheapest way to move a ratio is verbose setup and snapshot dumps, and a threshold invites exactly that.
The one alarm it does raise is an event, not a level: a day where production moved by
fifty lines or more, additions and removals together, and tests did not move at
all. It is marked above that day in the movement chart and counted beside it, and
/series?floor= moves the fifty. Padding the test side cannot make a past day go
away, which is what a ratio threshold gets wrong. Churn is split by the role a file
is filed as, so tests are judged twice: no test file changed, and the test lines in
the day's facts, which count Rust and Zig in-file tests, did not change either.
How a snapshot is built
For each UTC day, the last commit of that day is chosen; days with no commits carry the
previous tree forward, so the series is continuous. A merge can be that commit, since
its tree is the real state of the branch, but it is not counted as a commit or
credited to an author. gix lists the tree, each blob is
classified by path, and anything not already in the blob cache is counted in parallel.
The result rolls into a (day, role, language, zone, module) cube, and consecutive
snapshots are diffed for the daily movement figures. While a day's files are counted,
the rayon workers bump a counter the scan reads every tenth of a second, so the
progress stream moves through a large day instead of sitting at its start.
The cube stops at module. Per-path facts for every day would be orders of magnitude larger and almost never read, so the files inside a row are computed when asked for, for one day, from that day's tree and the blob cache.
Fact totals were verified against the previous Node implementation across 49,494 rows
and eleven repositories: identical everywhere except empty files, which the old
implementation counted as one line and this one counts as zero, matching wc -l.
Churn totals track git diff --numstat to within 0.003%; the residue is inherent,
since imara-diff and git's xdiff pick different but equally minimal edit scripts on a
handful of files.
Layout
crates/core/ classify, count, git, scan, coverage, report, store
crates/server/ axum routes, embedded assets
crates/cli/ the codegauge binary
web/ Solid dashboard
web-dist/ vite output, embedded at compile time
API: /api/repos (each row carries its family, shared by a repository's
branches), /api/repo/:slug, /api/repo/:slug/series (with an untested flag per
day, ?floor= to move it), /api/repo/:slug/breakdown?by=zone|module|lang,
/api/repo/:slug/files?by=…&key=…&day=… (the files inside one breakdown row, 404
for a day with no snapshot),
/api/repo/:slug/authors, /api/repo/:slug/coverage,
/api/repo/:slug/export (CSV), and POST /api/repo/:slug/rescan. Sent with
Accept: text/event-stream, the rescan answers with server-sent events instead:
progress while it runs, then done with the report or failed with the error. A
second rescan of a repository already being scanned answers 409 rather than
queueing behind the first.
A request that panics while holding the store poisons its lock. The server takes the lock back instead of failing every later request, because SQLite has already rolled back whatever transaction the panic abandoned.
The dashboard loads Big Shoulders Display and Karla from Google Fonts, falling back to system faces offline.
Node and pnpm are pinned in mise.toml. Rust is not, so it comes from whatever
toolchain is on the machine; rust-version in Cargo.toml records the floor.