Download CONTRIBUTING.md from mlanvvs/fitcheck: direct link, hf CLI and curl.
- Browser
- Download file 7.81 kB
-
https://huggingface.co/spaces/mlanvvs/fitcheck/resolve/main/CONTRIBUTING.md
- Command line
-
hf download hf://spaces/mlanvvs/fitcheck/CONTRIBUTING.md
-
curl -L -o CONTRIBUTING.md https://huggingface.co/spaces/mlanvvs/fitcheck/resolve/main/CONTRIBUTING.md
A newer version of the Gradio SDK is available: 6.29.0
Contributing to fitcheck
Fork, branch off main, open a PR. Please keep changes to one memory component per PR where
possible β the modules under fitcheck/memory/ are deliberately independent so a formula can be
argued about in isolation.
The most useful thing you can contribute
A measured row on hardware that is not a Tesla T4.
Every archived measurement comes from one T4 (sm_75), which means BF16 and real FlashAttention-2
β both need sm_80 or newer β have never been exercised. The CUDA context was measured at 140.875
MiB on 69 archived rows and 141.0 MiB on the other 10, all on that same card; there is still no
measured profile for another GPU. If you have an Ampere or newer GPU, one run of
scripts/measure.py is worth more to this project than any feature:
pip install -r scripts/requirements-measure.txt
python scripts/measure.py TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
--qlora --precision fp16 --lora-r 32 --batch-size 2 --seq-len 1024 --gpu t4
It prints prediction vs measurement at all three tiers, a per-component spot-check, and a markdown row ready to paste. Open it with the measurement issue template.
Calibrating your card
C_overhead β the CUDA context plus the caching allocator's fragmentation β is the one component
that cannot be derived from a config.json, because it belongs to a driver and a card rather than
to a model. The shipped fit is keyed by (GPU, attention kernel, quantization) and shipped as data in
fitcheck/overhead_db.py, so adding your card is a one-line source change:
python scripts/calibration_sweep.py --gpu <key> # ~20 rows, unattended, one process each
python -m fitcheck.calibrate runs/*.json # ad-hoc look at the fit and its residuals
Each row file is named after the whole identity of the run β card, kernel, quantization,
precision, optimizer, LoRA rank, checkpointing, model, batch size, sequence length β followed by a
fingerprint of that identity, so two configurations can never land on the same path. The identity
is written into the row as well (sweep.identity), the sweep refuses to skip or overwrite a file
whose identity does not match the row being asked for, and it rejects a row whose own run block
disagrees with what was requested. A sweep that did not complete every planned row exits
non-zero, so a grid with holes in it cannot be mistaken for a finished one.
Then commit the rows and declare them. data/measurements/manifest.json gives every archived
row a role β calibration, holdout, repeat or excluded (with a reason) β and only
calibration rows are fitted. Add one entry per row, then:
python -m fitcheck.calibrate data/measurements/manifest.json # the fit and its residuals
python -m fitcheck.calibrate data/measurements/manifest.json --emit-python # the OVERHEAD_DB literal
python -m fitcheck.calibrate data/measurements/manifest.json --check # grade what now ships
To grade an out-of-sample holdout, run --role holdout --check explicitly. Holdout rows are
deliberately absent from the default fit and from --emit-python; they measure generalization
after the coefficients are frozen.
Paste the --emit-python output over the OVERHEAD_DB literal in fitcheck/overhead_db.py.
Do not hand-edit it: tests/test_manifest.py re-runs that command and compares character for
character, so a typed coefficient fails CI. It also fails if the archive and the manifest disagree
by a single row, or if a manifest entry no longer matches the measurement it points at.
See the README under data/measurements/ for the row format and for what makes a row fittable. A
fitted profile without its rows in the repo is a number nobody can check; a row without a declared
role is a number that can change the fit without anybody deciding it should.
The paths with no measured row at all are listed under "What is not measured" in the README.
The largest gaps: any card that is not a T4 (which is also what C_overhead needs before a
per-card fit means anything), a second --quant int8 model, and an fitcheck infer concurrency
sweep.
The bar for a merge
pytest --cov=fitcheck --cov-report=term-missing -m "not network"is green. The standard suite has 877 offline tests, with 100% line coverage on all sevenmemory/modules; β₯80% there is the floor. Another 13 measurement-harness tests run when optionaltorchis installed. The-m "not network"filter is not optional: it skips the 7 tests markednetwork, which hit the Hub for real β two of them the gatedmeta-llama/Llama-3.1-8B, which fails without anHF_TOKEN. The offline tests cover the same parsing against a fixture.Numbers in the docs are checked, not trusted.
tests/test_docs_claims.pyrecomputes every quantitative claim inREADME.md,CLAUDE.mdand this file from the artifacts they describe β the measurement archive,fitcheck/overhead_db.py, and the test suite's own collection β and fails with the number you should have written. If you add a test or a measured row, run it.Any change to a formula updates its module, its test, and
docs/SPEC.mdin the same PR. The Llama-3.1-8B golden numbers in the SPEC appendix are the reference set β if a change moves them, say so explicitly in the PR description. A formula whose derivation and implementation disagree is how this project got a 36% error once already.Type hints and docstrings on public functions, dataclasses for configs, MiB returned as
float.ruff check .andmypy --strict fitcheck/are both clean. CI runs them as their own job, so a PR that fails either one does not merge. Both come withpip install -e ".[dev]":ruff check . # add --fix for the mechanical ones mypy --strict fitcheck/The configuration lives in
[tool.ruff]and[tool.mypy]inpyproject.toml. Two deliberate choices there: the notebooks are excluded, because they are dated measurement artifacts rather than maintained source, andRUF001-003are off, because the formulas are written with the same symbolsdocs/SPEC.mduses (Ξ³, Γ) and ASCII lookalikes would make the two disagree.--strictcoversfitcheck/only βscripts/measure.pyimportstorch, which is not installed in CI and must never become a dependency of the package.
Two constraints that are not negotiable
fitchecknever importstorch,peftorbitsandbytesβ not lazily, not inside atry. An estimate must cost a few KB ofconfig.jsonand no GPU; that is the whole product. Those libraries belong only inscripts/measure.py, which is not a runtime dependency and is not installed bypip install fitcheck-llm. The dependency runs one way:measure.pyimportsfitcheck, never the reverse.Only
config.jsonis ever fetched. Never weights, never a checkpoint.
Runtime dependencies are click, rich and huggingface-hub. Adding a fourth is a decision, not a
detail β raise it in an issue first.
Units
Report MiB (1024Β²) everywhere, never MB (10βΆ). Compute in bytes and convert once, at the boundary.
The 4.9% gap between the two is enough on its own to flip a fits/doesn't-fit verdict near the edge
of a card β and it is exactly how the T4 entry in gpu_db.py ended up claiming more usable memory
than the card physically has.
Where things live
docs/SPEC.md is the source of truth for formulas, derivations, and the golden numbers. Its
Component 5 contains the activation derivation. Section 3.8 explains what
scripts/measure.py does and why each part of it matters β read it before trusting a measurement,
including your own.