Alogotron commited on
Commit
b94c32f
Β·
verified Β·
1 Parent(s): b8e2703

Wire fine-tuned interpreter analysis

Browse files
Files changed (1) hide show
  1. AGENTS.md +177 -0
AGENTS.md ADDED
@@ -0,0 +1,177 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DOX framework
2
+
3
+ - DOX is highly performant AGENTS.md hierarchy installed here
4
+ - Agent must follow DOX instructions across any edits
5
+
6
+ ## Core Contract
7
+
8
+ - AGENTS.md files are binding work contracts for their subtrees
9
+ - Work products, source materials, instructions, records, assets, and durable docs must stay understandable from the nearest applicable AGENTS.md plus every parent AGENTS.md above it
10
+
11
+ ## Read Before Editing
12
+
13
+ 1. Read the root AGENTS.md
14
+ 2. Identify every file or folder you expect to touch
15
+ 3. Walk from the repository root to each target path
16
+ 4. Read every AGENTS.md found along each route
17
+ 5. If a parent AGENTS.md lists a child AGENTS.md whose scope contains the path, read that child and continue from there
18
+ 6. Use the nearest AGENTS.md as the local contract and parent docs for repo-wide rules
19
+ 7. If docs conflict, the closer doc controls local work details, but no child doc may weaken DOX
20
+
21
+ Do not rely on memory. Re-read the applicable DOX chain in the current session before editing.
22
+
23
+ ## Update After Editing
24
+
25
+ Every meaningful change requires a DOX pass before the task is done.
26
+
27
+ Update the closest owning AGENTS.md when a change affects:
28
+
29
+ - purpose, scope, ownership, or responsibilities
30
+ - durable structure, contracts, workflows, or operating rules
31
+ - required inputs, outputs, permissions, constraints, side effects, or artifacts
32
+ - user preferences about behavior, communication, process, organization, or quality
33
+ - AGENTS.md creation, deletion, move, rename, or index contents
34
+
35
+ Update parent docs when parent-level structure, ownership, workflow, or child index changes. Update child docs when parent changes alter local rules. Remove stale or contradictory text immediately. Small edits that do not change behavior or contracts may leave docs unchanged, but the DOX pass still must happen.
36
+
37
+ ## Hierarchy
38
+
39
+ - Root AGENTS.md is the DOX rail: project-wide instructions, global preferences, durable workflow rules, and the top-level Child DOX Index
40
+ - Child AGENTS.md files own domain-specific instructions and their own Child DOX Index
41
+ - Each parent explains what its direct children cover and what stays owned by the parent
42
+ - The closer a doc is to the work, the more specific and practical it must be
43
+
44
+ ## Child Doc Shape
45
+
46
+ - Create a child AGENTS.md when a folder becomes a durable boundary with its own purpose, rules, responsibilities, workflow, materials, or quality standards
47
+ - Work Guidance must reflect the current standards of the project or user instructions; if there are no specific standards or instructions yet, leave it empty
48
+ - Verification must reflect an existing check; if no verification framework exists yet, leave it empty and update it when one exists
49
+
50
+ Default section order:
51
+ - Purpose
52
+ - Ownership
53
+ - Local Contracts
54
+ - Work Guidance
55
+ - Verification
56
+ - Child DOX Index
57
+
58
+ ## Style
59
+
60
+ - Keep docs concise, current, and operational
61
+ - Document stable contracts, not diary entries
62
+ - Put broad rules in parent docs and concrete details in child docs
63
+ - Prefer direct bullets with explicit names
64
+ - Do not duplicate rules across many files unless each scope needs a local version
65
+ - Delete stale notes instead of explaining history
66
+ - Trim obvious statements, repeated rules, misplaced detail, and warnings for risks that no longer exist
67
+
68
+ ## Closeout
69
+
70
+ 1. Re-check changed paths against the DOX chain
71
+ 2. Update nearest owning docs and any affected parents or children
72
+ 3. Refresh every affected Child DOX Index
73
+ 4. Remove stale or contradictory text
74
+ 5. Run existing verification when relevant
75
+ 6. Report any docs intentionally left unchanged and why
76
+
77
+ ## User Preferences
78
+
79
+ When the user requests a durable behavior change, record it here or in the relevant child AGENTS.md
80
+
81
+ - User wants to be very ambitious for this hackathon
82
+ - User wants to go deep on manifold research, not just surface-level framing
83
+ - User prefers maximizing prize stacking across multiple categories
84
+
85
+ ---
86
+
87
+ ## Project: Build Small Hackathon (Gradio)
88
+
89
+ ### Purpose
90
+
91
+ Gradio "Build Small" hackathon project. Build an ambitious HF Space powered by ≀32B models using Gradio, deployed under the hackathon org. Aiming for multiple prize categories.
92
+
93
+ ### Concept Direction (current β€” pivoted)
94
+
95
+ **Activation Brain** β€” A live comparative interpretability visualizer: two Gemma streams drive dual EEGs, baseline-corrected emotion deltas, model-native meters, and a generated comparison analysis panel. The prior 3D brain view was replaced because a single brain was less clear once both models run simultaneously. The user picks between **two architecturally identical Gemma-4-12B models** β€” `google/gemma-4-12B-it` (base) and `OBLITERATUS/Gemma-4-12B-OBLITERATED` (abliterated/uncensored) β€” sharing ONE UMAP frame so firing differences are directly comparable (a censored-vs-uncensored interpretability demo).
96
+
97
+ **Pivot history:** Started as "Activation Avatars" (Qwen3-4B hidden states β†’ FLUX.2-Klein face + LivePortrait animation). Abandoned because per-token emotion mapping to a human face produced uncanny flicker (timescale/subspace mismatch β€” a face has affective inertia; fast multi-emotion change reads as noise). A brain/EEG is *expected* to fire fast across many regions simultaneously, so the pivot dissolved the core problem and deleted the fragile face stack (FLUX, LivePortrait, expression regressor).
98
+
99
+ **Why two Gemmas (not Qwen):** Qwen3-4B (36 layers, 2560 dim) is not comparable to Gemma (48 layers, 3840 dim). The two Gemma-4-12B models ARE identical, enabling a valid same-space comparison. Measured signature: abliteration drops the deep-shell (layer 36) activation norm (153.06β†’146.21) while surface/mid layers are unchanged.
100
+ - Goodfire AI's neural geometry research underpins the manifold approach (manifold steering, curved-manifold concepts, SAE limitations).
101
+
102
+ ### Hard Constraints
103
+
104
+ - Model ≀ 32B parameters
105
+ - Gradio app on HF Spaces under `build-small-hackathon` org
106
+ - Deadline: June 15, 2026
107
+ - Compute: $250 Modal credits + $20 HF credits + 40min/day ZeroGPU
108
+ - Submission: Space + demo video + social post
109
+ - Unlimited submissions allowed, max 10 ZeroGPU Spaces
110
+ - Can submit to both tracks
111
+
112
+ ### Models
113
+
114
+ - **google/gemma-4-12B-it** (base) β€” hidden-state source, 48 layers, 3840 dim, arch `gemma4_unified`, decoder hooks at layers 12/24/36, decoder path `model.language_model.layers`
115
+ - **OBLITERATUS/Gemma-4-12B-OBLITERATED** (abliterated/uncensored) β€” architecturally identical twin; shares hooks/dims/tokenizer β†’ valid same-space comparison
116
+ - Both 12B (≀32B OK, but **Tiny Titan ≀4B forfeited** β€” traded for the censored-vs-uncensored interpretability demo)
117
+ - No image model in the hot path (FLUX/LivePortrait removed in the brain pivot)
118
+
119
+ ### Prize Targeting Strategy
120
+
121
+ | Target | Prize | Path |
122
+ |---|---|---|
123
+ | πŸ„ Thousand Token Wood | $4,000 (1st) | Weird + delightful (live brain firing) |
124
+ | Modal | $10,000 credits (1st) | Serve both Gemmas on Modal |
125
+ | Community Choice | $2,000 | Viral/shareable |
126
+ | Best Demo | $1,000 | Real-time neuron firing + EEG |
127
+ | Judges' Wildcard | $1,000 | Censored-vs-uncensored interpretability art |
128
+ | Bonus Quest Champion | $2,000 | badges on main project |
129
+ | ~~BFL~~ | ~~$5,000~~ | forfeited (no FLUX in brain pivot) |
130
+ | ~~Tiny Titan~~ | ~~$1,500~~ | forfeited (12B > 4B cap) |
131
+
132
+ ### Badge Strategy (per project, "single sash")
133
+
134
+ Main project badge paths (brain pivot):
135
+ - βœ… 🎨 Off-Brand β€” custom 3D-brain + EEG visualization UI (Three.js)
136
+ - βœ… πŸ“‘ Sharing is Caring β€” publish per-model brain bundles + neuron clouds on Hub
137
+ - βœ… πŸ““ Field Notes β€” blog post on manifold research + the abliteration deep-shell finding + the build
138
+ - βœ… 🎯 Well-Tuned β€” fine-tuned `Ministral-8B-Instruct` LoRA interpreter published as `build-small-hackathon/activation-brain-interpreter`; it translates hidden-layer-derived telemetry into comparison analysis
139
+ - ❌ πŸ¦™ Llama Champion β€” dropped with Qwen (no 4B model in scope now)
140
+ - ❌ πŸ”Œ Off the Grid β€” conflicts with Modal ($10k prize); 12B is hard to run local anyway
141
+
142
+ ### Prior Art & Research
143
+
144
+ - **Milady project** (`/a0/usr/projects/milady`): Activation Avatars β€” 3 trained adapter checkpoints (CrossAttention, MultiToken), ~2M params each. Maps Qwen3-4B layers 9/18/27 hidden states β†’ FLUX.2-Klein prompt embeddings. 20 emotions. Results: milady-style face images per emotion.
145
+ - **Goodfire research** (May-June 2026): Neural geometry series. Key papers:
146
+ - "The World Inside Neural Networks" β€” concepts on curved manifolds
147
+ - "Steering Along Manifolds" β€” manifold steering >> linear steering on Llama 3.1 8B
148
+ - "Geometric Calculator" β€” general-purpose addition module in layer 18 computing over circular geometry
149
+ - "Can SAEs Capture Neural Geometry?" β€” SAEs shatter/dilute manifold structure
150
+ - **Hackathon winners research**: One-sentence pitch + instant wow, multi-agent clear roles, optimize experience not model, latency engineering wins, nail deliverables, community likes count
151
+
152
+ ### Architecture (deployed)
153
+ ### Public Technical Artifacts
154
+
155
+ - HF Dataset `build-small-hackathon/activation-brain-artifacts` publishes the reproducibility bundle: Gemma fingerprint JSONs, 627 affect-labeled probe prompts, manifold plots, `manifold_processed.npz`, summary report, fingerprinting script, analysis scripts, Modal backend reference code, and interpreter training/serving scripts.
156
+ - HF Model `build-small-hackathon/activation-brain-interpreter` publishes the fine-tuned Mistral-family LoRA interpreter adapter used for generated comparison analysis.
157
+ - The raw local `manifold_data.pt` hidden-state dump remains unpublished intentionally; use processed/public artifacts for submission evidence unless raw activations are explicitly requested.
158
+
159
+
160
+ - **Modal** `gemma-brain` app β€” two classes `BaseGemma` + `OblitGemma` (L40S), each loads its model + precomputed brain bundle, streams `fire`/`token`/`done` SSE. Hook path `model.language_model.layers`, hooks [12,24,36].
161
+ - **Interpreter** `activation-brain-interpreter` Modal app β€” serves the published `build-small-hackathon/activation-brain-interpreter` LoRA adapter on `mistralai/Ministral-8B-Instruct-2410`; `/api/analyze` proxies prompt + both responses + baseline-corrected deltas + native meters to generate varied plain-English analysis. Deterministic frontend analysis remains fallback.
162
+ - **Frontend** `brain_app.py` (Gradio + FastAPI) β€” per-model same-origin proxy routes `/api/{neurons,init,stream}/{model}` plus `/api/analyze`, serves `static/brain_engine.js` (dual EEG + comparison analysis UI). **Dual concurrent mode (no dropdown):** one Send opens TWO parallel SSE streams (base + oblit); each drives its own EEG strip (right column), both model responses are shown labeled in chat, and the left column shows a live plain-English comparison-analysis card beneath Stats instead of the older single 3D brain; the analysis emphasizes what the divergence means for tone, warmth, caution, uncertainty, and shared-manifold trajectory rather than only listing metric values. Each EEG includes baseline-corrected emotion activation deltas (positive excess over the first 8 fire events of that response, not argmax counts) plus a calibrated baseline-corrected model-native state meter derived from live weights: uniform 0–100 Valence, Activation, Uncertainty, Constraint, activity-gated Conflict, and Warmth. `mdLite()` renders safe markdown + collapses degenerate loops.
163
+ - **Fingerprints** built by `fingerprint_model.py` (model-agnostic; shared UMAP fit on base, transform on oblit) β†’ `gemma4_{base,oblit}_{neurons.json,brain_bundle.pt}` on Modal volume `avatars-cache`.
164
+ - Note: Gradio mounts `gr.HTML` asynchronously, so `brain_engine.js` must poll for `#ab-brain` before init (boot poller) β€” otherwise the canvas stays blank.
165
+
166
+ ### Key Documents
167
+
168
+ - `BRIEFING.md` β€” Full hackathon rules, prizes, timeline, strategy, badges, sponsor model catalog, winning playbook
169
+ - Tracker at `/a0/usr/workdir/office-manager/tracker.json`
170
+
171
+ ### Decision Log
172
+
173
+ Tracked in `BRIEFING.md` section 11.
174
+
175
+ ## Child DOX Index
176
+
177
+ No child directories yet. Will be created as project grows (e.g., `src/`, `adapters/`, `notebooks/`).