RJuro Claude Opus 4.6 commited on
Commit
3a6e6ca
Β·
1 Parent(s): 2dac300

Add comprehensive deployment guide to README

Browse files

Step-by-step instructions for students: install deps, run precompute,
test locally, create HF Space, push with git LFS, and wait for build.
Includes architecture diagram, file table, and how-it-works section.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Files changed (1) hide show
  1. README.md +167 -0
README.md CHANGED
@@ -6,3 +6,170 @@ colorTo: purple
6
  sdk: docker
7
  app_port: 7860
8
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
  sdk: docker
7
  app_port: 7860
8
  ---
9
+
10
+ # SciFact Multilingual Semantic Search
11
+
12
+ A deployable semantic search engine over 5,183 scientific abstracts (SciFact dataset), using **ChromaDB** for vector storage and **multilingual-e5-small** for cross-lingual search in English, French, German, and Spanish.
13
+
14
+ **Live demo:** [huggingface.co/spaces/RJuro/scifact-semantic-search](https://huggingface.co/spaces/RJuro/scifact-semantic-search)
15
+
16
+ ## Architecture
17
+
18
+ ```
19
+ LOCAL (one-time setup) HF SPACES (runtime, CPU)
20
+ ────────────────────── ────────────────────────
21
+ SciFact dataset Load ChromaDB from data/chroma_db/
22
+ ↓ Load multilingual-e5-small
23
+ Encode 5,183 docs with ↓
24
+ multilingual-e5-small /search?q=... β†’
25
+ ↓ encode query β†’
26
+ Save to ChromaDB ──── push via git ────→ ChromaDB query β†’
27
+ (data/chroma_db/) JSON results
28
+ ```
29
+
30
+ **Key idea:** Corpus encoding happens once on your machine. Only query encoding runs on HF Spaces (CPU). This keeps the Space fast and cheap.
31
+
32
+ ## Files
33
+
34
+ | File | Purpose |
35
+ |------|---------|
36
+ | `precompute.py` | Encodes all 5,183 SciFact docs and saves them to ChromaDB (run locally) |
37
+ | `app.py` | FastAPI server β€” loads ChromaDB + model, serves search API and frontend |
38
+ | `static/index.html` | Frontend β€” vanilla HTML/CSS/JS, no dependencies |
39
+ | `requirements.txt` | Python dependencies |
40
+ | `Dockerfile` | Container config for HF Spaces |
41
+ | `README.md` | This file (YAML header is required by HF Spaces) |
42
+
43
+ ## Step-by-Step Deployment Guide
44
+
45
+ ### Prerequisites
46
+
47
+ - Python 3.9+
48
+ - A free [Hugging Face](https://huggingface.co) account
49
+ - Git with [Git LFS](https://git-lfs.com/) installed (`brew install git-lfs` on macOS)
50
+
51
+ ### Step 1 β€” Install dependencies
52
+
53
+ ```bash
54
+ pip install -r requirements.txt
55
+ ```
56
+
57
+ ### Step 2 β€” Run precompute.py (local, one-time)
58
+
59
+ This downloads the SciFact dataset, encodes all 5,183 abstracts with `intfloat/multilingual-e5-small`, and saves the vectors + metadata into a persistent ChromaDB at `data/chroma_db/`.
60
+
61
+ ```bash
62
+ python precompute.py
63
+ ```
64
+
65
+ Takes ~2 minutes on CPU. When done you should see:
66
+
67
+ ```
68
+ ChromaDB persisted to: .../data/chroma_db
69
+ Collection 'scifact': 5183 documents
70
+ ```
71
+
72
+ Verify the output:
73
+
74
+ ```bash
75
+ ls data/chroma_db/
76
+ # Should show: chroma.sqlite3 and a UUID-named directory
77
+ ```
78
+
79
+ ### Step 3 β€” Test locally
80
+
81
+ ```bash
82
+ uvicorn app:app --port 7860
83
+ ```
84
+
85
+ Open [localhost:7860](http://localhost:7860) in your browser. Try searching:
86
+ - `effects of vaccination` (English)
87
+ - `effets de la vaccination` (French)
88
+ - `Auswirkungen der Impfung` (German)
89
+
90
+ The same English-language corpus should return relevant results regardless of query language.
91
+
92
+ ### Step 4 β€” Create a Hugging Face Space
93
+
94
+ Go to [huggingface.co/new-space](https://huggingface.co/new-space):
95
+ - **Space name:** choose any name (e.g. `scifact-semantic-search`)
96
+ - **SDK:** Docker
97
+ - **Visibility:** Public
98
+
99
+ Or use the CLI:
100
+
101
+ ```bash
102
+ pip install huggingface-hub
103
+ huggingface-cli login # paste your HF token
104
+ python -c "
105
+ from huggingface_hub import HfApi
106
+ api = HfApi()
107
+ url = api.create_repo('YOUR-SPACE-NAME', repo_type='space', space_sdk='docker')
108
+ print(url)
109
+ "
110
+ ```
111
+
112
+ ### Step 5 β€” Push to HF Spaces
113
+
114
+ Initialize git, enable LFS (needed because `chroma.sqlite3` is ~74 MB), and push:
115
+
116
+ ```bash
117
+ git init
118
+ git lfs install
119
+ git lfs track "*.sqlite3" "*.bin"
120
+ git add .
121
+ git commit -m "Initial deploy"
122
+ git remote add origin https://huggingface.co/spaces/YOUR-USERNAME/YOUR-SPACE-NAME
123
+ git push origin main
124
+ ```
125
+
126
+ If the push is rejected (HF creates a default commit), pull first:
127
+
128
+ ```bash
129
+ git pull origin main --rebase
130
+ # Resolve any conflicts in .gitattributes / README.md (keep your versions)
131
+ git add .
132
+ git rebase --continue
133
+ git push origin main
134
+ ```
135
+
136
+ ### Step 6 β€” Wait for build
137
+
138
+ HF Spaces will build the Docker image (installs PyTorch, sentence-transformers, etc.). This takes 5-10 minutes on the first deploy. Watch progress in the Space's **Logs** tab.
139
+
140
+ Once the status shows **Running**, your app is live.
141
+
142
+ ## How It Works
143
+
144
+ ### Embedding model
145
+
146
+ **`intfloat/multilingual-e5-small`** (118M params, 384 dimensions)
147
+
148
+ This is a compact multilingual retrieval model. Critical detail β€” E5 models require prefixes:
149
+ - Documents: `passage: {text}`
150
+ - Queries: `query: {text}`
151
+
152
+ Without these prefixes, retrieval quality drops significantly.
153
+
154
+ ### Vector database
155
+
156
+ **ChromaDB** with persistent storage and cosine distance. Documents are stored with precomputed embeddings so ChromaDB doesn't need to re-embed anything at runtime.
157
+
158
+ ### Search flow
159
+
160
+ 1. User types a query in any supported language
161
+ 2. FastAPI encodes it with `query: {text}` prefix using the E5 model
162
+ 3. ChromaDB finds the 5 nearest neighbors by cosine distance
163
+ 4. Results are returned as JSON: `{rank, score, title, text}`
164
+ 5. Score = 1 - cosine_distance (displayed as similarity percentage)
165
+
166
+ ### Cross-lingual search
167
+
168
+ The multilingual E5 model maps text from different languages into the same vector space. A French query about vaccination lands near English documents about vaccination β€” no translation needed.
169
+
170
+ ## Customization Ideas
171
+
172
+ - **Different dataset:** Replace `load_scifact()` in `precompute.py` with your own corpus
173
+ - **More languages:** The model supports 100+ languages β€” add more example chips in `index.html`
174
+ - **More results:** Change `top_k` parameter (default 5, max 20)
175
+ - **Reranking:** Add a cross-encoder reranker on top of the retrieval results for better precision