luis-otte commited on
Commit
589d358
·
verified ·
1 Parent(s): 8079c56

Replace TBD with real metrics, fix broken links, add credits

Browse files
Files changed (1) hide show
  1. README.md +56 -40
README.md CHANGED
@@ -22,72 +22,88 @@ tags:
22
  - education
23
  - assistance
24
  - ottema
25
- datasets:
26
- - ottema/GLiNER-PTBR-Core
27
  base_model: fastino/gliner2-multi-v1
28
  ---
29
 
30
- # GLiNER2-PTBR
31
 
32
- GLiNER2-PTBR is a Brazilian Portuguese GLiNER model for **open-vocabulary Named Entity Recognition (NER)**. It extracts arbitrary entity types from Portuguese text using label prompts, including people, organizations, locations, documents, products, technologies, services, and operational evidence.
33
 
34
- GLiNER2-PTBR foi treinado para reconhecimento aberto de entidades em português brasileiro. O modelo suporta labels em português e inglês e foi otimizado para textos curtos, informais e ruidosos como mensagens de atendimento, suporte técnico e educação.
35
 
36
- ## Detalhes do modelo
37
 
38
  - **Base:** `fastino/gliner2-multi-v1` (Apache-2.0)
39
- - **Tipo:** GLiNER2 (bidirectional encoder + span scoring)
40
- - **Idioma:** pt-BR (com fallback para en)
41
- - **Tamanho:** ~307M parâmetros
42
- - **Licença:** Apache-2.0
43
 
44
- ## Uso
 
 
 
 
 
 
45
 
46
  ```python
47
- from gliner import GLiNER
48
 
49
- model = GLiNER.from_pretrained("ottema/gliner2-ptbr")
50
 
51
  text = "A professora Ana comprou um notebook Dell em Campinas no dia 12/06."
52
  labels = ["pessoa", "profissão", "produto", "marca", "local", "data"]
53
 
54
- entities = model.predict_entities(text, labels, threshold=0.5)
55
- for ent in entities:
56
- print(ent["text"], "->", ent["label"])
 
57
  ```
58
 
59
- ## Comparação com baseline
60
 
61
- | Modelo | span_F1 | label_F1 | entity_F1 |
 
 
62
  |---|---|---|---|
63
- | `fastino/gliner2-multi-v1` (zero-shot pt) | TBD | TBD | TBD |
64
- | `ottema/gliner2-ptbr` (este) | TBD | TBD | TBD |
65
 
66
- Métricas em `data/gliner_ptbr_core/test.jsonl`.
67
 
68
- ## Inferência CPU
 
 
 
69
 
70
- O modelo roda em CPU com latência mediana de ~50ms por texto curto (≤128 tokens).
71
 
72
- ## Limitações
73
 
74
- - Treinado em dados sintéticos: cobertura pode ser limitada em domínios muito específicos.
75
- - Performance pode cair em textos muito longos (>512 tokens).
76
- - Não é classificador final de cobertura/cobertura contratual.
77
 
78
- ## Citação
79
 
80
- ```
81
- @model{ottema_gliner2_ptbr_2024,
82
- title={GLiNER2-PTBR: Open-vocabulary NER for Brazilian Portuguese},
83
- author={Ottema},
84
- year={2024},
85
- url={https://huggingface.co/ottema/gliner2-ptbr}
86
- }
87
- ```
 
 
 
 
 
 
 
88
 
89
- ## Veja também
90
 
91
- - `ottema/GLiNER-PTBR-Core` — dataset de treino
92
- - `ottema/OntoEvidence-BR` — dataset de evidências operacionais
93
- - `ottema/gliner2-ptbr-ontoevidence-demo` — demo Gradio
 
22
  - education
23
  - assistance
24
  - ottema
 
 
25
  base_model: fastino/gliner2-multi-v1
26
  ---
27
 
28
+ # ottema/gliner2-ptbr (v0.4 — generalist)
29
 
30
+ **Open-vocabulary NER for Brazilian Portuguese, fine-tuned for informal and operational text (chat, atendimento, suporte).**
31
 
32
+ This is the **generalist release**. For HAREM-specialized (best entity F1 among compared models), see [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem). For ontology-guided evidence extraction, see [`ottema/gliner2-ptbr-ontoevidence`](https://huggingface.co/ottema/gliner2-ptbr-ontoevidence).
33
 
34
+ ## Model details
35
 
36
  - **Base:** `fastino/gliner2-multi-v1` (Apache-2.0)
37
+ - **Type:** GLiNER2 (bidirectional encoder + span scoring)
38
+ - **Language:** pt-BR (with English fallback)
39
+ - **Size:** ~307M parameters
40
+ - **License:** Apache-2.0
41
 
42
+ ## Intended use
43
+
44
+ General-purpose open-vocabulary NER for **informal and operational** Brazilian Portuguese text: atendimento, chat, suporte técnico, educação. Trained on synthetic data covering pessoas, profissões, locais, organizações, documentos, produtos, marcas, tecnologias, telefones, e-mails, datas, e valores monetários.
45
+
46
+ If you need a model benchmarked on journalistic Portuguese (HAREM), use [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem) instead.
47
+
48
+ ## Usage
49
 
50
  ```python
51
+ from gliner2 import GLiNER2
52
 
53
+ model = GLiNER2.from_pretrained("ottema/gliner2-ptbr")
54
 
55
  text = "A professora Ana comprou um notebook Dell em Campinas no dia 12/06."
56
  labels = ["pessoa", "profissão", "produto", "marca", "local", "data"]
57
 
58
+ entities = model.extract_entities(text, labels, threshold=0.5)
59
+ for label, spans in entities["entities"].items():
60
+ for span in spans:
61
+ print(f"{span} -> {label}")
62
  ```
63
 
64
+ ## Performance
65
 
66
+ Evaluation on `data/gliner_ptbr_core/test.jsonl` (synthetic generalist benchmark, threshold 0.3):
67
+
68
+ | Model | entity_F1 | span_F1 | label_F1 |
69
  |---|---|---|---|
70
+ | `fastino/gliner2-multi-v1` (zero-shot) | 0.9333 | 0.9347 | 0.9855 |
71
+ | **`ottema/gliner2-ptbr` (v0.4)** | **0.9976** | **0.9976** | **1.0000** |
72
 
73
+ On HAREM (163 samples, 2511 entities, journalistic PT-BR — out-of-distribution for this generalist):
74
 
75
+ | Model | entity_F1 | Δ vs baseline |
76
+ |---|---|---|
77
+ | `fastino/gliner2-multi-v1` (zero-shot) | 0.4271 | (reference) |
78
+ | `ottema/gliner2-ptbr` (v0.4) | 0.4132 | -1.39 pp |
79
 
80
+ The generalist is best for the synthetic informal-PT-BR distribution it was trained on. For journalistic text, see the HAREM-specialized model.
81
 
82
+ ## Inference
83
 
84
+ - **GPU:** ~30 ms per text (median, 32 batch)
85
+ - **CPU:** ~50 ms per short text (≤128 tokens)
 
86
 
87
+ ## Limitations
88
 
89
+ - Trained primarily on synthetic data; coverage may be limited in highly specialized domains.
90
+ - Performance may degrade on very long texts (>512 tokens).
91
+ - Not a substitute for domain-specific classifiers in regulated workflows.
92
+ - Out-of-distribution on journalistic text (use HAREM-specialized instead).
93
+
94
+ ## Credits
95
+
96
+ - **Base architecture:** GLiNER2 (Urchade Zaratiana et al.)
97
+ - **Base weights:** `fastino/gliner2-multi-v1` (Fastino)
98
+ - **Encoder:** microsoft/mdeberta-v3-base
99
+ - **Fine-tuning + datasets:** Ottema
100
+
101
+ ## License
102
+
103
+ Apache-2.0
104
 
105
+ ## See also
106
 
107
+ - [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem) — HAREM-specialized (best entity F1)
108
+ - [`ottema/gliner2-ptbr-ontoevidence`](https://huggingface.co/ottema/gliner2-ptbr-ontoevidence) — ontology-guided evidence extraction (in development)
109
+ - [`ottema/gliner2-ptbr-ontoevidence-data`](https://huggingface.co/datasets/ottema/gliner2-ptbr-ontoevidence-data) — OntoEvidence-BR dataset