view post Post 4715 Big update to llm-datasets, my curated list of datasets and tools for post-training LLMs.> Added many new datasets> New "thinking" column> Refreshed recommended tools.Thanks to everyone who told me they used it for their research at ICLR, you motivated this update! See translation 2 replies · 👍 4 4 👀 3 3 🤗 3 3 + Reply
Zero-Overhead Introspection for Adaptive Test-Time Compute Paper • 2512.01457 • Published Dec 1, 2025 • 3
view post Post 10510 New family of 1B models just dropped!> LiquidAI/LFM2.5-1.2B-Base: 10T → 28T tokens> LiquidAI/LFM2.5-1.2B-Instruct: new large-scale multi-stage RL> LiquidAI/LFM2.5-1.2B-JP: our most polite model> LiquidAI/LFM2.5-VL-1.6B: multi-image multilingual> LiquidAI/LFM2.5-Audio-1.5B: 8x times faster, no quality lossSuper proud of this release 🤗 See translation 3 replies · 🚀 18 18 👀 1 1 + Reply
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation Paper • 2406.14971 • Published Jun 21, 2024
Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit Paper • 2506.06607 • Published Jun 7, 2025 • 3
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation Paper • 2511.22989 • Published Nov 28, 2025 • 17
Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer Paper • 2510.05846 • Published Oct 7, 2025 • 3
view post Post 8501 LiquidAI/LFM2-8B-A1B just dropped!8.3B params with only 1.5B active/token 🚀> Quality ≈ 3–4B dense, yet faster than Qwen3-1.7B> MoE designed to run on phones/laptops (llama.cpp / vLLM)> Pre-trained on 12T tokens → strong math/code/IF See translation 1 reply · 🔥 9 9 🚀 3 3 + Reply
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Paper • 2510.01284 • Published Sep 30, 2025 • 37
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models Paper • 2509.23233 • Published Sep 27, 2025 • 4
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition Paper • 2509.19768 • Published Sep 24, 2025 • 7
view post Post 3939 ⚛️ New drop of tiny task-specific models!Want to do data extraction, translation, RAG, tool use, or math on a Raspberry Pi? We got you covered! ✅These tiny models were fine-tuned to perform narrow tasks extremely well, making them competitive with much larger models.You can deploy them today on-device or even on GPUs for big data operations! LiquidAI/liquid-nanos-68b98d898414dd94d4d5f99a See translation 1 reply · 🔥 5 5 👍 2 2 😎 1 1 + Reply
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments Paper • 2509.14233 • Published Sep 17, 2025 • 23
Granite Code Models: A Family of Open Foundation Models for Code Intelligence Paper • 2405.04324 • Published May 7, 2024 • 26