N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
Abstract
We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train N_0-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a physically grounded contact signal that conditions action generation. To improve long-horizon and multi-stage manipulation, we introduce tactile contact events for task staging and advance through them during execution. For real-time efficiency, we adopt an asymmetric Mixture-of-Transformers architecture that pairs a full-width expert for video prediction with slim experts for downstream action and tactile prediction. Evaluations on both real and simulated benchmarks justify the capabilities of N_0-TWAM across a range of contact-rich tasks, and demonstrate the benefit of data scaling for precise tactile and action prediction. In summary, N_0-TWAM endows a world-action model with predictive capabilities to foresee vision, touch and action, building a solid foundation for fine-grained manipulation on open contact-rich tasks. The codebase and model checkpoints will be made publicly available to foster further research and development in tactile-enabled robotic manipulation.
Community
Hi everyone โ author here! ๐
Excited to share Nโ-TWAM, a tactile-native world-action model for contact-rich manipulation. Code and pretrained checkpoints are already out:
Code: https://github.com/neoteai/N0-TWAM
Project page: https://research.neoteai.com/n0-twam/
Instead of treating touch as an auxiliary input channel, we make the model predict future contact together with future vision, under the same objective and at the same causal step, and read actions off that jointly predicted future.
Tactile-native world model! To our knowledge the first tactile world-action model trained at scale โ visuo-tactile joint pre-training on tactile-rich demos spanning 6 embodiments and 450 tasks.
Happy to answer questions here or at github, and feedback or criticism is very welcome, especially from people working on tactile sensing, VLA, or world models. ๐
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- $N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens (2026)
- Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention (2026)
- VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation (2026)
- TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction (2026)
- T-Rex: Tactile-Reactive Dexterous Manipulation (2026)
- TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation (2026)
- TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.23783 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper