ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Paper • 2607.28625 • Published • 34
The Kling Team is building next-generation multimodal world models across video, audio, text, 3D, and beyond. We are continuously looking for exceptional talent to join us. Feel free to reach out!
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Beacon: Knowing When and How to Perform Agentic Visual Reasoning