-
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
Paper β’ 2504.01016 β’ Published β’ 29 -
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
Paper β’ 2503.05638 β’ Published β’ 20 -
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
Paper β’ 2409.07447 β’ Published β’ 1 -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37
Collections
Discover the best community collections!
Collections including paper arxiv:2409.02095
-
Controllable Text Generation for Large Language Models: A Survey
Paper β’ 2408.12599 β’ Published β’ 65 -
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
Paper β’ 2408.12590 β’ Published β’ 35 -
Real-Time Video Generation with Pyramid Attention Broadcast
Paper β’ 2408.12588 β’ Published β’ 17 -
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Paper β’ 2408.11039 β’ Published β’ 63
-
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
Paper β’ 2405.20222 β’ Published β’ 11 -
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
Paper β’ 2406.00908 β’ Published β’ 12 -
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Paper β’ 2406.02509 β’ Published β’ 10 -
I4VGen: Image as Stepping Stone for Text-to-Video Generation
Paper β’ 2406.02230 β’ Published β’ 18
-
LocalMamba: Visual State Space Model with Windowed Selective Scan
Paper β’ 2403.09338 β’ Published β’ 8 -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
Paper β’ 2403.09394 β’ Published β’ 26 -
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Paper β’ 2402.19479 β’ Published β’ 35 -
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper β’ 2405.10300 β’ Published β’ 31
-
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37 -
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Paper β’ 2409.01704 β’ Published β’ 83 -
CDM: A Reliable Metric for Fair and Accurate Formula Recognition Evaluation
Paper β’ 2409.03643 β’ Published β’ 19 -
UniDet3D: Multi-dataset Indoor 3D Object Detection
Paper β’ 2409.04234 β’ Published β’ 9
-
Kolors Virtual Try-On
π10kGenerate a virtual tryβon image of a person wearing a garment
-
Finegrain Object Cutter
β516Create HD cutouts from any image with just a prompt
-
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37 -
Diffusers Image Outpaint
π2.52kEasily expand image boundaries
-
DepthFM: Fast Monocular Depth Estimation with Flow Matching
Paper β’ 2403.13788 β’ Published β’ 18 -
Learning Temporally Consistent Video Depth from Video Diffusion Priors
Paper β’ 2406.01493 β’ Published β’ 23 -
NeuFlow v2: High-Efficiency Optical Flow Estimation on Edge Devices
Paper β’ 2408.10161 β’ Published β’ 15 -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper β’ 2402.04252 β’ Published β’ 30 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper β’ 2402.03749 β’ Published β’ 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper β’ 2402.04615 β’ Published β’ 44 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper β’ 2402.05008 β’ Published β’ 23
-
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
Paper β’ 2504.01016 β’ Published β’ 29 -
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
Paper β’ 2503.05638 β’ Published β’ 20 -
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
Paper β’ 2409.07447 β’ Published β’ 1 -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37
-
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37 -
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Paper β’ 2409.01704 β’ Published β’ 83 -
CDM: A Reliable Metric for Fair and Accurate Formula Recognition Evaluation
Paper β’ 2409.03643 β’ Published β’ 19 -
UniDet3D: Multi-dataset Indoor 3D Object Detection
Paper β’ 2409.04234 β’ Published β’ 9
-
Controllable Text Generation for Large Language Models: A Survey
Paper β’ 2408.12599 β’ Published β’ 65 -
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
Paper β’ 2408.12590 β’ Published β’ 35 -
Real-Time Video Generation with Pyramid Attention Broadcast
Paper β’ 2408.12588 β’ Published β’ 17 -
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Paper β’ 2408.11039 β’ Published β’ 63
-
Kolors Virtual Try-On
π10kGenerate a virtual tryβon image of a person wearing a garment
-
Finegrain Object Cutter
β516Create HD cutouts from any image with just a prompt
-
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37 -
Diffusers Image Outpaint
π2.52kEasily expand image boundaries
-
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
Paper β’ 2405.20222 β’ Published β’ 11 -
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
Paper β’ 2406.00908 β’ Published β’ 12 -
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Paper β’ 2406.02509 β’ Published β’ 10 -
I4VGen: Image as Stepping Stone for Text-to-Video Generation
Paper β’ 2406.02230 β’ Published β’ 18
-
DepthFM: Fast Monocular Depth Estimation with Flow Matching
Paper β’ 2403.13788 β’ Published β’ 18 -
Learning Temporally Consistent Video Depth from Video Diffusion Priors
Paper β’ 2406.01493 β’ Published β’ 23 -
NeuFlow v2: High-Efficiency Optical Flow Estimation on Edge Devices
Paper β’ 2408.10161 β’ Published β’ 15 -
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Paper β’ 2409.02095 β’ Published β’ 37
-
LocalMamba: Visual State Space Model with Windowed Selective Scan
Paper β’ 2403.09338 β’ Published β’ 8 -
GiT: Towards Generalist Vision Transformer through Universal Language Interface
Paper β’ 2403.09394 β’ Published β’ 26 -
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Paper β’ 2402.19479 β’ Published β’ 35 -
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Paper β’ 2405.10300 β’ Published β’ 31
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper β’ 2402.04252 β’ Published β’ 30 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper β’ 2402.03749 β’ Published β’ 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper β’ 2402.04615 β’ Published β’ 44 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper β’ 2402.05008 β’ Published β’ 23