Import a video and choose a background
Import MP4, WebM or MOV, then choose transparent, solid, local image or blurred-source background without stretching or cropping the frame.
Your video is not uploaded. This lab-grade Beta prioritises hair, fast motion, occlusion and temporal stability across scene cuts. Models download on demand; decoding, matting, compositing, encoding and audio muxing remain on-device.
These facts describe the current product, not unshipped roadmap work.
Reviewed by the DataDance product and engineering team · Published · Updated
Confirm important settings before export; the processing location is always disclosed.
Import MP4, WebM or MOV, then choose transparent, solid, local image or blurred-source background without stretching or cropping the frame.
Pick 7.2 MB RVM compatible, 51.3 MB RVM quality, or the MatAnyone2 non-commercial test tier calibrated on its first frame by RVM ResNet50.
WebCodecs encodes the frame-by-frame result and remuxes source audio. Opaque backgrounds default to MP4; transparency uses WebM.
ImgIng treats video matting as a quality-first lab-grade Beta. Evaluation covers more than a still outline: it inspects hair, crossing limbs, fast motion, low-contrast edges, subject-free openings and temporal stability after scene cuts. On the current representative internal samples, the MatAnyone2 ultimate tier produced the best overall matting result among the comparable browser-side solutions ImgIng tested. This is a scoped, provisional result—not an absolute ranking of every product, device or possible clip—and will be updated as the public benchmark set expands.
Independent frames can make hair and translucent edges flicker. ImgIng’s RVM path carries recurrent temporal state and resets it on detected scene cuts. The experimental tier first uses RVM ResNet50 to calibrate first-frame hair and the complete silhouette, then passes that seed to MatAnyone2 memory propagation to reduce initial omissions, jitter and previous-shot residue.
RVM MobileNetV3 FP16 is about 7.2 MB and suits quick previews, typical devices and longer video. RVM ResNet50 FP16 is about 51.3 MB and trades speed for steadier detail. MatAnyone2 automatically selects a landscape, square or portrait profile of about 294 MB and reuses the 51.3 MB RVM ResNet50 model for first-frame calibration, about 345 MB in total; it is for non-commercial testing only. Every model uses ModelScope CDN first, falls back to the same pinned version on ImgIng, and enters browser cache only after verification.
No. The workspace preserves the original frame ratio and export dimensions. It does not force the source into a square or crop the visible frame for inference. MatAnyone2 only selects the internal landscape, square or portrait input whose usable analysis area is largest. First-frame calibration proportionally caps ultra-high-resolution input at a 1920-pixel long edge to control memory, without changing final output dimensions. Background images independently support cover or contain.
The experimental tier no longer aborts a full export merely because the opening frame has no subject. Those frames output the selected background while the pipeline keeps looking; temporal memory starts when a person is detected. A scene cut clears the previous memory and runs calibration again.
Current desktop Chrome and Edge are the primary targets. The page needs HTTPS or localhost, WebCodecs and WebAssembly. RVM uses the stable WASM inference path; MatAnyone2 prefers WebGPU and can fall back to WASM. Older devices, phones and long 4K video may be impractical, so the workspace detects capabilities and reports download and processing ETA.
Beta limits: preview the current frame before processing the full video. Complex occlusion, fast motion, translucent objects and scene cuts can still produce edge errors. “Any aspect ratio” means no stretching or cropping; it does not promise identical matting accuracy for every extreme composition. Transparent video exports as WebM VP9 alpha; MP4 is for solid, image or blurred opaque backgrounds.
Updated 2026-08-25 · Live capability detection inside the tool is authoritative
These visible answers match the current product behaviour and structured data.
No. The complete video-matting pipeline stays in the browser; the network is used only to download the runtime and selected model on first use.
No. It means the MatAnyone2 ultimate tier produced the best overall result on the current representative internal samples among the comparable browser-side solutions evaluated. Difficult footage should still be previewed, and the benchmark scope will continue to expand.
It loads an approximately 294 MB temporal profile for the current aspect class and reuses the roughly 51.3 MB RVM ResNet50 model for first-frame calibration, about 345 MB total. It takes longer and is for non-commercial testing only.
Yes by default. Compatible audio is copied; otherwise it is encoded to Opus for WebM or AAC for MP4 and remuxed.
Each task page documents real settings, limits and format advice—not keyword-swapped duplicates.