Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
3D GenerationActive

tencent / HY-World-2.0-World-Reconstruction

HY-World-2.0-World-Reconstruction

bytencent

Reconstruct a 3D scene from multi-view photos or a short video. Produces 3D Gaussian splats with depth maps and cameras, plus a colored point cloud.

3D Generation
Model ID
HY-World-2.0-World-Reconstruction
Provider
tencent
Updated
1776858372
1
Comments
Average rating : 5 (1 users)
Providertencent
ModelHY-World-2.0-World-Reconstruction
3D Generation
wiro playground—tencent/HY-World-2.0-World-Reconstruction
Reset to defaults
Delete All
0 / 100
Maximum 100 image allowed
Drop image to upload

OR

Click to browse your device

Supports: JPG, JPEG, PNG, GIF, WEBP, HEIC

Image files for multi-view 3D reconstruction. Supported formats: .jpg, .jpeg, .png, .webp. Ignored if video is also provided.

Maximum inference resolution along the longest image edge.

Sample outputs
Updated 1776858372
## Overview HY-World 2.0 World Reconstruction is Tencent’s WorldMirror 2.0 scene reconstruction pipeline. It turns multi-view images or a casual video into a 3D scene representation. It runs a single feed-forward pass that predicts depth, surface normals, and camera parameters for each view. It then writes a 3D Gaussian Splatting scene and a colored point cloud you can render and inspect. This is useful when you want a fast “capture to 3D” path for digital twins, previews, or simulation scenes. ## What you can build - Digital-twin captures of rooms, streets, or small outdoor scenes from a phone video - 3D Gaussian Splat assets for real-time viewers and novel-view preview - Camera pose and intrinsics estimation for multi-view pipelines - Depth and normal map generation for VFX, relighting prep, and geometry cleanup - Prototype scene geometry for robotics simulation and embodied AI testing ## Inputs - A set of multi-view images of the same scene, captured from different viewpoints (JPG, JPEG, PNG, or WEBP). Use consistent exposure and avoid big moving objects. - Or a single video clip of the scene (MP4, AVI, MOV, WEBM, or GIF). If you provide a video, the run uses the video and ignores uploaded images. - A maximum processing resolution for inference, measured on the longest image edge (224 to 1920 pixels). The pipeline resizes and center-crops to a multiple of 14. - Video frame extraction controls: - A motion-aware strategy that skips near-static frames. - A uniform sampling strategy where you choose frames-per-second. - Minimum and maximum frame counts to extract from the clip. - Optional output cleanup filters: - Sky removal for outdoor scenes. - Edge and discontinuity filtering to reduce broken geometry around object boundaries. - Low-confidence filtering to drop the noisiest points. - A sky masking method choice that can fuse model-based masks with an ONNX segmenter. ## Outputs - A packaged results folder with multiple files and subfolders. - A 3D Gaussian Splatting file in standard PLY format. It stores per-Gaussian attributes like position, scale, rotation, opacity, and color coefficients. - A colored point cloud in PLY format. It is derived from the predicted per-view depth. - A camera parameters JSON file with camera intrinsics and camera-to-world extrinsics for each view. - Per-view depth outputs: - PNG depth visualizations. - NPY files with raw float32 depth arrays. - Per-view surface normal outputs as RGB normal-map PNGs. - Optional extras when enabled: - A COLMAP-style sparse reconstruction folder. - A rendered fly-through MP4 (RGB, and optionally a depth visualization video). ## Recommended settings - Keep the longest-edge inference resolution near 952 pixels for a good quality and memory balance. - For handheld videos, use the motion-aware frame selection first. - Start with 1 to 32 extracted frames for video reconstruction. Raise the cap only when the scene needs more coverage. - Turn sky removal on for outdoor captures with large sky regions. Turn it off for indoor scenes. - Keep edge filtering on for cleaner geometry around boundaries. - Enable low-confidence filtering for noisy inputs like low light, heavy compression, or fast motion. ## Limitations - This workflow expects multi-view coverage. One photo with no viewpoint change won’t reconstruct a stable scene. - Dynamic objects can create ghosting and floating artifacts. People, cars, and foliage in strong wind are common issues. - Reflective, transparent, or textureless surfaces often produce wrong depth and scattered points. - The pipeline resizes and center-crops to meet model constraints. Important content near the image edges can get cropped. - The main outputs are 3D Gaussian Splats and point clouds, not a watertight, textured triangle mesh. - Low-quality inputs can break reconstruction. Blurry frames, strong motion blur, rolling shutter, and heavily compressed video are high risk. ## Safety & compliance - Tencent’s HY-World 2.0 Community License Agreement limits use to a defined territory and explicitly excludes the European Union, the United Kingdom, and South Korea. - The license includes an Acceptable Use Policy. It bans harmful use cases, attempts to bypass safety safeguards, and misuse involving minors. - The Acceptable Use Policy also restricts election-related harmful misinformation, impersonation without consent, malware generation, and using personal data to harm others. - When you publish generated content, you must clearly disclose it is machine-generated, as required by the Acceptable Use Policy.

API quick start

Run HY-World-2.0-World-Reconstruction with a single API call.

POST https://api.wiro.ai/v1/Run/tencent/HY-World-2.0-World-Reconstruction
{
  "inputImage": "https://your-cdn.com/input.png",
  "targetSize": 952,
  "inputVideo": "https://your-cdn.com/input.mp4",
  "fps": 1
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion