Image2Mesh: Transformer-based Single-Image 3D Reconstruction with Triplane Representation and Neural Rendering
Recent advancements in transformer-based deep learning have significantly improved the accuracy and efficiency of single-image 3D reconstruction. The transformer-based feed-forward architecture for single-image 3D reconstruction, system integrates DINOv1 vision transformers, triplane repre-sentations and Neural Radiance Fields (NeRF) principles to generate textured 3D meshes from single RGB images. Key contributions include efficient triplane decoding, robust pre-processing with background removal (rembg) and edge de-tection, and a Gradio-based web interface enabling real-time deployment on consumer GPUs (6GB VRAM). Preliminary deployment experiments suggest that Image2Mesh is capable of near-real-time inference on high-performance GPUs while demonstrating a favorable speed–accuracy trade-off compared to traditional photogrammetry-based workflows. Image2Mesh provides a practical foundation for AR/VR, gaming, and content creation applications.