This Chinese company took center stage at the prestigious international conference SIGGRAPH.

He will be featured at the SIGGRAPH main venue keynote, joining Disney, NVIDIA, and Bolt Graphics in the lineup of four keynote speakers at this year's conference.

This Chinese company took center stage at the prestigious international conference SIGGRAPH.

From July 19th to 23rd, SIGGRAPH 2026 was held in Los Angeles, USA. This annual event, which brings together leading scholars, artists, and engineers in the fields of computer graphics and visual computing from around the world, attracted nearly 10,000 attendees from 69 countries. Rendering, geometric modeling, animation simulation, and generative AI—almost all the cutting-edge explorations on how to use computers to create the visual world were presented in these five days.


This year, under the spotlight of this grand event,VAST, a leading global general artificial intelligence company(Developed the Tripo series of large 3D models) and became one of the most frequently mentioned Chinese teams at the conference:


Keynote address at SIGGRAPH main venueHe joined Disney, NVIDIA, and Bolt Graphics in the lineup of four keynote speakers at this year's conference.Five papers selected for Technical PapersIt covers mesh generation, 3D Gaussian generation, texture generation, interactive 3D asset generation, and cross-topology animation.


0a4074f03176e1a2bdcaf67483900067.png
0a4074f03176e1a2bdcaf67483900067.png

VAST Chief Scientist Dr. Cao Yanpei's Keynote Speech


The team also won the top prize in the conference's signature Real-Time Live! segment.Best in Show.


046527c88b3872ce445da55a11fcef7f.png
046527c88b3872ce445da55a11fcef7f.png


Outside the venue,VAST and World Labs, founded by Fei-Fei Li, jointly hosted a hackathon.And the World Model Industry Night. Concurrent with the conference,VAST Chief Scientist Cao Yanpei Selected for MIT Technology Review TR35 China List.


Several events combined to make this company, which was only three years old, attract a lot of attention.


Teapot test on the main stage


If there's one segment of SIGGRAPH that best represents the most noteworthy voices in the field of computer graphics this year, it's the Keynote address. This year, there were only four Keynote presentations, from Disney, NVIDIA, Bolt Graphics, and VAST.


On the afternoon of July 22,Dr. Cao Yanpei led the algorithm team to the Main StageHe delivered a speech entitled "The Teapot Test: First It Tested Showing. Now It Tests Making".


686a8cb3c7925773c099add365c9bf41.png
686a8cb3c7925773c099add365c9bf41.png


There's an old story behind this question in computer graphics: In 1975, Martin Newell, a pioneer in computer graphics, created a teapot model by hand measurements. This was one of the earliest 3D models in human history and later became the standard test object for evaluating rendering capabilities in the entire field of graphics. For half a century, the teapot test has consistently examined whether a computer can realistically represent a three-dimensional object.


However, Cao Yanpei believes that in the era of generative AI, this question that has lasted for half a century needs to be redefined.


Today, generative models can skip rulers, graph paper, and manual input. Input a sentence, and the model can generate a teapot with a smooth surface, realistic lighting, and the ability to be viewed from different angles—a result that looks complete enough.


Cao Yanpei then made three requests of the teapot: to put it into a game, to manufacture it using a 3D printer, or to use it to train robots to pour tea. None of these requests could actually be fulfilled. The reason is that there wasn't actually a teapot in the image.


9f86a1144c9488b6f9399d5c531527a1.png
9f86a1144c9488b6f9399d5c531527a1.png

VAST Chief Scientist Dr. Cao Yanpei's Keynote speech. Speech link: https://www.youtube.com/watch?v=zZ8fbEeaG-8


The model can generate a large number of beautiful images, but it may not have a reliable mesh, editable topology, realistic geometric details, movable joints, and physical properties that can be read by the simulator.


As Cao Yanpei said: You can look at it from all sides, but you can't really touch it.


Fifty years ago, the teapot test assessed whether a machine could display an object. Today, the new teapot test asks: Can a machine truly create an object?


The creation mentioned here means that the generated results can enter the game engine, be 3D printed, and become objects for robot training and physical simulation.


This is actually the second time VAST has been included in the SIGGRAPH Keynote speaker lineup.


Three years ago, VAST founder and CEO Song Yachen shared the stage with world-class entrepreneurs such as Jensen Huang, becoming the first Chinese entrepreneur to deliver a keynote speech at SIGGRAPH. At that time, generative 3D was still in its early stages, and the industry's most direct vision was to significantly lower the barrier to entry for modeling.


Three years later, the same company took to the main stage again, and the discussion had expanded from how to generate a model to how to build a full-stack 3D production system and how to drive a continuously operating world.


In the Real-Time Live! segment, VAST won the top prize.


In SIGGRAPH's Real-Time Live! segment, teams need to demonstrate their technical effects in real time in front of a real audience. Any lag or mishap will be exposed, making it one of the most anticipated segments of the entire conference.


f4afbce1e10823e426f822c0cf9f8eca.png
f4afbce1e10823e426f822c0cf9f8eca.png


This year, a total of 8 projects were selected for this segment, ranging from a major presentation such as Disney bringing the animated character Olaf to life as a robot, to Runway showcasing a single image generating a real-time AI character.


The VAST team's project, "Create Interactive 3D Assets in Seconds!", demonstrated how generative AI can create fully textured, skeletally bound, and directly animable 3D assets in seconds. Audience prompts submitted on the spot instantly generated corresponding assets and were displayed in a shared virtual world in real time.


d012bc759d8e3125a5330cc5cec36ac5.jpg
d012bc759d8e3125a5330cc5cec36ac5.jpg

The prompt submitted by the user at the event generated a treasure chest with "SIGGRAPH 2026" written on the front.


The capabilities truly demonstrated in this demonstration are far more complex than simply generating a model from a given sentence.


The system needs to understand user input, generate a 3D structure, complete texture mapping, predict reasonable skeletons and skinning, enable the character to move, and finally put multiple results into a scene that can run in real time.


cd6e0e26319148d72fc81b51e3343e38.jpg
cd6e0e26319148d72fc81b51e3343e38.jpg


If any step goes wrong, the final presentation may fail.


final,This project won the top prize in the Best in Show—Real-Time Live! category, as selected by the judging panel.


73d6bd2103546a683875448cbf914793.jpg
73d6bd2103546a683875448cbf914793.jpg

Audience members participate in the creation process, and the generated 3D props will randomly appear on the track. Cars that hit the props will score points, and the car with the higher score within 30 seconds will win.


Winning the top prize in such a low-tolerance on-site setting proves that VAST's capabilities, as described in its paper, are indeed capable of running without editing in front of thousands of people.


Sharing the stage with World Labs, a public dialogue with the world's top world model teams.


Another notable trend at SIGGRAPH 2026 is that the boundary between large 3D models and world models is rapidly disappearing.


During the conference, NVIDIA showcased its advancements in neural rendering, simulation, and the Cosmos world model, and positioned 3D graphics as a crucial foundation for physical AI to understand, generate, and simulate the world.


World Labs also positions itself as a spatial intelligence company, hoping to build models that can perceive, generate, reason, and interact with the three-dimensional world.


VAST's position in this global technology discussion has also become clearer.


From July 18th to 19th, prior to the opening of the conference,VAST and World Labs, founded by Fei-Fei Li, jointly hosted a hackathon called Worlds in Action.The event attracted over 500 creators to build prototypes of next-generation AI-native 3D interactive experiences within two days. The judging panel also included industry experts from organizations such as Lionsgate, Sony, Meta, and the NBA. The prize pool, combined with the opportunity to showcase the products at SIGGRAPH, made it a significant draw.


2a1020d506a5b0d6360f9080b17aa8bf.png
2a1020d506a5b0d6360f9080b17aa8bf.png


On July 20, the two parties jointly launched an industry networking night called World Models & GenAI Mixer, which brought together about 200 practitioners from AI platforms, basic model companies, visual effects studios, virtual production teams and creative agencies. VAST, World Labs and Groupcore Technology also held a dialogue on AI-driven next-generation 3D content production.


For a Chinese company, being able to co-host an event at a forum like SIGGRAPH with World Labs, a company founded by top academics, is itself a sign of peer recognition. This speaks volumes about its industry standing, more so than a single product launch or ranking achievement.


This has also changed VAST's role: it is no longer just showcasing a 3D generation technology developed by a Chinese team, but has begun to act as an important participant in the global field of AI 3D basic models, discussing the next stage of technological boundaries and industry direction with leading companies.


SIGGRAPH Keynote and TR35

Cao Yanpei received double recognition


Almost simultaneously with SIGGRAPH, Cao Yanpei received another recognition: on July 25th,MIT Technology Review released its 2025 list of 35 Innovators Under 35 (TR35) in China, and Cao Yanpei was among them..


e0ec81bc9f303664894ce8d14eaf7dfa.jpg
e0ec81bc9f303664894ce8d14eaf7dfa.jpg


Founded in 1999, TR35 aims to identify young scientific and technological talents under the age of 35 worldwide who possess outstanding innovation capabilities and development potential. Notable figures in their respective fields, including Linus Torvalds, Lisa Su, Song Han, Zhenan Bao, and Xiaowei Zhuang, have been selected.


MIT Technology Review's evaluation of Cao Yanpei focuses on his continuous technological accumulation in the fields of 3D vision and generative AI: from 3D reconstruction and neural rendering to AI 3D generation, he has promoted related technologies from generating static appearances to further moving towards usable structures, movable assets, and interactive scenes, providing underlying 3D generation capabilities for spatial intelligence and physical AI.


This selection coincides almost simultaneously with Cao Yanpei's keynote address at the SIGGRAPH 2026 main venue. The simultaneous appearance of individual honors and team achievements also reflects VAST's technological accumulation in AI 3D basic models, animable asset generation, and world modeling.


Five papers: breaking down every step of 3D asset creation.


What truly reflects a company's technological prowess is often the quantity and quality of its publications. SIGGRAPH's Technical Papers are renowned for their extremely high peer review standards, with an acceptance rate consistently around 25%, making it one of the most prestigious tracks among the top computer science conferences.


It's already quite an achievement to get one of them accepted.VAST secured five papers this year.It covers almost all the core aspects of 3D assets, from geometric representation, mesh generation, and skeleton binding to cross-topology animation migration and texture generation.


What's even more noteworthy about this set of results is that they are interconnected, almost covering the entire process of a 3D asset from its generation to its actual animation.


《Nexus: Native Mesh Generation with Diffusion》This paper is the core support behind VAST's flagship model, Tripo P1.0.


It abandons the previous mainstream serialization generation approach and decouples the generation of vertices and topology: vertices are regarded as sparse voxels in an octree and are generated globally from coarse to fine using a hierarchical diffusion model; at the same time, it proposes the concept of spatiotemporal intervals and encodes the topology of arbitrary edges and non-manifold surfaces into continuous per-vertex embeddings.


This paper is not just a laboratory achievement. The Tripo P1.0, released in March of this year, is the productization of Nexus's research results. It is the industry's first AI 3D generation model that can directly output a clean topology within 2 seconds and can be directly imported into a game engine.


876fc02b20497825f2378a4878b4f691.jpg
876fc02b20497825f2378a4878b4f691.jpg

Tripo P1.0 generation results


Another paper that also went beyond the academic paper - the product closed loop is"Generative 3D Gaussians with Learned Density Control", abbreviated as DeG.


It proposes the Density-Sampled Gaussians method, which remodels the density control of 3D Gaussians as an end-to-end learnable probabilistic sampling process and introduces the gradient contribution of rendering loss, transforming the previous discrete, heuristic rules for adding and deleting points into a differentiable density optimization process.


The results of this paper have been transformed into the open-source project TripoSplat. After its launch, it quickly topped the HuggingFace Space Trending list and received official workflow support from ComfyUI, sparking considerable discussion in the 3D generation community.


17df074367832fc7181e28aa1c9a6d93.jpg
17df074367832fc7181e28aa1c9a6d93.jpg

TripoSplat can convert a single 2D image into a high-quality, adjustable-number 3D Gaussian representation.


The remaining three papers each address a different technological gap:


  • "PixTex: Consistent 3D Texturing via Pixel-Space Multi-View Diffusion"By using a multi-view diffusion mechanism in pixel space, the old problems of misaligned seams and blurry artifacts in 3D native textures under multi-view projection are solved, allowing the materials of generated assets to directly achieve the clarity that can be imported into game and film engines. Currently, this technology has been seamlessly integrated into the VAST full-stack generation ecosystem.

  • 《AniGen: Unified S³ Fields for Animatable 3D Asset Generation》A unified representation method called S³ Fields is proposed, which puts the three things that are usually processed sequentially into the same shared space for joint generation. It has a significant lead over the current strong baseline methods in terms of topological correctness, skin distribution and other indicators, and can generalize to a wide range of categories such as animals, people, cartoon characters and even robotic arms.

  • "TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation"They then targeted the video-to-animation link and built the first topology-insensitive motion extraction framework, which can directly and with zero samples transfer motion from monocular videos to 3D characters with any unknown skeletal structure. The team also released a large-scale dataset, MOBJAVERSE, containing over 5,000 skeletal topologies and 2 million frames of animation.


Putting these five papers together reveals a clear logical thread: VAST has never aimed to solve the problem of simply generating a beautiful 3D model, but rather the end-to-end usability of 3D assets, from geometry and texture to skeleton and animation. This also echoes the question raised by Cao Yanpei in his keynote speech—can the generated content actually be used?


Conclusion


VAST is a leading global general artificial intelligence company dedicated to building general AI 3D foundational models and world models.


In just three years, the company has built a relatively complete technical system around AI 3D, with more than 20 million users worldwide.


Flagship ModelTripo P1.0It can generate topologically clean, low-polygon mesh assets that can be integrated into game engines and real-time workflows directly within 2 seconds;Tripo H3.1To achieve high-precision asset generation, we further enhance geometric density, surface detail, and structural accuracy.Project EdenThis extends the scope of exploration to the world model, attempting to break away from the generation paradigm centered on continuous pixels or video frames, decoupling the underlying world state from visual rendering, so that the generated world can run continuously, record interactions, and remain consistent under different observation perspectives.


5bf37a4c3ec24c2d898e2ca943c12161.jpg
5bf37a4c3ec24c2d898e2ca943c12161.jpg

Tripo H3.1: High-Fidelity 3D Generation


Looking back at VAST's numerous appearances at SIGGRAPH 2026, it's clear that they have established a fairly clear technological line.


The Nexus paper underpinned the flagship model Tripo P1.0, the DeG paper spurred the open-source project TripoSplat, and PixTex, AniGen, and TopoCap respectively filled the gaps in texture, skeletal rigging, and animation transfer—areas that were previously relatively weak in the field of 3D generation. Academic research, productization, and the open-source community have formed a mutually reinforcing closed loop within this company.


Fifty years ago, the Newell teapot tested whether computers were capable of displaying the three-dimensional world.


Today, a new test for teapots has been laid out before all AI 3D companies: once generated, can it truly enter the world?


From Song Yachen's first appearance at the SIGGRAPH main venue three years ago to Cao Yanpei's return this year with the new question of the "teapot test," coupled with the flourishing of papers, awards, open-source projects, and industry collaborations, this can perhaps be seen as a signal: in the highly technical field of computer graphics and 3D generation, Chinese companies have begun to join international leading institutions such as NVIDIA in the core agenda of the conference, and are engaging in direct dialogue with cutting-edge teams such as World Labs, participating in raising questions that need to be answered in the next stage.