ML//3D generation

3D generation is the use of generative models to produce three-dimensional shapes, usually as a polygon mesh ready for a game engine, a simulator or a CAD tool, from a text prompt, an image or a rough point cloud, and it is what promises to cut the hours an artist or an engineer spends modelling assets: props for a training simulator, environments for a drone's synthetic data, a first draft of a bracket. Two families compete. One generates a field (a density or a signed distance at every point of space) and extracts the surface afterwards as an isosurface; the other writes the mesh directly, vertex by vertex and face by face.


3D generation is the use of generative models to produce three-dimensional shapes, usually as a polygon mesh ready for a game engine, a simulator or a CAD tool, from a text prompt, an image or a rough point cloud, and it is what promises to cut the hours an artist or an engineer spends modelling assets: props for a training simulator, environments for a drone's synthetic data, a first draft of a bracket. Two families compete. One generates a field (a density or a signed distance at every point of space) and extracts the surface afterwards as an isosurface; the other writes the mesh directly, vertex by vertex and face by face.

The direct family treats a mesh like a sentence. The faces are put in a fixed order and their vertex coordinates, rounded to a grid, become tokens, so an autoregressive transformer can generate a shape the way an LLM generates text, conditioned on a point cloud or a description. NVIDIA's Meshtron (December 2024) is the clearest example: the paper reports meshes of up to 64,000 faces at a 1,024-step coordinate grid, against roughly 1,600 faces for earlier artist-style generators, with about half the training memory and 2.5 times the throughput (the authors' own comparisons). The trick is sequence length: a 32,000-face mesh is about 300,000 tokens, so it uses an hourglass architecture that compresses the sequence in stages and a sliding window when generating.

A generated mesh is judged by its topology as much as by its look.

Artists build meshes with clean, regular faces that deform and texture well; surface extraction from a field gives dense, irregular triangles that look right and are painful to edit. Writing the mesh directly aims at the artist's kind.

The direct approach controls detail where it is needed (more faces on curves, few on flat walls) but inherits the failures of long autoregressive sequences: a wrong token late in the sequence leaves holes, intersecting edges or a face that points inward.

A shape for a simulator or a printer must be watertight and physically sensible; a pretty render does not prove either. Checking a generated asset (closed surface, no self-intersections, scale) is part of the pipeline.

For engineering parts, a generated mesh is at most a starting point: tolerances, materials and loads live in CAD and simulation, which a mesh does not carry.

Field-based methods connect to diffusion models and to reconstruction from lidar scans; mesh-writing methods connect to the transformer and tokenization machinery of language models.