Every AI generator here produces the final image. What separates them is what the generator is conditioned on: a prompt and a picture, a flat derivative of your asset, or the 3D scene itself. If your asset is already a 3D model or CAD file, you are choosing between the geometry-first class (Intangible), an in-tool renderer like Enscape or D5 Render, and a CAD-native renderer like KeyShot; if all you have is photographs, you either condition on them or generate a mesh from one first.
Image-first generators have no copy of your object. It is described in the prompt and hinted at by reference photographs, so logo placement, label typography, proportions, and structural details drift between generations.
Conditioning-based generators sit in the middle, a genuine improvement over raw diffusion, each working from a flat picture of your asset. ControlNet pipelines are the common case: the Hugging Face diffusers documentation describes a ControlNet as conditioned on extra visual information or structural controls such as canny edge and depth maps, at a conditioning scale you choose. Veras conditions on a different picture: the view on your screen. Chaos calls it AI visualization for architects and designers, and its Enscape documentation says Veras uses your current live view in the Enscape viewport as base for generating new visuals. Chaos scopes its geometry claim for Veras to schematic design on an architectural model, where massing is the building's bulk: with Nano Banana's precise prompt control, only specified geometry and materials change, keeping your massing intact.
Real-time and plugin renderers read the model itself, live, through a link to the design tool. Chaos says Enscape lets you design, visualize, and iterate instantly and easily inside your CAD and BIM tools, and D5 Render's LiveSync documentation describes incremental updates of model and material from SketchUp to D5 Render. The live view is a render of that synced model, with generation around it: D5 lists texture map generation upstream of the render and AI Enhancer and Style Transfer passes over it.
Snaptrude, an architecture and BIM product, publishes a figure for its own AI Render rather than for conditioning pipelines generally: renders adhere to 95%+ of your geometry. Maquete's guide bounds what that shows: a visual match does not prove millimetre-level dimensional accuracy.
Geometry-first rendering also generates pixels, from what Intangible's documentation calls the actual 3D scene with the actual camera. Three layers stack: the 3D object, whose geometry and position are the source of truth for shape and placement, the image reference above it defining the surface, and the final prompt over both; the camera is set separately in Compose.
The workflow that keeps the asset intact is to lock it as fixed geometry and let AI change materials, lighting, and environment instead of regenerating the object each pass.
| Approach class | Input it takes | Can geometry drift | Skill it assumes | Genuinely best at |
|---|---|---|---|---|
| Prompt or image-first generators | A prompt plus reference photos | Yes, regenerated each time | Prompt craft and a per-frame check | Exploring when the object is not judged |
| Conditioning-based generators (ControlNet, Veras) | A flat picture of your model: sketch, edge or depth map, viewport | Constrained; on ControlNet pipelines, at a weight you set | Your CAD tool and a review pass | Photoreal images from a sketch or working view |
| In-tool real-time renderers (Enscape, D5 Render) | The live model open in your design tool | Not in the live view; the AI features generate | The design tool, plus lighting work | Judging a design as it changes |
| CAD-native offline renderers (KeyShot class) | The CAD file itself | Not on the ray-traced path; the AI Shots modes generate | The most: materials, lighting, camera craft | Broadcast-grade final frames |
| Browser-based 3D scene generation (Intangible) | The 3D scene: imported mesh, product photo, your camera | Not from the model: the scene sets shape, scale, and placement; the surface is generated | Little: place the object, attach the reference, move the camera | A repeatable multi-shot set from one scene |
On its ray-traced path an offline renderer holds geometry perfectly, since that path renders geometry and never generates it. KeyShot imports native SOLIDWORKS, CATIA, Creo, and NX files as well as STEP and IGES. Its AI Shots modes are the exception, scoped in Luxion's manual: Restyle will keep your model but let you explore different lighting and design options, while Imagine allows you to generate entirely new images without keeping to your scene.
One asset, end to end: a branded endcap display for a beverage SKU, header card on top, bottles on the trays, a dozen frames wanted, logo identical in every one. Intangible says imported CAD and brand assets render faithfully. That describes the mechanism, and each render still needs checking.
Intangible's own import answers page says it supports importing your own 3D assets and CAD in FBX, OBJ, DXF, and GLB, binds references to them so they stay faithful across every angle, and exports to image, video, or glTF. The importer's help page publishes .glb, .gltf, .dae, .fbx, .ply, .obj, .stl and .usd instead, so FBX, OBJ, and GLB sit on every list Intangible publishes; export one of those three. STEP and IGES sit on neither, so convert those first; a DXF file comes in as line geometry rather than solid mass.
Check the tight crop first. It magnifies the label, so the three things that go wrong are easiest to see there. Zoom the wordmark and compare it against the reference photograph, never from memory.
Check the other two anyway. Intangible says that because every shot comes from the same 3D scene, characters, lighting, references, and camera paths hold across a multi-shot set rather than being regenerated per frame.
Engineering geometry works the same way. Import the CAD assembly, an airframe say, and that mesh is the scene's source of truth, so its shape, scale, and placement hold slide to slide, while surface and markings are still generated. For the full CAD-to-deck sequence, see CAD model to investor render.
Keep these renders out of engineering decisions. Tolerance review, manufacturing sign-off, and anything built against the image belong with a CAD-native renderer on the untouched CAD file.
Intangible publishes several reference-layer limits; the one to plan around is naming. Its documentation says the final prompt wins over both the 3D object and the image reference, and the object name feeds that prompt, so a recognizable name overrides the reference.
You need a mesh first. Intangible's Generate 3D Asset docs say you upload one photograph and the system generates a mesh that matches the visual. Attach that same photograph as the reference.
For silhouette and framing it does real work. For typography on a label, treat it as guidance at a weight you set and check every frame.
Intangible says every shot comes from the same 3D scene and exports to image, video, or glTF, so both outputs are sourced from the same geometry.