How to Use Blender to Train AI Vision Models

2026-09-30 • Maciej Maciejak

Computer vision models now sit inside a growing number of systems: quality control on industrial production lines, self-driving vehicles, robots that pick parts from bins, medical imaging tools, self-checkout counters in retail and crop monitoring in agriculture. Computer vision itself is not new. What has changed is the method. Classic approaches built on hand-crafted features, such as HOG (Histogram of Oriented Gradients), SIFT keypoints, Haar cascades, edge detection and template matching, are giving way to neural networks.

Whether you use a convolutional neural network (CNN) like YOLO or a vision transformer, one requirement does not go away: you need labelled training images, and a lot of them.

Why Render Training Data Instead of Photographing It

The usual route is to take thousands of photos and then pay someone to do the tedious job of drawing boxes or masks around every object in every image. It is slow, it is expensive, and the labels are only as good as the attention of the person drawing them.

There is another option: render the images from 3D models. Because the software knows exactly where every object is, the labels are generated together with the image and are pixel accurate. You also get full control over the scene, so you can deliberately produce the edge cases that rarely show up in real photos, such as unusual lighting, odd angles or rare defects. The result is lower cost and better labelling quality at the same time. We covered the case for synthetic data in more detail in https://csmx.eu/blog/why-3d-synthetic-data-is-the-secret-to-vision-models. This post is about how we actually build it.

Our Tool of Choice: Blender and Its Python API

At CSMX we generate synthetic data with Blender, the open source 3D software. Blender comes with a Python API called bpy, which gives scripted access to almost everything you can do in the interface: loading models, changing materials, moving lights and cameras, and rendering.

We use Python and bpy to build a pipeline that generates annotated datasets for training, validation and testing. Once the pipeline is set up, producing another thousand images is a matter of running it again with new random parameters.

Let us look at how we do it in Blender, and which tools we rely on.

Lighting and Environment Augmentation

Augmentation, meaning deliberate variation of everything in the scene, is crucial if you want a model that generalizes from renders to real photos. A model that has only ever seen one lighting setup will fail the moment the light changes.

We use hundreds of HDRIs (high dynamic range panoramic images) as background lighting, and we randomize their exposure and rotation for each render. Depending on the use case, we add other light sources as needed and randomize their:


Shape Variation in the 3D Mesh

Depending on the object, some variation has to happen in the geometry itself. Real parts bend, dent and differ slightly from one another, and the model needs to see that. Blender gives us several tools for it:


Material Variation Beyond Random Numbers

Almost every real object differs from the next one in its surface: colour, gloss, scratches, dirt. Simply plugging random numbers into the channels of a PBR shader (base colour, roughness, metallic) usually does not cut it, because the result looks like a set of clean, slightly recoloured copies.

Instead, we build custom node groups for specific effects, such as ageing, wear and tear, and deposits of other materials, like dust or grime, building up on the object's surface. Each effect has its own controls, so the pipeline can randomize them in a way that stays physically plausible.

Geometry Nodes for Complex Objects and Defects

Some objects are really complex, or the anomalies we want the model to catch are. For those we bring out the big guns: Geometry Nodes, Blender's procedural modelling system. It lets us generate complex objects and defects from rules and parameters, instead of modelling each variant by hand.

This is also the way to go when a model has to work on a general class of objects rather than on one particular object. A quality control model for a single factory line only ever needs to recognize one specific part, for example one type of industrial connector. A model for a self-driving car has to recognize any car, any pedestrian or any road sign, including ones it has never seen. Procedural generation can produce a wide range of shapes within a class, so the model learns the class and not one example of it. You can see a simpler version of the idea in our project where we rendered many variations of donuts to train a model that counts them on a production line: https://csmx.eu/portfolio/project/synthetic-ai-training-from-renders-for-counting-donuts.

Camera, Pose and Background Randomization in Python

You might ask about randomizing the object's rotation and position, the camera's focal length, or the background images. We do all of that too, but usually directly in Python, since these are simple parameters that bpy can set before each render. The same script also writes out the annotations for each image and sorts the results into training, validation and test sets.

For a real example of the whole pipeline, see our case study on synthetic data for quality control of industrial connectors, built together with Panda GmbH: https://csmx.eu/portfolio/project/demo-synthetic-data.

Need Training Data for Your Vision Model?

If you are looking to train a computer vision model but are struggling to source the data for it, visit our dedicated page at https://csmx.eu/synthetic-data for more information. Or better, schedule a call today at https://calendly.com/maciejak28/30min, or reach us through our contact page at https://csmx.eu/contact.