A self-driving car in the near future has one job where it needs to be highly accurate: identifying traffic signs. Mistaking a stop sign for a speed limit can put people's lives in danger. However, the AI models doing this can be tricked by adding changes so small a human wouldn't notice them.

These tricks are called adversarial attacks. You take a real image, add a tiny bit of carefully calculated noise, and the model confidently calls a stop sign something else entirely. The picture could still look identical to you and me.

Back in 2024, I ran an experiment (github repo here) to test a simple idea: if I top up the training data with synthetic images (fake traffic signs generated by an AI) does the model get more resistant to adversarial attacks?

The problem with the dataset

I used the GTSRB dataset, which is the standard German traffic sign benchmark. It has around 39,000 training images across 43 sign types, but it's lopsided. Some signs show up over 2,000 times, others barely 200. Models trained on data like this get good at the common signs and unreliable on the rare ones. That's a problem when the rare sign is the one that matters.

Training image counts across the traffic sign classes tested
Image counts across the 8 traffic sign types I tested. Some have way more training data than others.

The usual fix is data augmentation such as rotating, scaling, or adding random noise to the images you already have. I wanted to test something different: generating synthetic images to fill the gaps, and seeing if that did anything for security on top of fixing the imbalance.

To keep it manageable on my hardware, I narrowed the test to 8 sign types instead of all 43.

Generating the fake signs

To make the synthetic images, I used a diffusion model, which is the same family of tech behind image generators like DALL-E and Midjourney. Specifically a Diffusion Transformer (DiT), which was close to state-of-the-art when I did this experiment. I trained the transformer on the real signs, then had it generate new ones so that each of the 8 traffic sign types had an equal 2,250 images.

Sample synthetic traffic sign images generated with DiT
Sample synthetic images generated using DiT.

This was by far the slowest part: about 165 hours of generating on an RTX 3090 GPU.

The two attacks

I then trained two identical traffic sign classifier models: one on the original imbalanced data and one on the balanced data with synthetic images mixed in, and attacked both.

I used two well-known attack methods:

  • Carlini-Wagner (CW): slow and computationally heavy, but produces a stronger attack. The added noise can be slightly more visible.
  • DeepFool: fast and efficient in finding the smallest possible nudge to flip the model's decision. The noise is nearly invisible.
Traffic sign adversarial perturbation examples
What a successful attack looks like. Top to bottom: the original image, the added noise, and the result that fooled the model. CW on the left, DeepFool on the right.

The results

Traffic sign predictions before and after adversarial perturbations
The attacks in action: each original sign next to its perturbed version, showing the real label and the model's (wrong) guess. CW on the left, DeepFool on the right.
Original model Model with synthetic data
Accuracy on clean images 97.67% 96.22%
Accuracy under CW attack 15.11% 21.46%
Accuracy under DeepFool attack 7.51% 80.82%

First, the synthetic data did almost nothing to improve accuracy (it actually dropped slightly). So if your only goal is a more accurate model, this is probably not the best approach.

But second, look at the attack columns. Against DeepFool, the original model collapsed to 7.5% accuracy, or basically useless. The model trained with synthetic data held at 80.8%, which is a massive jump in model robustness. Against the stronger CW attack, the gain was smaller but still real.

There was also an interesting pattern in the attacks themselves: the slower, more visible CW attack was both harder to defend against and easier for a human to spot (subjectively). The fast, invisible DeepFool attack was the one synthetic data prevented most effectively.

What it means

Adding synthetic training data made the model meaningfully harder to fool, even though it didn't make it more accurate. My guess is that the extra variety in the training data prepares the model for inputs it wouldn't otherwise have seen, including subtly manipulated ones.

This doesn't replace more established defenses like adversarial training. But those methods are expensive and have their own weaknesses, while synthetic data is a cheaper tool that does double duty: fixing dataset imbalance and adding a layer of robustness. For safety-critical systems like self-driving cars, having another option in the toolbox matters.

A 2024 caveat

This experiment is from 2024, and the field of generative AI moves fast. The diffusion models available today are far better and cheaper to run than what I used, so those 165 hours of generation would be a fraction of that now. The attack methods and defenses have also evolved. So treat the specific numbers as a snapshot, not a final word. The core finding that synthetic data can improve robustness without significantly affecting accuracy is the part I'd still stand behind. It might be worth re-running the experiment at full scale with current tools.