Industry Voices | The Road to 99.999%: Eliminating Blind Spots in Automotive Vision AI
Automotive vision AI needs to achieve 99.999% accuracy, but most projects fail due to data problems. This article analyzes the causes of data blind spots, such as manual annotation bottlenecks and the disconnect between data and models, and uses the "phantom pothole" case to show how integrated data management can quickly identify issues, helping achieve the "five nines" goal.

As the automotive industry races toward a more autonomous and intelligent future, visual AI is critical in shaping this transformation. Despite our efforts to build trust in this emerging technology, the reality is that most AI projects fail—with failure rates as high as80%—and data issues are often the root cause.
To successfully compete in this space with visual AI, it is essential to eliminate the blind spots around data and AI models. When we clearly see where the problems lie, we can iterate quickly to improve, reduce failure rates, and deliver continuously optimized experiences for human drivers and road sharers.
AI's data challenges are not unique to the automotive industry. However, the consequences of AI failure in automotive applications can be a matter of life and death. Autonomous vehicles and driver assistance features must achievefive nines of accuracy and reliabilityto avoid being deemed a failure—meaning surpassing human behavior and operating safely and accurately 99.999% of the time.
For example,in accidents at intersections with stop signs, 70% are due to failure to come to a complete stop. Anyone with driving experience has likely shaken a fist at a driver who rolled through a stop sign without slowing down. An autonomous vehicle that stops at only 95% of stop signs, while a significant improvement over human performance, still poses a clear danger.
To meet this high standard and achieve "five nines," we must train systems to handle the widest possible range of scenarios—including those we cannot foresee. Furthermore, by accumulating the right visual data and clearly understanding this data and its impact on AI models, we can quickly and efficiently find solutions when problems arise, rather than blindly throwing more data into the dark.
Illuminating the Data Problem in Visual AI
As a graduate student researching AI in the early 2000s, I spent an entire year testing the effects of algorithm parameter changes on datasets of 60 to 80 images. Today, datasets are massive, reaching hundreds of millions or even billions of samples.
At such scale, relying on manual processing of all data is no longer feasible. For example, we have long sincepassed the tipping point of relying solely on manual data annotation, let alone manually managing the interaction between data and algorithms, or manually combing through data to find target matches. The result is that when projects go off track, we often grope in the dark, unable to determine why.
Consider the case of a visual AI innovation company: its lane-following software began identifying "phantom potholes." The model had not changed, yet no one knew why the car suddenly swerved to avoid obstacles that did not exist. To find the root cause, the team had to obtain the visual data—all 3 million images.
Due to the lack of a unified view of the situation, this blind spot could bring the project to a halt. The company struggled to collect and process massive amounts of data, passing it among multiple busy teams, then repeatedly reviewing and testing data and models through a series of complex steps before pinpointing the specific data causing the issue.
Addressing Visual AI's Data Challenges
Improving data quality and tuning model parameters is an iterative and lengthy process. Adding to the complexity, AI development is often siloed, with different teams, tools, and processes managing data and models separately. Add external vendors into the mix, and you have bottlenecks where teams struggle in the dark to identify the cause of each new issue.
Ultimately, if you cannot identify the data "culprit," it becomes difficult or impossible to improve the model to meet performance requirements, stalling new projects in R&D and making maintenance of existing models (like the "phantom pothole" case above) complex and costly. When teams have visibility into how data and models interact, they can also use algorithmically generated synthetic visual data to fill gaps, further enhancing models to represent real-world scenarios.
Driving Visual AI Project Success
Adopting an integrated approach to building data models empowers teams to take control of the development process, eliminating obstacles that slow progress—and the blind spots that cause projects to stall or fail. Beyond simply viewing data, teams equipped with tools to query, analyze, and optimize data, and to reveal its impact on model performance, can iterate quickly, improving both data and algorithms simultaneously to bring powerful AI models into production.
In the "phantom pothole" case, the organization placed teams, data, and models on the same platform. They quickly inspected the anomalies and easily identified that an unusual storm had blown branches onto the ground, casting shadows. Because this storm scenario was not included in the training data, the model misclassified the shadows as potholes.
Instead of stalling, the team quickly iterated on the model to resolve the issue—allowing them to continue meeting the "five nines" standard and maintaining customer confidence in their autonomous driving products. To create and maintain AI systems that achieve "five nines" and fulfill the promise of autonomous vehicles and intelligent driver assistance, we must iterate faster and take control of the data and algorithms driving visual AI. By replacing fragmented, rigid tools and cumbersome manual processes with modern technology, and unlocking the insights that connect data and models, we can begin to reduce failure rates and make road travel safer and more accessible.
About the Author
Jason Corso is a professor of robotics, electrical engineering, and computer science at the University of Michigan, and co-founder and chief science officer at Voxel51. Voxel51 helps enterprises build production-grade AI models and applications using visual data.