Edge Cases and the Long Tail of Safety
Remote assistance is the unlikely hero, but we need to get serious about its effectiveness
By Phil Koopman
Autonomous vehicle safety is won or lost on the so-called edge cases. Recent recalls and news reports have made it start to feel like we might be losing. But we can get back on track if we acknowledge some hard truths about what edge cases really are and what it takes to beat them. The pieces are already on the playing board, but we’re not getting straight talk about what it takes to win.
What do we mean by edge case? Is a kid in a yellow raincoat and red boots amidst reflecting red lights an edge case for a robotaxi? And if so, what should we do about it?
The generic meaning of the term “edge case” is an operational problem that occurs in an extreme situation of some sort.1 But in the Autonomous Vehicle (AV) world, the term has a slightly different meaning of a situation that is “rare” or “unusual” — with some confusion as to whether that necessarily means it is somehow extreme or a boundary condition.
The problem with the use of the term “edge case” in the world of AVs comes when someone tries to imply that an edge case is no big deal for safety, due to it being said to be rare or unusual. Far from it. Any time an edge case is encountered there needs to be a robust risk mitigation approach in place, involving both an operational response and an engineering feedback response. At least some engineering teams take those responses seriously. But too often, public relations postures use “edge case” as a deflection excuse for an embarrassing news story. Nothing to see here. Move along.
So let’s dig in to edge cases. How we should think about them. Why they matter to safety. And why they are a problem that is never going away.
Out-of-distribution & rainy days
A common way of looking at an edge case in the machine learning world is that it involves a situation that is “out-of-distribution.” Machine learning involves collecting training data and extracting statistical characteristics. But that training data has limitations based on the conditions in which it was collected. Data collected in different conditions might be part of a different statistical “distribution.”
For example, driving data might only be collected on sunny days. Driving in rainy days would have different statistical characteristics for camera images and vehicle control responses, resulting in an out-of-distribution situation. For that reason, we expect that an AV trained only on sunny day driving might well struggle on rainy days.
At this point, a picture might be helpful. There are two flavors of edge cases shown in the figure below.2 AVs usually have many more dimensions, so this is a simplified example of data in a two-dimensional space.
Figure: Simplified visualization of clustered training data and two types of edge cases.
The black, clustered points in the figure show training data. The presumption is that any new data point that appears inside a cluster is in-distribution, and will be handled properly. In this case there is a group of training data taken on dry days, and a different group of training data taken on wet days. Perhaps this training data has to do with how slippery the road is, or maximum advisable vehicle speed.
Points A and B in the figure are two somewhat different out-of-distribution data points that might be seen in real-world operation.
Point A lies between the two clusters of training data. This means there is an in-between situation with insufficient training data. For example, this might be damp road surface after rain has stopped, with no puddles, but not entirely dry either. Without training data, the machine learning system could arbitrarily treat point A as entirely dry or entirely wet rather than its own separate environmental condition.
Point B lies far outside the clusters of training data, and is clearly an outlier. Data point B might be a system encountering a frozen road, or perhaps flooding. If those situations are missing from training data, the AV will have no basis for responding to those conditions.
Both points A and B might not be handled well by the AV because it has not been trained on that type of data. For point A there might need to be some sort of in-between response, or it might be an entirely new cluster of related points missing from training (damp roads without active rain) that needs significantly more training data. For point B, there might be no basis to extrapolate from training data, so new training data for that novel situation would be needed.
Both data points A and B represent the usual discussion about out-of-distribution situations. And it is easy for people to relate to these issues. If you do not have training data in rain, snow, damp roads, or floods, then you can’t expect an AV to do well in those conditions. The usual refrain is: get more data.
But there is a much tricker problem to consider. Sometimes a data point all out on its own is legitimate but rare part of the real-world statistical distribution that training data is meant to represent. It is simply so rare that it is missing from the training data. Sometimes pavement is slippery even in dry conditions3, but those rare situations are missing from training data.
“Tail” data matters too
Consider collecting training data on Powerball lottery tickets to determine what the payoff probabilities are.4 You look at outcomes from a million purchased tickets (sounds like a lot!) and use that for your training data. But your training data is unlikely to include a Grand Prize, or even a $1 million prize payout. Those things exist in the real world. And in some sense they are the most important data points in the system (do people dream of winning $100 in the lottery vs. the Grand Prize?) They are just too rare to be included in your training data.
Severe car crashes amount to being an inverse lottery winner.5 They are rare on a per-mile basis. But that does not mean they don’t matter.
These rare events are often referred to as “tail” events because they live in the low probability section of a statistical distribution that looks sort of like a tail hanging off to the right of a statistical distribution diagram.
Edge cases in open world operation, such as those encountered by AVs, are expected to be a particular kind of statistical distribution that has a “heavy tail” or “long tail.” That means there are huge numbers of rare events. Each one is individually unlikely, and so is probably not in the training data. The issue is that, when aggregated, there are so many the risk from them cannot be ignored.6
Here’s a quick explainer video (3.5 minutes):
Edge cases and contextual clues
Even this explanation under-states the complexity of edge cases. Edge cases might not involve any single thing that is novel, but rather, a previously un-trained novel combination of ordinary characteristics.
Some edge cases are crazy things missing from training data, such as a stop sign to yield to crossing jet fighters.7 Other edge cases are attributes missing from training data, such as fair weather data missing people wearing yellow rain coats, resulting in the system concluding that anything yellow must be a road sign.8
However, another type of edge case involves a combination of already-known attributes providing context that dramatically affects expected vehicle behavior. For example, consider a pedestrian, a pole, and a stop sign. Individually, all of these are going to be in any respectable training data set. But it is the combinations and context that matter. Here are two similar (but not identical) situations with dramatically different vehicle behavior.
Stop sign on a fixed pole with person resting a hand on the pole: vehicle stops, waits for pedestrian to let go of the pole and cross (assuming crossing is intended), then goes.
Stop sign on an unfixed pole with a person holding the pole: vehicle stops, and stays stopped until the construction worker flips the sign to say “slow” instead of “stop”
For these examples the essential difference is whether the pole with the stop sign is permanently embedded in the ground or on a hand-held pole. There are other contextual clues that might help: whether the person with a hand on the pole is in hi-viz clothing (hopefully yes), whether there is a construction zone visible (it might be around a corner), whether the base of the pole is in contact with the road surface rather than on the side of the roadway (depends), and whether the placement and type of pole are representative of a permanent stop sign vs. a hand-held sign. A high-definition map might also help, but can be imperfect and might not be available for some vehicles in some locations.
Similar subtleties are involved at interpreting the meaning of a stop sign involving school children: hand-held stop paddle extended (stay stopped), hand-held stop paddle held loosely at side of body, potentially inverted (disregard), stop sign mounted on school bus extended (stay stopped), stop sign mounted on school bus folded back (disregard), and young children not paying full attention walking to school at a 4-way stop sign intersection (exercise extra caution; many localities deem 4-ways stops to be sufficient traffic control without a crossing guard).
And what if a student at a street corner has a large stop sign graphic on their t-shirt, or a stop-sign decal on their backpack? Does the AV recognize them as a pedestrian, or does the stop-sign detection over-rule pedestrian detection?
Edge cases are in the eye of the beholder
It is common when looking at a photo or video of an embarrassing robotaxi incident see AV supporters blaming it on a “rare” edge case, accompanied by a plea to take pity because they’re still learning. Other commenters might insist that it can’t be an edge case because it is happened to them personally, and therefore should be expected. Commenters defending the robotaxi might well reply with some combination of no harm was done so no big deal, or robotaxis are always improving so this one shouldn’t count, or We’re Saving Lives so stop complaining. And so on.
While the above types of comments might have some merit, such discussions tend to miss fundamental issues related to AV safety:
It does not matter whether a person thinks a situation is rare, or even if the situation is objectively common. If it is under-represented in the robotaxi training data, it is an edge case, and might be handled improperly in the wild.
Risk matters more than frequency in life-critical systems. A rare but high-severity event must still be handled to avoid a severe loss event.
A potentially severe event that does not happen to result in harm is still a safety issue. Getting lucky doesn’t scale.
A useful definition of edge cases for AV safety is that they are things under-represented (or entirely missing) in the training data. Those things might be rare objects, rare events, rare combinations of common objects & events, or even common objects that were missing from the training environment. Think heavy parkas in summer, shorts in winter, leaves on deciduous trees in winter, ice cream trucks in winter, and hi-viz construction vests in training that avoids construction zones.
Absent an equipment defect, if an AV makes a mistake, it has by definition encountered an edge case in the form of something it was under-trained on.9
Whether someone thinks the edge case should have been addressed before an incident has occurred is a legitimate concern, but does not enter into whether it is an edge case from the AV’s point of view.
Fixing edge cases — you can’t catch them all
The obvious approach to fixing a problem caused by an edge case is to identify the nature of the edge case and create training data to address the problem. When you find out it does sometimes rain in the desert, train with rain. If you have trouble with going by school buses, train more on school buses.
But there are two problems. The first problem is that there are, for practical purposes, an infinite number of edge cases. The world is filled with a bunch of crazy stuff. But there is also a huge number of context-sensitive combinations of ordinary stuff that informs good driving behavior. That context includes not just the environment, but also the local driving culture and knowledge of local events.10
The second problem is, the training to fix an edge case might not work.11 The jury is still out on why this might happen, but I speculate this has to do with the use of end-to-end (E2E) machine learning. It looks like E2E approaches make it easier to get sophisticated behaviors with less up-front engineering effort. However, there are reasons to believe that if the system learns a bad behavior, it can be arduous to fix the situation by figuring out what training data will change the bad behavior.
When AVs were starting, the solution to edge cases was to collect more data. If one had enough data, the theory went (and still goes sometimes), you’ll eventually capture all the edge cases, and life will be good.
A few tens of billions of dollars later, companies have come to realize that there are too many edge cases to catch them all.
However, the safety question was never about catching them all. The relevant question is can we catch enough of them that we are willing to live with the residual risk of novel edge cases causing mishaps? Perhaps. But looking at robotaxi news it is obvious we are not there yet. And getting there is easily a decade away. If it can be done at all.
I think we will likely not solve the heavy-tail training problem in the foreseeable future.12 The world is always changing, and the distribution of rare events might be so heavy-tail that it is just impractical to catch enough rare events to handle everything autonomously.
In fact, the AV industry has already admitted defeat on this topic. They just don’t like to talk about it that way.
Compensating for edge cases with remote assistants
Humans have edge cases too. Drivers eventually see something they have never encountered before, and find a way to muddle through. Some of it is based on robust world models that allow accurate extrapolation beyond training data. But some of it is simply being aware enough to know that something doesn’t feel right, and reducing risk exposure to buy time to figure out how to muddle through.13
People are imperfect, but common sense on average is better than some give credit for. In contrast, machine learning has no common sense, and is uneven at best at faking common sense.
So what if we use people to provide common sense to solve novel edge cases? AV companies have been doing exactly that. They’re called remote operators and/or remote assistants.
Faced with an urgent need to expand operations without continuous safety drivers, AV companies have pivoted to a hybrid computer/human approach. Computers handle most of the driving most of the time, but when they need that human ability to handle edge cases, they contact a remote human operator to get help.
There are various roles the humans play, which I would argue all have safety responsibilities.14 An especially critical role is that of providing decisions for situations in which the AV has low confidence, such as whether a traffic light is red or not. It is important to keep in mind that, despite the protestations of AV CEOs, these remote assistants are not just providing “phone a friend” optional advice. Rather, they play an essential role in ensuring acceptable operational safety when edge cases are encountered during real-world operation.
There’s nothing wrong with using remote assistants — if done safely. In fact, I would argue that outside of highly constrained operational limits, robotaxi technology is not currently viable for general use without remote assistants. And it will be a long time before that problem is solved for operation at scale.
The question is not whether remote assistants are needed to handle edge case. Rather, the questions that matter are: (1) Can a remote assistant approach handle edge cases effectively enough to provide acceptable safety? and (2) Can we get by with few enough remote assistants for the services to be economically viable?
Current robotaxi services are a computer+human driver hybrid. The computer mostly drives. A remote assistant weighs in sometimes to provide decision support or for other roles. The net safety of that approach depends on getting four things right:
The robotaxi has to know when to ask for help. This is a huge technical challenge, because machine learning by nature tends to be over-confident when it is acting on edge cases. There is work in this area, but this is far from a solved problem. If the robotaxi does not know to ask for help, it is taking actions it has not been trained for, and can arbitrarily fail in sometimes-spectacular ways.
The remote assistant has to have adequate situational awareness. This might need to include reviewing events that led up to the request for help since not all important information might be obvious in a real-time view. Was a pedestrian on the ground struck or just tripped? Was there a “road closed” sign the robotaxi blew by a quarter-mile back? Might there be something under the vehicle’s tires?
The remote assistant needs the time and the control ability to resolve the problem. It will take remote assistants time to gain situational awareness. Potentially tens of seconds or longer depending on how complex the situation is. Meanwhile things might still be happening on the road. As a practical matter there might be limits on control authority such as a policy preventing remotely moving vehicles due to other safety concerns. So the question is, even if the remote assistant has the time, can they remotely resolve the situation?
Remote assistants will sometimes make mistakes that the computer driver has no way to mitigate. As a simple example, if a computer driver isn’t sure about traffic light color and asks a remote assistant, and that remote assistant gets it wrong, it is unreasonable for the computer driver to somehow recognize what the true traffic light color is in time to prevent it from accelerating into a red light. There needs to be training, qualification, monitoring, and other mitigations in place to keep dangerous remote actions to an acceptably blow level.
Safe enough handling of edge cases at scale
In a very real sense, every AV incident we see in the news is a failure of one of these four issues: the AV fails to ask for help in time, the remote assistant fails to accurately build a mental model of the situation, the remote assistant does not have the time and/or control authority to resolve the situation, or the remote assistant makes a mistake. And don’t forget that we are asking those human remote assistants to help out at precisely the time that the robotaxi is confused, so they might well be dropping in on a complex, messy situation under pressure to resolve things quickly.
It us unrealistic to require the remote assistance approach to handling edge cases to be perfect. But it needs to be robust enough to provide acceptable safety along all the different aspects of safety relevant to stakeholders.15
The reason we’re seeing more concerns about edge cases relate to the attempts to scale up operation on public roads. More vehicles with more miles in more cities means more exposure to edge cases. This was always going to be an issue with robotaxi scaling, and now we are starting to live through it.16
In the early stages of scaling perhaps it was enough to collect data and fix things as you saw them. That’s what they did in the early days of aviation as well. But we would never have gotten to where we are on aviation safety by waiting for planes to crash before we fixed problems. Safety of large complex systems comes from being proactive, following industry safety standards, and independent checks. For edge cases that means fixing edge cases when you see one for the first time rather than waiting for an embarrassing news story. And getting ahead of the curve to address inevitable edge cases even if they haven’t happened or the engineering team has not noticed them happening.
AV companies should also think harder about at-scale failures, including how to ensure adequate remote assistant staff are available for a large surge in robotaxis encountering edge cases at the same time due to an incident such as a power outage. And they shouldn’t ignore what the plan needs to be for hurricane and earthquake response just because they haven’t been through one — yet.
As long as we see business pressure to continue scaling up robotaxi service, we’ll see increasing problems from edge cases. Success will require the humility to acknowledge that remote assistants are a safety-critical part of the edge case handling process.
We need to see a more proactive approach to treating edge case incidents as central to safety rather than as publicity gaffes to be quashed. Additionally there is more hard work required to figure out how to get edge case retraining to be effective in practice for end-to-end machine learning.
Edge cases are not an embarrassment should be swept under the rug. They are the core problem that must be addressed to have safe autonomous vehicles. And they are here to stay.
Phil Koopman has been working on self-driving car safety for about 30 years, and embedded systems for even longer. For more on applying AI, see his new book: Embodied AI Safety.
Alternate meanings are: (1) a boundary between two different behaviors causing an issue, and (2) a value is at an extreme maximum or minimum. “Corner case” is a similar term that involves two or more simultaneous data values encountering an edge case. See: https://en.wikipedia.org/wiki/Edge_case
I’m not attempting to be mathematically rigorous in this piece. Rather, I’m trying to convey intuition as to the types of edge case situations that can arise for a non-specialist audience.
Spilled liquids, oil, or sand can reduce road surface friction even in dry conditions.
With a million tickets you have about good odds at a single $50k prize, and are unlikely to win a $1M or grand prize. This is a simplified example. Training machine learning with lottery tickets won’t help you win a fair lottery.
The official odds are published here: https://www.powerball.com/powerball-prize-chart
Useless fact: the odds of dying in a car crash are roughly equivalent to winning the Powerball Grand Prize if you bought three Powerball tickets for every mile you drive. Powerball Grand Prize odds are 1 in 292,201,338. This is not investment advice.
See: Koopman, The Heavy Tail Safety Ceiling, 2018: https://users.ece.cmu.edu/~koopman/pubs/koopman18_heavy_tail_ceiling.pdf
I’ve actually seen that at runway crossings at naval air stations.
My team found that type of failure in an early perception system. It had apparently not been trained in rain or in construction zones, and so had trouble detecting people wearing bright yellow clothing.
Some might object to this statement, but I believe it is true. In my opinion the objections have to do with whether it is possible to improve the training, which is covered in the next two sections.
Among examples I’m familiar with: parking chairs, families walking to worship on Friday evenings, the chaos near schools on the first day of school dropoff line, construction detours for never-ending bridge repairs, students racing across a heavy traffic route from public buses if it happens to be a few minutes before standard university class start times, road closures for foot races, Anthrocon, Halloween, and the infamous Pittsburgh Left.
Marshall, A., “A School District Tried to Help Train Waymos to Stop for School Buses. It Didn’t Work,” Mar. 29, 2026. https://www.wired.com/story/a-school-district-tried-to-help-train-waymos-to-stop-for-school-buses-it-didnt-work/
Attempting to use Generative AI to create training data might help with some combination-based edge cases. But it will have limits based on its statistical nature. It is unclear if GenAI will solve the edge case problem well enough, and I’m betting probably it will help but not be a sufficient solution.
I once saw a person on the side of the road and a dog in the middle of the road. If I had not thought something was off, I might have driven between them, but I slowed down because something about it felt risky. After taking a closer look I realized the dog was on hind legs only, which is odd. I deduced, and later confirmed that it was straining against a not-visible-to-me cable leash. If I had proceeded I would have snagged the leash and dragged it down the road, along with the dog and potentially the owner. I’m glad I slowed in response to my “that’s odd!” feeling. This is yet another example of how very many combination-based edge cases might be out in the real world, and how people muddle through novel situations.
See: Koopman, What’s the Deal with Robotaxi Remote Assistants?
See: Koopman, Robotaxi Scaling Is Just Beginning



“…we would never have gotten to where we are on aviation safety by waiting for planes to crash before we fixed problems. Safety of large complex systems comes from being proactive, following industry safety standards, and independent checks. For edge cases that means fixing edge cases when you see one for the first time rather than waiting for an embarrassing news story.”
Exactly.
In aviation we typically find that minor incidents outnumber major incidents/accidents, often by several orders of magnitude. For every Waymo that got stuck in floodwaters (and publicized), there were probably several that drove through smaller amounts of water and did not get stuck or lose control. To use the wet/dry training data example, in those cases there would be a few dots in between the wet data and point B (Flood) where the road was under a manageable amount of water (but that fact may not have been known prior to entering it).
If companies can address those situations sooner through a proactive process then they can start to address some of the most severe risks instead of waiting for the next (hopefully only) embarrassing news story. Has there been any discussion of an aviation Safety Management System-type framework for AV manufacturers/operators?
Thanks for the post! As I’ve mentioned, I assign your work to my students. You have inspired me to move my own content onto Substack, and I’ll circle back in some future post along the following lines…
I see a need to look at AVs from a different perspective. It’s the basic premise of scientific research that if you understand a system, you can predict its behavior. (Think of residuals and regression analysis.) Famously, any research performed on swans prior to the 19th century would come to the certain conclusion that “Swans are always white”, until the Dutch (?) explored Australia and found black swans. Systems are easily predictable if we simply ignore relevant facts in the surrounding environment.
Except for the “edge cases”.
That these keep appearing implies that the AV mfrs cannot conceive of Black Swans. Notably, the mfrs are unable to emulate human drivers’ cognitive ability in situational awareness (your dog leash example). Waymo is quick to discuss reaction times (girl falling from a scooter, braking before hitting a child in a school zone), but fail to discuss how to avoid the situation altogether. That AV mfrs attempt to address an infinite supply of “edge cases” implies that they cannot predict (and thus do not comprehend) the more systemic behavior of our auto-based transportation system.
Which brings us to Daniel Kahneman and the human brain’s reliance on easily available information, to the exclusion of objectively valid counterpoints. When asked a hard question, we instinctively choose to provide an answer to an easier, unstated, question.
I see mega-billion dollar business models based on solving the driving problems (reaction times) that technologist believe that they can solve, rather than the transportation solutions required by society.
Crashes lead to $1B judgements, eventually the money dries up, and we can attempt a more systemic solutions.