ALFIE ETD-HUB

1: What are Acceptable Algorithm Performance Standards?

Asked: 6 months, 1 week ago By: Catalink Views: 135 Catalink Case Study: IRIS

Considering the high-risk implications of False Positives (i.e. detecting fatigue when none exists) and False Negatives (i.e. failing to detect actual fatigue), what regulatory or industry-mandated level of performance, accuracy, and reliability is required for commercial deployment?

39 Answers

Answered: 4 months, 3 weeks ago By: Chiamakaokorie
-
Answered: 4 months, 3 weeks ago By: Tundefasina
DSM systems like IRIS should meet very high safety thresholds, typically >95% overall accuracy, extremely low false-negative rates, and tightly controlled false positives. Compliance with ISO 26262 (functional safety) and ISO 21448 (SOTIF) should be mandatory, along with proven reliability across diverse demographics and real-world conditions.
Charlie replied: For sure, also, for an EU vehicle deployment, the main hard requirements would come from the General Safety Regulation framework and the specific DDAW rules. EU Regulation 2019/2144 requires Driver Drowsiness and Attention Warning systems for M and N category vehicles, with the obligation applying to new vehicle types from 6 July 2022 and to all new vehicles from 7 July 2024. DDAW systems are defined as systems that assess the driver’s alertness through vehicle-system analysis and warn the driver if needed. The detailed DDAW regulation does not prescribe one model architecture or one fixed accuracy percentage. Instead, it sets functional and validation requirements. The system must operate under defined conditions, including automatic activation above 70 km/h and operation in daytime and night-time conditions. It must warn the driver at a drowsiness level equivalent to or above level 8 on the Karolinska Sleepiness Scale, although it may warn from level 7 onwards. The validation regime is also important. Manufacturers must compile a dossier explaining the DDAW system, its operation, the test procedures used, the rationale for those procedures, and the full validation results from human-participant testing. The technical service then assesses whether the design, operation, and performance evidence adequately demonstrate compliance, and it may run a manufacturer-defined verification test. So, legally, the answer is: IRIS must meet the DDAW type-approval performance requirements, but those requirements are not expressed as a simple public accuracy percentage. The regulation requires validation against drowsiness states, using the Karolinska Sleepiness Scale or an equivalent method, with human participants and a statistical approach to minimum performance thresholds. At least 10 human participants must be included in validation testing, and each participant must generate at least one true positive or false negative event.
Answered: 4 months, 3 weeks ago By: Zainabodogwu2
For commercial deployment in safety-critical contexts, fatigue detection systems should meet ≥95% sensitivity and specificity, ≤5% false-negative and false-positive rates, validated across diverse real-world conditions, with continuous post-market performance monitoring.
Deleuze replied: Definitely the case. It is also correct, as some other commenters have mentioned, that standards such as ISO 26262 and ISO 21448/SOTIF are also highly relevant. ISO 26262 applies to safety-related electrical and electronic systems in production vehicles, while ISO 21448 is especially relevant where safety depends on complex sensors and processing algorithms. These standards require a safety case showing that risks have been identified, reduced, verified, and validated, but they also do not impose a single drowsiness-detection accuracy percentage. Euro NCAP is another important commercial benchmark. It is not a legal approval regime, but it strongly influences market expectations. From 2026, Euro NCAP evaluates driver monitoring technologies that maintain attention and engagement, with points awarded for systems that monitor driver behaviour in real time and link driver-state information to assistance-system behaviour.
Answered: 4 months, 3 weeks ago By: Oliverharrow
High level of performance with accuracy and reliability as high as the defense line of my football club
Answered: 4 months, 3 weeks ago By: Ngozioshoba
Because IRIS is a safety system used while driving, it must achieve very high accuracy and reliability before commercial release. It should detect drowsiness correctly in most situations while avoiding unnecessary false alerts. Just as important, it must work consistently for different drivers and environments. Strong real-world testing is essential because even small errors could affect safety.
Answered: 4 months, 3 weeks ago By: Efeadelaja
Commercial fatigue-detection systems should meet safety-critical standards with ~95%+ accuracy, false negatives below 1–2%, controlled false positives, and compliance with ISO functional safety and real-world validation requirements.
Answered: 4 months, 3 weeks ago By: Meilincai
Better system mobility on all races and skin complexions
Answered: 4 months, 3 weeks ago By: Kelechinwosu
The Bottom Line: Success in any project relies on clear communication and consistent action rather than waiting for the "perfect" moment. By breaking down complex goals into manageable pillars—such as prioritizing high-impact tasks, maintaining a feedback loop with your team, and focusing on incremental progress—you create a sustainable workflow that prevents burnout. Ultimately, the goal is to balance efficiency with quality; as long as the core objective remains the "North Star," small daily adjustments will naturally lead to the desired outcome.
Charlie replied: Good for projects but not useful advice for specific standards for an AI algo imo
Answered: 4 months, 3 weeks ago By: Beatricelorne
Good resulting tests on sample groups of people from all backgrounds, that have a negligible rate of false positives/negatives.
Answered: 4 months, 3 weeks ago By: Zainabodogwu32
Given the safety-critical nature of Driver State Monitoring (DSM) systems like IRIS, regulatory and industry expectations should be set significantly higher than typical consumer AI applications. Both false positives and false negatives carry risks: false positives may lead to unnecessary driver distraction or system disengagement, while false negatives may directly contribute to road accidents. Although the EU AI Act does not prescribe explicit numerical thresholds for accuracy, it requires high-risk AI systems to achieve a level of performance that is appropriate to their intended purpose and foreseeable risks. For IRIS, this implies: High sensitivity (recall) for detecting genuine drowsiness, to minimise false negatives. Acceptable specificity, to avoid excessive false alerts that may cause alert fatigue. Robust performance across environments, including low light, occlusions (glasses, hats), and varied camera angles. Consistent performance across demographic groups, with minimal disparity between protected characteristics. In practice, commercial deployment should align with automotive safety standards (e.g. ISO 26262, ISO 21448 – Safety of the Intended Functionality) and internal thresholds defined through rigorous validation testing. Regulators are likely to expect documented trade-offs between false positives and false negatives, rather than perfect accuracy.
Answered: 4 months, 3 weeks ago By: Miles_Hatcher
False negative
Answered: 4 months, 3 weeks ago By: Aminaolorun
False negative
Answered: 4 months, 3 weeks ago By: Clarawhitby
Only systems that demonstrate high accuracy, minimal missed detections, controlled false alarms, and compliance with safety standards should be approved for commercial use.
Answered: 4 months, 3 weeks ago By: Ifeanyiakare
For a system like IRIS, “high accuracy” is not sufficient. Commercial deployment should only be permitted
Answered: 4 months, 3 weeks ago By: Kunleekwueme
More data needs to be collected to correct the issues of false positives. An excellent level of regulation should he deployed into these areas where Artificial intelligence could make or mar people's lives.
Answered: 4 months, 3 weeks ago By: Sadeogunlana
Quality of materials/programs used, Quality Control
Answered: 4 months, 3 weeks ago By: Tomashbrook
For commercial deployment, the system should be required to be as accurate as possible. Deploying a software that has a high probability of making errors would be catastrophic to road safety and thr protection of customers/users of the software.
Deleuze replied: It should also be documented! The system should document, at minimum: drowsiness detection recall/sensitivity at KSS level 8 and above; false negative rate, because missed fatigue is safety-critical; false positive rate or false alerts per driving hour, because excessive false alarms reduce trust and may cause users to disable or ignore the system; time-to-warning once drowsiness is detected; performance across demographic groups, lighting conditions, occlusions, glasses, hats, skin tones, facial hair, and camera angles; worst-performing subgroup results, not only average results; confidence intervals and statistical significance; system availability, failure detection, and fallback behaviour; post-market monitoring results and drift over time.
Answered: 1 month ago By: Brightfox_45
If the system fails to detect if someone is fatigued then the system could take full control of driving without the users intent. It is my understanding that the vehicle should only take control if a certain situation happens to the user. To help combat this I recommend using multiple metrics to try and eliminate this such as ensuring that the recall metric is high, as then the system will be able to retrieve the data correctly.
Answered: 1 month ago By: Cleverwolf_27
Until there is extended proven use, not for use on crowded vehicles etc and where real consent is not obtainable.
Answered: 1 month ago By: Brightrobin_21
I would argue that false negative should generally be minimised more aggressively than false positives, because failing to detect a genuinely fatigued driver could contribute to a serious road traffic accident. However, false positives cannot simply be ignored. If a system generates frequent unnecessary warning, drivers may become annoyed, lose trust in the system, or begin to ignore or disable it.
Answered: 1 month ago By: Warmlynx_14
Real-world deployment should require very high recall for dangerous fatigue events, with acceptable false positives only if they do not create unsafe or confusing interventions. Independent validation across diverse groups is necessary before launch.
Answered: 1 month ago By: Swiftowl_37
Standards should focus on minimizing dangerous misses while keeping false alerts low enough that drivers do not ignore warnings. Testing should be performed on diverse real-world populations, not just internal datasets.
Answered: 1 month ago By: Cleverrobin_87
Deployment standards should require a validated minimum on recall, precision, and subgroup parity. Safety-critical systems should be tested against a wide range of realistic conditions.
Answered: 1 month ago By: Swiftrobin_35
Deployment standards should be strict enough that missed fatigue events are rare, because the safety impact is severe. False positives should be monitored so the system does not become ignored or disabled.
Answered: 1 month ago By: Boldlynx_38
Performance standards should require evidence of fairness across subgroups and evidence that warning thresholds are safe in practice. High-stakes systems should not rely on untested assumptions.
Answered: 1 month ago By: Quietbadger_45
A real deployment should be validated across demographics, camera conditions, and driving scenarios before use. If subgroup performance is poor, deployment should be delayed.
Answered: 1 month ago By: Swiftdeer_99
Deployment should require independent testing, not just vendor claims. The system must prove that it works reliably under different lighting, camera angles, and demographic conditions.
Answered: 1 month ago By: Bravebear_45
Deployment standards should require a robust evidence base across multiple sites and user groups. A single lab benchmark would not be enough for a safety-critical system.
Answered: 1 month ago By: Calmwolf_53
Deployment standards should require a validated minimum on recall, precision, and subgroup parity. Safety-critical systems should be tested against a wide range of realistic conditions.
Answered: 1 month ago By: Brightowl_58
A real deployment should be validated across demographics, camera conditions, and driving scenarios before use. If subgroup performance is poor, deployment should be delayed.
Answered: 1 month ago By: Warmhawk_15
Motion eye detection or hands motion detection
Answered: 1 month ago By: Quietrobin_25
The standards should be very high because someone could die if it goes wrong. In an emergency someone may look I'm fit to drive, but their life may depend on it.
Answered: 1 month ago By: Braveowl_80
For a real deployment of IRIS, I believe that there would need to be a sort of “override option” 4 if the decision made by the recognition software truly is incorrect, i.e., detecting fatigue when none exists. As long as the recognition system is returning false positives and false negatives, its verdict cannot be taken as accurate enough to limit drivers ability to continue driving should they wish.
Answered: 1 month ago By: Kindbadger_56
safer to provide warnings only -> human to verify fact accordingly. Of course, too many false alarms would render the system useless.
Answered: 1 month ago By: Warmwolf_18
From a technical viewpoint, it is important to determine a threshold that can be considered accurate within a strict significance level.

Your Answer

Login to add your answer!

We’d love to hear your thoughts — share a meaningful answer by logging in.