{"system":{"name":"SupplyScore AI","version":"2.3","provider":"Acme Supplies SRL","high_risk_category":"employment_worker_management"},"assessment":{"reference":"AIRMS-2025-Q1","assessment_date":"2025-02-18","version":"2025.1","methodology":"Hybrid qualitative/quantitative assessment combining (a) structured hazard analysis (FMEA adapted for ML systems) run by the Responsible AI Committee, (b) quantitative evaluation of fairness and drift metrics on the held-out test set and on production traffic, and (c) review of post-market monitoring signals accumulated since the previous iteration. Likelihood and severity are rated on a three-point scale (low/medium/high) per risk; residual ratings are produced after the mitigation measures listed in section 6 are assumed to be operational. The record is reviewed quarterly or whenever a substantial modification under Art. 43(4) AI Act is planned.","next_review":"2025-05-18"},"intended_purpose":"SupplyScore AI produces a 0-100 onboarding-risk score with a short natural-language explanation for each prospective or existing third-party supplier of Acme Supplies SRL. The score is used as decision-support within the procurement workflow to rank suppliers and to flag decline-recommended cases for human review by a procurement lead. It is never used as the sole basis for contract termination or for any decision producing legal effects concerning natural persons.","reasonably_foreseeable_misuse":"Reasonably foreseeable misuse includes: (i) procurement users treating the score as an automated decision and approving/rejecting suppliers without performing the mandated human review; (ii) the system being applied to supplier categories outside the training distribution (e.g. sole-trader individuals rather than legal entities), where outputs would not be reliable; (iii) suppliers attempting to game known features by disclosing favourable but unrepresentative financial indicators; (iv) use of the score for purposes other than onboarding risk (e.g. price negotiation) for which the system has not been validated.","identified_risks":[{"id":"R-1","category":"discrimination","description":"Systematic under-scoring of small suppliers established in under-represented regions (Africa, Latin America, parts of South-East Asia), where the training set is thin and historical outcome labels may reflect prior biased procurement decisions rather than genuine onboarding risk.","who_affected":["Small and medium-sized suppliers in under-represented regions","Individual owners of such suppliers (indirectly, via commercial consequences)"],"likelihood":"medium","severity":"high","impact":"Unfair exclusion from procurement opportunities, reduced access to the EU single market for affected suppliers, reputational and legal exposure for the provider under Art. 9 AI Act and under EU non-discrimination law."},{"id":"R-2","category":"safety","description":"Model drift: the onboarding-risk distribution of newly onboarded supplier categories (for instance newly regulated sectors or new geographies) diverges from the training distribution, reducing the accuracy of the score on those segments.","who_affected":["Suppliers in newly onboarded segments","Procurement function of Acme Supplies SRL (relying on degraded signal)"],"likelihood":"medium","severity":"medium","impact":"Reduced decision quality, increased false-positive and false-negative rates in the affected segments, potential downstream operational and compliance exposure if undetected."},{"id":"R-3","category":"security","description":"Adversarial manipulation: a sophisticated supplier reverse-engineers which features drive the score and tailors disclosures to game known features, producing an artificially favourable score that does not reflect real onboarding risk.","who_affected":["Procurement function of Acme Supplies SRL","Honest competing suppliers"],"likelihood":"low","severity":"medium","impact":"Incorrect prioritisation of high-risk suppliers, possible onboarding of suppliers that later fail or expose the buyer to sanctions or compliance issues."},{"id":"R-4","category":"other","description":"Training-data staleness: as the economic and regulatory environment evolves, risk signals learned from 2018-2024 data become gradually less predictive, degrading overall model performance if retraining does not occur on a sufficient cadence.","who_affected":["All suppliers assessed by the system","Procurement function of Acme Supplies SRL"],"likelihood":"medium","severity":"medium","impact":"Silent degradation of the score's predictive value across the entire supplier base, leading to progressively noisier decision support."}],"mitigations":[{"risk_id":"R-1","measures":["Fairness constraint in the training loss: demographic-parity difference across supplier-country clusters bounded to <= 0.05 (introduced in v2.3)","Quarterly demographic-parity audit across four supplier-country clusters and two firm-size clusters, with published summary report","Mandatory human review of every decline-recommended score (<= 40), with override-with-reason logged for retraining feedback","Country-of-establishment removed as a direct feature; retained only as a fairness-monitoring slice"],"residual_likelihood":"low","residual_severity":"medium","owner":"Responsible AI Committee (chair: AI Governance Officer)"},{"risk_id":"R-2","measures":["Monthly drift detector running Population Stability Index on each feature and per-segment AUROC on a rolling 90-day evaluation set","Automatic retraining trigger when PSI exceeds 0.2 on any feature or per-segment AUROC drops by more than 0.05 from the release baseline","Segment-level confidence flags surfaced to procurement users when a prospective supplier falls in a low-coverage segment"],"residual_likelihood":"low","residual_severity":"medium","owner":"ML Platform team lead"},{"risk_id":"R-3","measures":["Monitoring of disclosed-feature distributions for suspicious patterns consistent with gaming (e.g. boundary-hugging values on monotonicity-constrained features)","Mandatory human review of high-score edge cases (score >= 85 with short commercial history) by a procurement lead before onboarding","SHAP-based explanation displayed alongside every score so that procurement leads can detect implausible feature contributions"],"residual_likelihood":"low","residual_severity":"low","owner":"Procurement Director"},{"risk_id":"R-4","measures":["Rolling six-month retraining cadence on the latest closed 12-month outcome window","Backtesting of each candidate release against the previous-release production traffic before promotion","Version-history documentation under Art. 43(4) AI Act for every substantial modification"],"residual_likelihood":"low","residual_severity":"low","owner":"ML Platform team lead"}],"testing_evidence":[{"test_name":"Demographic-parity audit on 2024-H2 test set (EU vs non-EU clusters)","result":"Demographic-parity difference 0.037 (target <= 0.05) — passes","date":"2025-02-05"},{"test_name":"Drift-detector dry run on 2024-Q4 production traffic","result":"Maximum per-feature PSI 0.11 (threshold 0.20) — no drift-triggered retrain needed","date":"2025-02-10"},{"test_name":"Adversarial feature-dropout robustness test (up to 20% dropout)","result":"AUROC degradation 0.03 at 20% dropout (acceptable per release criterion of <= 0.05)","date":"2025-02-12"}],"post_market_feedback":[{"source":"user_report","summary":"Procurement lead in the Brussels office reported a pattern of false-positive declines affecting small suppliers in two non-EU countries. Investigation confirmed the pattern and directly motivated the v2.3 fairness constraint.","date":"2025-01-14","triggered_update":true},{"source":"drift_monitoring","summary":"Automatic drift alert on feature `declared_headcount_bucket` (PSI 0.22) following onboarding of a batch of logistics suppliers in a new NACE sub-sector. Handled by a targeted retrain on the 2024-Q4 data; no substantial modification under Art. 43(4) required.","date":"2025-01-28","triggered_update":true},{"source":"fairness_audit","summary":"Quarterly fairness audit (Q4 2024) across supplier-country and firm-size clusters: all clusters within the 0.05 demographic-parity bound. No corrective action required.","date":"2025-02-03","triggered_update":false}],"conclusion":{"decision":"proceed_with_conditions","rationale":"Residual risks for R-2, R-3 and R-4 are assessed as low and adequately controlled by the measures in section 6. Residual risk for R-1 (discrimination against small suppliers from under-represented regions) remains medium in severity despite mitigation; placing the system on the market is therefore permitted only under the condition that the quarterly fairness audit is performed and its summary report is published to the Responsible AI Committee. Any failure of the demographic-parity bound triggers suspension of the system under the incident process referenced in the Annex IV technical documentation."},"approval":{"signatory_name":"Marie Laurent","signatory_title":"AI Governance Officer, Acme Supplies SRL","date":"2025-02-20"}}