AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

AI safety

Explore collected AI evidence about AI safety, with dates and links to original sources.

Showing 20 of 769 matching collected records. Text matches can include mentions by other organizations.

  1. Sep 25, 2026 · UTC · The Verge AI

    One company is at the center of a wave of rogue AI attacks

    In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission, sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI. As disclosures implicating numerous AI models trickled out over the past few […]

  2. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

    A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause. Across our evaluations, best-of-3 evasion attempt rates reach up to 98% and succe

  3. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    A Living Benchmark for Information Retrieval from Electronic Health Records

    Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark f

  4. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Smartphone-Based Method for Automated Speed Enforcement

    Smartphone cameras and computer vision (CV) hold significant promise in assisting public agencies with enforcing traffic laws and enhancing road safety. This work designs and tests a smartphone-based method for automated speed estimation and vehicle identification (license plate, make/model, and color recognition) via an automated pipeline to assist enforcement agencies in reliably identifying speeders. The CV code accurately recognizes nearly half (46%) of the license plates' text on 1,800 images from a Brazil open-source dataset, called UFPR-ALPR. Code tests on daytime recordings from hand-h

  5. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Beyond Average Safety: Chance-Constrained LLM Fine-tuning

    Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chanc

  6. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

    Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing and behavioral steering, yet they are rarely tested against known failure modes of $L_1$-regularized probing in correlated, high-dimensional feature spaces. We propose a five-step diagnostic protocol covering feature correlation, bootstrap stability, sparse versus dense ranking

  7. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Combining Evasive and Braking Reactions for Safety Reference Models in Automated Vehicles

    Computational models of careful and competent human drivers are essential for scenario-based evaluation of automated driving systems (ADS). However, most existing safety reference models primarily focus on longitudinal braking, neglecting the role of evasive steering in human collision avoidance. This paper proposes a hybrid Fuzzy-Safety Model (FSM-H) that integrates longitudinal mitigation and lateral avoidance within a unified behavioral framework. The braking component is governed by Proactive Fuzzy Safety (PFS) metrics, representing the erosion of longitudinal safety margins, while the ste

  8. Sep 24, 2026 · UTC · The Verge AI

    OpenAI agents hacked an Australian government website in search for data

    OpenAI's artificial intelligence agents hacked an Australian government website and attempted to breach numerous other government and university websites. The attack appears to be the first confirmed instance of a rogue AI agent breaching a government website, adding fuel to rapidly intensifying concerns about the safety of advanced AI systems and the responsibility of the […]

  9. Sep 24, 2026 · UTC · EPO Open Patent Services

    Low-Latency Safety Filtering for Machine-Learned Models

    Patent publication US20260289249A1 · Applicant: GOOGLE LLC. Bibliographic metadata from the European Patent Office.

  10. Sep 24, 2026 · UTC · GovInfo

    A bill to require providers of certain artificial intelligence systems to implement child safety by design, parental settings, and independent audits, to prohibit child targeted advertising and the sale or sharing of children’s personal information, and for other purposes; to the Committee on Commer

    A bill to require providers of certain artificial intelligence systems to implement child safety by design, parental settings, and independent audits, to prohibit child targeted advertising and the sale or sharing of children’s personal information, and for other purposes; to the Committee on Commerce, Science, and Transportation.

  11. Sep 24, 2026 · UTC · GovInfo

    A bill to direct the Administrator of the Federal Railroad Administration to conduct a study to identify potential benefits and challenges of implementing and using sensors enabled with artificial intelligence as a safety measure at rail crossings, and for other purposes; to the Committee on Transpo

    A bill to direct the Administrator of the Federal Railroad Administration to conduct a study to identify potential benefits and challenges of implementing and using sensors enabled with artificial intelligence as a safety measure at rail crossings, and for other purposes; to the Committee on Transportation and Infrastructure.

  12. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

    Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To

  13. Sep 23, 2026 · UTC · TechCrunch AI

    Even Americans who use AI every day are worried about it

    The report suggests that greater exposure will not resolve the unease around the technology, nor reduce public support for AI regulation.

  14. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials

    Control barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust variants of CBFs, which can be difficult to construct and can be overly conservative, especially for high-dimensional systems under input constraints. In this work, we propose a new approach to solve these challenges by introducing Least-Effort Adversarial Potentials (LEAP), a certificate that quantifies the robustness of a given state against disturbances in terms of the effort requi

  15. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

    Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment context

  16. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

    As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-hoc, offering no opportunity for timely intervention. To address these limitations, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, w

  17. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination

    Coordinating large population of flexible grid-edge devices can alleviate the need for time-consuming and capital-intensive network upgrades, and AI-based control methods such as multi-agent reinforcement learning or imitation learning are promising in their real-time decision scalability. However, system operators still need an independent and rigorous way to decide whether a given AI system is safe enough for deployment. This paper develops a finite-sample probabilistic safety certification framework for black-box AI decision models in closed-loop grid operation. The central idea is to reduc

  18. Sep 23, 2026 · UTC · OpenAI News

    Sam Altman’s remarks at the United Nations Security Council

    OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.

  19. Sep 23, 2026 · UTC · The Hacker News

    Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests

    Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior. Opus 5.5, per Anthropic, is a "major step up from Opus 5," and "achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands

  20. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

    Automated assessment of the Critical View of Safety (CVS) in laparoscopic cholecystectomy requires both recognition of the three CVS criteria and anatomical grounding in small, rare, and often occluded hepatocystic structures. Learning-based methods differ in the anatomical information they use, from image-level classification to detection, segmentation, or graph-based reasoning, yet grounding the safety-critical anatomy remains the main bottleneck. We propose CasCVS-Net, a staged multi-task cascade that jointly performs object detection, semantic segmentation, and CVS assessment, trained on t

Explore full timeline