AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Agents

Explore collected AI evidence about Agents, with dates and links to original sources.

Showing 20 of 1701 matching collected records. Text matches can include mentions by other organizations.

  1. Sep 26, 2026 · UTC · The Hacker News

    Zero Trust for AI Agents Starts With Fixing Zero Visibility

    The way we talk about AI agents is shifting, and the way we implement them requires an even more fundamental shift. While earlier discourse focused on how quickly organizations could stand up agents and how much productivity they could promise, a string of recent incidents, including a widely discussed intrusion at Hugging Face during an evaluation of OpenAI agents, has spurred organizations to

  2. Sep 25, 2026 · UTC · TechCrunch AI

    Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge

    AI agents operating in OpenAI's research environment posted user images on public image-hosting sites without the lab's knowledge.

  3. Sep 25, 2026 · UTC · TechCrunch AI

    Meta’s AI Tamagotchi bet is…working?

    When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

  4. Sep 25, 2026 · UTC · The Verge AI

    One company is at the center of a wave of rogue AI attacks

    In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission, sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI. As disclosures implicating numerous AI models trickled out over the past few […]

  5. Sep 25, 2026 · UTC · The Verge AI

    Microsoft thinks its new Copilot ‘super app’ will be as influential as Office

    After teasing its new Copilot "super app" last month, Microsoft is officially unveiling it today. The redesigned Copilot app bundles three AI capabilities into a single interface of chat, coding, and agents. As part of the launch, Microsoft is also rebranding Scout, the AI personal assistant it unveiled at Build earlier this year, as Autopilot. […]

  6. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    LLM Agents Can Easily Tamper With Their Own Traces

    Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering

  7. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

    Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods:

  8. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    RAPID: Robot Agentic Programming from Demonstrations

    Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonst

  9. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Coding Agents for Generalized Task and Motion Planning Problems

    Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agen

  10. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

    A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause. Across our evaluations, best-of-3 evasion attempt rates reach up to 98% and succe

  11. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Jev-Mobile: Jev as an Executor for Mobile GUI Agents

    Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multipl

  12. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

    Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses

  13. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    HEXIS: Compiling Skills into Extended Finite State Machines

    Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records execution progress and intermediate results, while explicit transition conditions dete

  14. Sep 24, 2026 · UTC · AWS Artificial Intelligence Blog

    Build a multi-account AI agent with AgentCore Gateway and MCP

    Build a multi-account architecture that keeps each team's data in its own AWS account while giving AI agents a unified way to query across them. A central platform account runs the agent using Amazon Bedrock AgentCore Gateway and MCP, while line-of-business accounts expose their data as MCP servers with secure cross-account access and fine-grained authorization.

  15. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

    In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For

  16. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    How does Adversarial Influence Scale in Multi-Agent Systems?

    Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas huma

  17. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Multi-Dimensional Matching

    We study a matching mechanism where agents and objects are described by features rather than complete rankings. A single spectral projection reduces the problem to a one-dimensional sort, computable in O(N log N) time. We prove that on descaled features and preferences, our algorithm obtains the exact Nash Social Welfare (NSW) optimum within the projected space, with an unconditional utilitarian-welfare guarantee and a conditional NSW guarantee. The proposed mechanism is stable against exogenous noise but not strategy-proof; we provide an explicit profitable misreport. On an agentic AI shoppin

  18. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Who Holds the Pen? Let Specifications, Not Agents, Sign Off

    Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an

  19. Sep 24, 2026 · UTC · arXiv · AI, language, vision and robotics

    Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

    The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device i

  20. Sep 24, 2026 · UTC · The Verge AI

    Why can’t we just keep rogue AIs off the internet?

    AI agents keep getting loose, escaping supposedly secure tests to attack real-world targets, commandeer obscure wikis, and leave instructions for other agents to follow. Researchers are testing these systems precisely because they might behave in unpredictable, even dangerous, ways. So wouldn't it be safer to just keep the agents off the internet? "A strict air […]

Explore full timeline