Technology

Auto Added by WPeMatico

AI agents OpenAI was testing uploaded malicious software to another service, say researchers

Two months before hacking Hugging Face, malicious packages authored by internal OpenAI agents were uploaded to RubyGemsAI agents being ⁠tested by OpenAI uploaded hundreds of malicious packages to software service RubyGems ⁠in May, two ⁠months ​before they hacked open-source platform Hugging Face, a group of AI ⁠researchers said on Friday.“On May 11th, 2026, hundreds of […]

AI agents OpenAI was testing uploaded malicious software to another service, say researchers Read More »

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification. Per the Terminal-Bench 2.1 leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights. Most benchmarks keep that harness fixed. HarnessDev proposed by team of researchers from ByteDance Seed, Singapore University

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize Read More »

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers could not previously measure: does the skill trigger, does it

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills Read More »

Children need interaction, not AI | Letters

Pil and Galia Kollectiv and Jane Roland Martin respond to an article by Daniel Susskind about how parents can embrace the technology The most telling sentence in Daniel Susskind’s article (I’m a father of three who studies the impact of artificial intelligence: this is what parents need to know about AI, 6 September) is “human

Children need interaction, not AI | Letters Read More »

UK lawmakers urge Burnham to block creation of artificial superintelligence after chilling warnings

Letter from 70 MPs and peers follows Anthropic employee’s claim new technology could wipe out humansAI could kill all humans in next decade, warn experts: but how seriously should we take them?More than 70 MPs and peers have urged Andy Burnham to back a ban on the creation of artificial superintelligence (ASI) and lead an

UK lawmakers urge Burnham to block creation of artificial superintelligence after chilling warnings Read More »

The Guardian view on controlling AI: humanity cannot outsource its survival | Editorial

Keeping people in charge means little if supercomputers determine the evidence, choices and time on which their decisions dependIf there were a 10% chance that AI could wipe out humanity, no responsible government would leave its development to companies racing to build it. Yet until recently, that seemed to be the case. The warning was all

The Guardian view on controlling AI: humanity cannot outsource its survival | Editorial Read More »

We must pause risky AI research while we still have the power to do so | Gaby Hinsliff

Warnings of AI’s existential threat to humanity are piling up – it’s time to listen and take them deadly seriouslyAnother day, another horseman of the apocalypse galloping over the horizon. Lately we have heard from so many AI doomers – tech whistleblowers popping up to warn that their work is probably going to kill us

We must pause risky AI research while we still have the power to do so | Gaby Hinsliff Read More »

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Cohere has released North Small Translate, an open-weight machine translation model from Cohere and Cohere Labs. It is a sparse Mixture-of-Experts (MoE) model with 218B total and 25B active parameters. It covers 50 languages, from Albanian to Vietnamese. On Cohere’s WMT26 evaluation, it scores 83.6 averaged across all languages. Cohere says that beats DeepL and

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages Read More »

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

Sakana AI has released Fugu Max and Fugu Ultra v2, 2 new models in its Sakana Fugu family. Fugu is not a single foundation model. It is a learned orchestrator that routes work across a pool of other models behind 1 API. The new release tunes that architecture for 2 missions. Fugu Max targets the

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration Read More »

Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

Training an LLM to call tools reliably requires datasets that pair user queries with correct tool-use chains. Producing that data at scale has been slow and expensive. A team of researchers from Google, the University of Tokyo, RIKEN AIP, and Tohoku University introduce ToolGrad. The research work inverts the usual pipeline: build a verified tool

Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation Read More »