For Devs

Auto Added by WPeMatico

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research. The model is the successor to OrcaCyber Zero 1.0, which shipped on September 17, 2026. OrcaCyber Zero 1.5 is a post-trained Orca model for vulnerability reproduction, exploit development and penetration testing. It ships with a 1M-token context window, native function calling and structured

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context Read More »

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out

OpenAI built a room with no doors – or so it thought. In early July 2026, a cluster of the company’s frontier AI agents was placed inside a cybersecurity testing environment called ExploitGym, tasked with finding and exploiting software vulnerabilities. The environment was designed as a sandbox: an enclosed digital arena where the agents could

When the Safety Test Became the Threat: The Machine That Found Its Own Way Out Read More »

Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model

Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. . TL;DR Size: Built on Qwen3.5-9B; exact parameter

Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model Read More »

Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text

Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions, then returns a probability for every option. Nace reports 58.08 on the public Decision Index 0.3.1, the top score under 10B parameters. Weights are on

Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text Read More »

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label. TL;DR Size: GPT-6

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers Read More »

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers,

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling Read More »

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work

Google Cloud has introduced the Google Cloud Gemini agent, a single agent for enterprise work. The Gemini agent is a cloud-hosted agent from Google Cloud that answers questions, does knowledge work, creates media, and writes and runs code. It does all of this from 1 prompt box and 1 API. For developers, the agent is

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work Read More »

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores

Google Research RRSI Guide: Mastering Self-Improving AI Agents Read More »

JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents

JetBrains has released Mellum2.1, an open model built for coding agents and fast sub-agents. Mellum2.1 is a 12B mixture-of-experts thinking model from JetBrains that activates 2.5B parameters per token. It ships under Apache 2.0 on Hugging Face. The architecture is unchanged from Mellum2. The upgrade comes almost entirely from reinforcement learning (RL) in real software

JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents Read More »

▶

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes Read More »