agentic ai

Auto Added by WPeMatico

Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model

Alibaba’s Qwen team has released Qwen-Image-2.1-Turbo, an accelerated checkpoint of its open-weight Qwen-Image-2.1 model. It generates and edits images in 8 denoising steps instead of the base model’s 40-step default. For developers, that means 5x fewer denoising steps on the same 7B architecture, plus a hosted API option. TL;DR Size: 7B parameters in the visual

Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model Read More »

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label. TL;DR Size: GPT-6

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers Read More »

Desktop AI: Why the Hardware Built for Tomorrow Is Finally Finding Its Real-World Job Today

For those of us who have spent decades building, modifying, and analyzing personal computers—going all the way back to my early days at IBM in the 1980s—watching the tech industry try to invent a new hardware category out of thin […] The post Desktop AI: Why the Hardware Built for Tomorrow Is Finally Finding Its

Desktop AI: Why the Hardware Built for Tomorrow Is Finally Finding Its Real-World Job Today Read More »

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers,

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling Read More »

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work

Google Cloud has introduced the Google Cloud Gemini agent, a single agent for enterprise work. The Gemini agent is a cloud-hosted agent from Google Cloud that answers questions, does knowledge work, creates media, and writes and runs code. It does all of this from 1 prompt box and 1 API. For developers, the agent is

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work Read More »

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores

Google Research RRSI Guide: Mastering Self-Improving AI Agents Read More »

JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents

JetBrains has released Mellum2.1, an open model built for coding agents and fast sub-agents. Mellum2.1 is a 12B mixture-of-experts thinking model from JetBrains that activates 2.5B parameters per token. It ships under Apache 2.0 on Hugging Face. The architecture is unchanged from Mellum2. The upgrade comes almost entirely from reinforcement learning (RL) in real software

JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents Read More »

Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference

Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL,

Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference Read More »

▶

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes Read More »

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality. Both models retrieve text, images and rendered PDF pages, and they share one embedding space. Is it deployable? Yes, if you host it yourself. Both models are on Hugging

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA Read More »