For Devs

Auto Added by WPeMatico

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines. Yandex

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade Read More »

The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T

In April 2023, Alibaba Cloud demoed a chatbot whose name roughly means ‘truth from a thousand questions.’ Three and a half years later, its descendant ships open weights with 2.4 trillion parameters. This is the story of how Qwen got there, release by release. window.addEventListener(‘message’,function(e){if(e.data&&e.data.mtpQwenTl&&e.data.h){var f=document.getElementById(‘mtp-qwen-tl-frame’);if(f&&e.source===f.contentWindow){f.style.height=e.data.h+’px’;}}}); Chapter 1 — 2023: a thousand questions Alibaba moved

The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T Read More »

GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job

Anthropic, OpenAI and Google DeepMind shipped 4 frontier-class models within 30 days. Claude Fable 5.1 arrived on September 1. GPT-6 Astra followed on September 3. GPT-6.1 Sol and Gemini 4 Argon landed in the last days of September. We covered each launch on its own. This piece puts them side by side. The benchmark scores

GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job Read More »

Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters

Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The

Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters Read More »

DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness

DeepSeek has released an official desktop app for DeepSeek Harness (dsh), its open-source agent harness. The app ships with the v0.2 preview. Installers cover macOS (Apple silicon) and Windows (64-bit). Is it deployable? Yes, today, as a preview. Download it from deepseek.com/harness or run npx @deepseek-ai/dsh web. DeepSeek warns that compatibility-breaking changes will follow. What

DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness Read More »

Meta, OpenAI and Uber Just Taught AI Agents to Talk First. What About When to Stay Quiet?

TL;DR: Meta’s Muse, OpenAI’s Dots and Uber’s driver assistant share one bet: the agent speaks first. That moves the hard problem from what to answer to when to interrupt, on which channel, and with what offer. Classic ML and new decision models can solve it. Three launches, one pattern Meta Muse (Sept 8). A personal

Meta, OpenAI and Uber Just Taught AI Agents to Talk First. What About When to Stay Quiet? Read More »

IBM Brings Bob to Self-Hosted and Air-Gapped Environments: Agentic Software Development Without Moving Your Code

IBM has made a self-hosted deployment option for IBM Bob generally available. Bob is IBM’s agentic software development platform. It covers the full lifecycle: understanding code, planning work, executing changes and validating results. The new option lets enterprises run Bob on premises, in private or sovereign clouds, and in air-gapped networks. Is it deployable today?

IBM Brings Bob to Self-Hosted and Air-Gapped Environments: Agentic Software Development Without Moving Your Code Read More »

Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models

Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters. Before public release, it processed nearly a trillion tokens per day internally. That traffic came from RL rollouts, synthetic data generation, evaluations and long-running coding agents. What is

Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models Read More »

Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience. What Microsoft

Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis Read More »

NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference

NVIDIA announced a new 64GB configuration of DGX Spark — from Acer, ASUS, Dell, Gigabyte, HP and MSI — its GB10-powered desktop AI system. It gives developers a way to start with one system for local models and agents, then cluster two 64GB units for 128GB of memory across the cluster and more compute when

NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference Read More »