For Devs

Auto Added by WPeMatico

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality. Both models retrieve text, images and rendered PDF pages, and they share one embedding space. Is it deployable? Yes, if you host it yourself. Both models are on Hugging

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA Read More »

What Happens When a Trusted Model Repo Changes? Unsloth Studio Re-Checks Before It Runs

Beta Lessons Learned After over 500 million downloads, years of requests from the open-source community, and being a top product on Hugging Face, Unsloth launched their beta desktop App, Unsloth Studio. Unsloth which makes it faster, easier, and more affordable to fine-tune and run AI models, including locally on your own hardware. The app centralizes

What Happens When a Trusted Model Repo Changes? Unsloth Studio Re-Checks Before It Runs Read More »

Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens

Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model to date. It targets high-volume work like summaries, compaction, classification and subagent tasks. It keeps a 1M token context window and up to 128K output tokens. Pricing starts at $0.10 per million input tokens and $0.50 per million output tokens. That is 90%

Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens Read More »

Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens

Liquid AI has released Open d1, two open-weight multimodal models in its d1 decision model family. d1-3B reads text and images. d1-omni-600M reads text with an image, or text with audio. Neither model writes text. Each returns calibrated, typed answers in one forward pass with zero output tokens. The target is real-time decisions on the

Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens Read More »

↔

Meta AI Open-Sources Rebalancer: A C++ Assignment Solver That Runs About 40 Million Placement Problems a Day

Meta has open-sourced Rebalancer, a C++ library with a Python interface for solving assignment problems. It decides which objects go into which bins under constraints and objectives. According to the Engineering at Meta’s post, Rebalancer has handled resource allocation across Meta for over 9 years. The release ships under Apache 2.0 with documentation, a PyPI

Meta AI Open-Sources Rebalancer: A C++ Assignment Solver That Runs About 40 Million Placement Problems a Day Read More »

Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4

Google DeepMind has released EmbeddingGemma 2, an open model that embeds text, code, images, video and audio into one 768-dimensional space. It has 740M parameters, an 8K token context window and an Apache 2.0 license. It targets on-device search, classification and privacy-first RAG. This article analyzes, compares and showcase how EmbeddingGemma 2 fits in the

Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4 Read More »

Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model

Mistral AI has just announced the release of Mistral Large 4 (ML4), internally nicknamed Le Chonk, as a public preview. ML4 is a granular Mixture of Experts model with 1.05 trillion total parameters, 49 billion active per token, a 1.6 billion parameter vision encoder, and a 1 million token context window, per the model documentation.

Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model Read More »

Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode

Together AI has released Together Link, a free, MIT-licensed CLI now in beta. It connects the coding agents developers already use to open models hosted on Together AI. Supported tools include Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi. The idea is simple: keep the harness, swap the model, and shrink the bill.

Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode Read More »

Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads

Reflection AI has introduced Beam, its first open-weight model. Beam is a sparse Mixture-of-Experts (MoE) model with 501B total parameters and 23B active per token, built for coding, reasoning and agentic workloads. As per the Reflection AI team, Beam directly competes with larger open models like GLM 5.2 while using 3 to 4x less inference

Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads Read More »

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines. Yandex

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade Read More »