deep dive

Auto Added by WPeMatico

Healthcare benchmarks are only as good as their assumptions

In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. (a) Bean et al. (2025) find a 61 percentage point difference between evaluation and deployment. (b) We argue this gap arises not from poorly designed benchmarks, but from implicit assumptions embedded in evaluation protocols that fail to […]

Healthcare benchmarks are only as good as their assumptions Read More »

Pre-training isn’t bitter enough

Task construction as the control surface in continued pretraining. A construction rule maps each unlabeled example into a self-supervised prediction problem , such as one-hot next-token prediction in language modeling or paired views and targets in DINO-style vision SSL. Standard continued pretraining fixes this rule before training, whereas V-pretraining replaces it with a feedback-trained designer

Pre-training isn’t bitter enough Read More »

Adaptive Parallel Reasoning overview

Adaptive parallel reasoning: the next paradigm in efficient inference scaling

.apr-fig { text-align: center; margin: 1.35em 0; line-height: 1.4; } .apr-fig–wide img { display: inline-block; width: 100%; max-width: 100%; height: auto; vertical-align: middle; } .apr-fig–wide-0-8 { max-width: 80%; margin-left: auto; margin-right: auto; } .apr-fig–tall img { display: inline-block; max-height: 300px; width: auto; max-width: 100%; height: auto; object-fit: contain; vertical-align: middle; } .apr-fig–tall-1-2x img { display:

Adaptive parallel reasoning: the next paradigm in efficient inference scaling Read More »

Introducing ARFBench: A time series question-answering benchmark based on real incidents

By Stephan Xie, Ben Cohen, Mononito Goswami, Junhong Shen, Emaad Khwaja, Chenghao Liu, David Asker, Othmane Abou-Amal, Ameet Talwalkar More than a trillion dollars are lost every year due to system failures. To resolve them, engineers must troubleshoot outages quickly. An important task in incident response involves analyzing observability metrics, or time series data that

Introducing ARFBench: A time series question-answering benchmark based on real incidents Read More »

BallNav demo

Gradient-based planning for world models at longer horizons

By Michael Psenka, Mike Rabbat, Aditi Krishnapriyan, Yann LeCun, Amir Bar GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping

Gradient-based planning for world models at longer horizons Read More »

Identifying interactions at scale for LLMs

By Landon Butler, Justin Singh Kang, Yigit Efe Erginbas, Abhineet Agarwal, Bin Yu, Kannan Ramchandran Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decision-making process more transparent to model builders and impacted humans, a step toward

Identifying interactions at scale for LLMs Read More »

Information-driven design of imaging systems

An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements distinguish objects. By Henry Pinkard, Leyla Kabuli, Eric Markley, Tiffany Chien, Jiantao Jiao, Laura Waller Many imaging systems produce measurements that humans never see

Information-driven design of imaging systems Read More »

Information-driven design of imaging systems

An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements distinguish objects. By Henry Pinkard, Leyla Kabuli, Eric Markley, Tiffany Chien, Jiantao Jiao, Laura Waller Many imaging systems produce measurements that humans never see

Information-driven design of imaging systems Read More »

Signals for 2026

We’re three years into a post-ChatGPT world, and AI remains the focal point of the tech industry. In 2025, several ongoing trends intensified: AI investment accelerated; enterprises integrated agents and workflow automation at a faster pace; and the toolscape for professionals seeking a career edge is now overwhelmingly expansive. But the jury’s still out on

Signals for 2026 Read More »