When cities build new subway lines, they face an impossible dilemma long before construction begins: how do you identify the cheapest, safest path through hundreds of city blocks without spending years on costly field surveys? Urban planners typically assume that the only way to make an optimal choice is to collect as much data as possible — often far more than budgets, timelines, or logistics allow.
A new study from researchers at the Massachusetts Institute of Technology (MIT) challenges that assumption with a radically different idea: for many complex decisions, you don’t need more data — you just need the right data.
The team has developed a mathematical and algorithmic framework that can identify the smallest possible dataset required to guarantee an optimal decision, even in problems involving thousands of uncertainties. Their findings suggest that a city planning a subway line under Manhattan, or a utility operator optimizing an electricity grid, may be able to cut their data collection needs dramatically.
And the breakthrough doesn’t just reduce the burden of field surveys — it rewrites long-held beliefs about AI and the data economy.
“Data are one of the most important aspects of the AI economy. Models are trained on more and more data, consuming enormous computational resources. But most real-world problems have structure that can be exploited. We’ve shown that with careful selection, you can guarantee optimal solutions with a small dataset,”Asu Ozdaglar, head of MIT’s EECS department, in a statement issued as part of this research.
Rethinking the ‘Big Data’ Era
For over a decade, modern AI has pushed the narrative that “more data is always better.” But many real-world optimization problems — from supply chains and transit networks to energy markets — have predictable structural patterns. The MIT researchers argue that this structure can be used to determine exactly which data points matter.
This is crucial in systems like:subway route selection, supply chain diversification, electricity network optimization, construction planning, logistics and resource allocation.
In all these cases, practitioners often drown in unnecessary data collection, hoping volume will compensate for uncertainty.
The new algorithm takes the opposite approach: start with no data, and add only what is provably essential.
The Key Breakthrough
The team’s method begins by mathematically defining what it means for a dataset to be “sufficient.” They break the decision space into “optimality regions” — scenarios in which a particular route, price, or configuration becomes the best choice.
A dataset is sufficient if it can accurately determine which region the real world belongs to. “When we say a dataset is sufficient, we mean that it contains exactly the information needed to solve the problem. You don’t need to estimate all parameters accurately; you just need data that can discriminate between competing optimal solutions,” said Amine Bennouna, co-lead author.
Once the structure is defined, the algorithm repeatedly asks a critical question: “Is there any scenario in which the optimal decision could change, and my current data would fail to detect it?”
If the answer is yes, the algorithm identifies precisely which new measurement would fill that gap. If no, the dataset is complete — and provably sufficient.
This approach can shrink data collection from thousands of measurements to a handful.
“The algorithm guarantees that, for whatever scenario could occur within your uncertainty, you’ll identify the best decision,” Omar Bennouna, co-lead author, added.
From Manhattan Tunnels to Global Supply Chains
Consider the example of a subway planned beneath New York City. Every city block might hide different soil conditions, underground utilities, or hazard profiles. The traditional assumption: investigate everything.
The MIT model: investigate only the blocks that can change the optimal route, and ignore the rest.
The same applies to: selecting shipping routes in a congested supply chain, configuring electrical grid nodes during volatile energy prices, deciding where to place sensors in large infrastructure projects. This could save millions in surveys, simulations, and engineering assessments.
The researchers were able to show not only that minimal datasets exist — but that their algorithm can find them systematically, preserving optimality with mathematical certainty.
“We challenge this misconception that small data means approximate solutions. These are exact sufficiency results with mathematical proofs,” Saurabh Amin, co-senior author, said.
Why This Matters: Cost, Carbon, and Computational Savings
If AI models can be trained with fewer, smarter data points: computation becomes cheaper, energy consumption drops, model training becomes faster and policymakers can rely on faster decision cycles.
For large infrastructure projects, this could mean shaving months — even years — off planning timelines.
The team plans to expand the framework to more complex scenarios, including situations where data are noisy or partially observable — a common challenge in the real world.
Their work will be presented at the Conference on Neural Information Processing Systems (NeurIPS) — a major stage for breakthroughs that redefine how AI systems learn and make decisions.
If successful, this research could push the global AI ecosystem to rethink its data obsession. Instead of hoarding massive datasets, the future may lie in asking sharper questions — and collecting only the data that truly matters.