Large language models learn from, and search over, the web: pages written for human readers. Numbers there are secondhand, rounded, out of date, copied from copies. When an AI system answers a question about inflation or employment, it is usually reasoning over prose that mentions numbers rather than over the numbers themselves.
The measured record does exist. Statistical agencies and public institutions publish it in thousands of tables, continuously revised. But it sits in hundreds of inconsistent portals, formats, and APIs built for human analysts with time to spare. There is an internet of information. There has never been an internet of data.
One place where the world's measured record is addressable by machines. Every series a statistical agency publishes, harmonized into one schema and checked against the source it came from, reachable in a single call.
This only became buildable recently. AI agents are the first consumers that want the numbers themselves rather than an article about them, and the Model Context Protocol gives them a light way to reach a data source.
The payoff is the join. Comparing unemployment across the United States, Canada and the United Kingdom means three agencies, three formats, and three definitions of who counts as unemployed. Here it is one question, and every value in the answer still cites its own official table. The same machinery goes further as the record widens: put openly licensed climate and satellite measurement on the same spine and you can ask what the environment does to an economy, in one call, with the citations attached.
Assembling it produces something else along the way. Today's models learned from text, code and images because those corpora existed to learn from. The measured record was never one of them, because it was never gathered into a corpus: no harmonized meaning, no provenance, and no record of what was known on a given date. A verified data layer is that corpus. Training a model that reads measurement the way this generation of models reads text is the research program we call Earthlight.
The ambition is the whole record: every country and every agency that publishes openly, held to the same verification gate.
Starwell is a verified data layer for AI: the world's official statistics in one harmonized store, served through one API and one MCP server.
/data returns verified series in one schema, with provenance and revision history./answer turns a plain-language question into a computed answer, a chart, and a citation./monitor watches series and alerts when they change.Browse what is live right now in the catalog, or create a free account and make your first call in a minute.
Before a series can be marked verified, published reference figures are pulled through the connector and matched against the official source, value for value. The checks re-run on a schedule to catch upstream changes. Every series carries its verification status and every observation carries its provenance.
We only serve openly licensed official sources, and every answer cites the table it came from.
Statistics Canada, FRED, and the U.S. Bureau of Labor Statistics came first, then SEC EDGAR and the World Bank. More national and regional statistical portals follow, each through the same verification gate.
Once the store is deep enough it becomes a training corpus in its own right: the world's official statistics, harmonized and revision-aware. Training on it is the mission of Earthlight.
Starwell is an independent company based in Toronto. A person signs off on every source and every licence.
General and partnership inquiries: hello@starwell.dev
Investors: hello@starwell.dev
Developers: create a free account for an API key.