There's a curious problem emerging as artificial intelligence reaches beyond screens and servers. These systems—increasingly called "agents" by the people building them—are supposed to make decisions about the physical world. Where to site a factory. Whether to insure a property. How to route a delivery truck through wildfire country. And they're running into a data problem that's only now becoming impossible to ignore.
The agents need to know things. Not vague, probabilistic things scraped from the internet, but concrete, defensible facts about terrain, flood zones, power lines, property boundaries. The kind of information that can hold up when an insurance commissioner asks where a premium calculation came from, or when a utility regulator wants to see the data behind a site rejection.
Goldman Sachs has forecast roughly $7.6 trillion in global AI infrastructure investment between 2026 and 2031, spanning compute, data centers, and power—much of it aimed at factories, utilities, and industrial facilities. NVIDIA, meanwhile, spent the first half of this year rolling out what CEO Jensen Huang called "Physical AI" toolkits—software explicitly designed for robots and industrial workflows. In March 2026, Huang spoke about an "agent inflection point." The message from the industry is consistent, if a bit breathless: AI is going operational.
But operational systems need operational data. And that's where things get interesting.
A Two-Person Bet on Provenance
Mireye is barely a company. Two founders—Ansh Chokshi and Shashwat Kapoor—working out of Y Combinator's summer cohort this year. What they're building, though, cuts to the heart of the problem. Feed their system a U.S. latitude and longitude, and you get back structured geospatial information: terrain type, land cover, nearby utilities, climate hazard scores, building footprints. More importantly, perhaps, you get a source URL for every single field, a timestamp showing when the data was fetched, and a confidence score.
Chokshi talks about wanting to "index every inch of the earth" and make it as queryable as the web. For now, the scope is narrower—U.S. coverage only, federal sources exclusively. Think USGS, NOAA, FEMA, EPA. The field count sits somewhere between 175 and 200, depending on when you check their catalog. They're giving away access for free during this early phase, with pricing details still to come.
The technical architecture is deliberately minimal. Two HTTP endpoints, one that answers natural-language questions and another that returns structured fields with full provenance. There's also a Model Context Protocol server—more on that shortly—that plugs into Claude Desktop and a few other tools. The whole thing runs on what Mireye describes as a planner-fetch-synthesizer pipeline using Claude Sonnet 4.6.
It's a small operation building infrastructure for what might be a very large moment.
The Protocol That Changed the Conversation
To understand why this matters now, you need to know about the Model Context Protocol. Anthropic introduced it in November 2024. By early this year, OpenAI and Microsoft had adopted it. Academic papers from earlier this year cite upward of 10,000 active MCP servers and something like 97 million monthly SDK downloads, though those numbers should be taken as directional until official registry data becomes available.
What MCP does, in essence, is let AI agents call external tools and data sources without requiring custom integration work for every combination of agent framework and data provider. For geospatial data, this turned out to be significant. A wave of geospatial MCP servers appeared this year. GIS-MCP bridges Python spatial libraries to agents. Others handle geocoding, routing, OpenStreetMap queries. Felt built a server for handling WMS and ArcGIS feeds. TomTom and Mapbox servers started showing up in agent directories.
Esri—the 800-pound gorilla of enterprise GIS—rolled out MCP support in beta during the second quarter. Their messaging was careful, positioning agents as complements to traditional workflows rather than replacements. Which tells you something about how established players are thinking about this shift.
The ecosystem is fragmenting, as ecosystems do. Some servers prioritize breadth and real-time data. Others emphasize open-source flexibility. Mireye's wager is on provenance and federal data integrity—betting that for certain use cases, knowing where your data came from matters more than having the most data.
When the Source Matters More Than the Answer
Consider insurance underwriting. The National Association of Insurance Commissioners issued a brief in March 2026 outlining expectations for AI use in pricing and claims decisions. The message was straightforward: existing laws apply, data sources must be documented, model outputs must be tested, third-party tools must be governed.
Moody's announced its acquisition of CAPE Analytics on January 13, 2025, consolidating AI-powered property risk assessment into one of the industry's leading model providers. ZestyAI has secured multi-state approvals for wildfire models. Floodbase launched an instant parametric flood quoting API in March, promising near-instant risk assessment.
The common thread isn't just technical capability—it's defensibility. When an insurer denies coverage or sets a premium based on geospatial risk factors, regulators and policyholders want answers. Where did that data come from? How recent is it? What methodology determined that flood score? An agent that can cite FEMA's National Flood Hazard Layer, properly timestamped and versioned, carries a different liability profile than one pulling from an undocumented third-party source.
Or take energy infrastructure siting. Lawrence Berkeley National Lab's Electricity Markets & Policy group published its "Queued Up" report in early July, showing over 2,000 gigawatts of generation and storage capacity seeking interconnection at the end of last year. Wait times remain brutal: roughly eight years in PJM, more than nine in CAISO, four in MISO. For developers screening thousands of potential sites, the ability to quickly check terrain constraints, utility proximity, and climate hazards with citable federal sources can compress months of preliminary work into hours.
The friction has intensified around data center siting, too. Tom's Hardware reported in June that two-thirds of 809 planned U.S. AI data centers are headed for drought zones. More than 75 data center projects worth approximately $130 billion got blocked in the first four months of this year, according to industry reporting—driven largely by local opposition over power and water demands. Google and Microsoft have made public commitments to water-efficient designs, but the scrutiny isn't letting up. Site selection agents that can surface water-stress indicators and power grid capacity with federal data provenance aren't solving a hypothetical problem.
Real-World Use, Faster Than Expected

The applications are moving from concept to production quicker than many anticipated. GrowthFactor published a walkthrough in mid-June showing how retail site selection agents can chain geocoding, trade area analysis, and cannibalization scoring through MCP. A February paper on arXiv described AURA, a reinforcement learning agent that reportedly cut affordable housing site selection time in New York City from 18 months to 72 hours in a case study.
Floodbase announced its collaboration with Liberty Mutual on February 24, 2026, enabling parametric flood quoting during underwriting. Industry case studies from consulting firms WNS and Datamatics—both published this year—detail commercial underwriting projects where agents handle document review and initial risk assessment, though these should be read with the usual vendor-report skepticism.
McKinsey's June analysis on what it calls "thinking machines" argues that robotics and physical AI will create value first in manufacturing and logistics. Field-deployed autonomous agents, the firm suggests, will face safety, regulatory, and infrastructure constraints through at least 2030. The World Economic Forum and Deloitte, in a June report, estimated that Earth observation contributes around $440 billion to global GDP, with roughly $263 billion per year in unrealized value. The constraint isn't the technology—it's integration and access.
Esri devoted considerable space in its summer ArcNews issue to "agents and agentic workflows," framing MCP as a bridge that lets agents query ArcGIS capabilities without displacing the underlying stack. It's revealing positioning, perhaps acknowledging that enterprise spatial workflows are too embedded in organizational infrastructure to be replaced wholesale. The agents augment and automate. The core systems stay put.
Regulatory Pressure, Rising
All of this is happening as the regulatory environment tightens. The EU's AI Act received final Council approval on June 29. High-risk AI obligations—which may cover certain spatial analytics and autonomy applications—begin phasing in from August. Draft guidance was still under public consultation mid-year, so specifics remain fluid, but the direction is unmistakable.
In the U.S., California's Privacy Protection Agency launched its "Delete Act" on January 1, requiring data brokers to honor statewide deletion requests. The state's consumer privacy law treats precise geolocation as sensitive personal information. In May, the FTC banned Kochava and a subsidiary from selling sensitive location data without affirmative consent, signaling enforcement focus on geospatial brokers.
The FAA's Part 107 and Remote ID requirements for commercial drones, updated through June, define the regulatory boundary for field agents involving unmanned aircraft. The U.K. adopted stricter Remote ID standards early this year. Any agent workflow that coordinates drone operations must meet certification rules.
These overlapping frameworks create demand for defensive data architecture—systems where every value carries a lineage, a timestamp, a source. The OECD published a blog post in February warning that expanding Earth observation access brings privacy and security risks, particularly as foundation models lower the barrier to analyzing satellite imagery. Regulators are paying attention. Compliance burdens are more likely to increase than ease.
What Happens Next

Grand View Research valued the geospatial analytics market at $102.7 billion last year, projecting growth to $234 billion by 2033. Market-sizing methodologies vary, so treat those figures as directional. But the trajectory is consistent across sources.
Agent frameworks will likely standardize provenance and audit trails over the next few years. Enterprise buyers in regulated industries already demand governance and traceability. NVIDIA's toolkit releases this year included explicit governance messaging. Esri's summer materials stressed that agents are complements, with MCP as the integration layer.
For physical AI, value will accrue first where constraints are clearest. Manufacturing and logistics have defined spaces and measurable ROI. Field-deployed agents—drones inspecting infrastructure, robots navigating outdoor environments—face longer timelines. Safety, regulatory, and infrastructure dependencies are harder to resolve. Siting friction for energy projects will likely persist through 2027 or beyond, maintaining demand for fast, sourced geospatial screens that can defensibly rule out bad locations before expensive studies begin.
Mireye is early. Two people, early-access product, no public pricing or funding announcements. Their catalog is evolving; the field count shifted during this year's documentation. But the architectural choice—federal sources, per-field provenance, U.S.-only coverage, agent-native delivery—maps onto problems becoming more acute as AI moves from simulation to operation.
Whether Mireye succeeds or someone else's approach wins, the infrastructure gap is real. AI agents can't function in the physical world on embeddings and vibes alone. They need terrain data, hazard layers, infrastructure proximity, regulatory boundaries. All of it timestamped, sourced, defensible.
Somebody has to build that layer. Mireye represents one attempt at an answer—small, focused, perhaps premature. Or perhaps right on time.
