空间数据基础设施(SDI)长期存在运维成本高、使用难度大的问题。云原生格式、人工智能(AI)与数据主权理念正改变这一现状。本文介绍 Portolan 与 CARTO SDI。
Spatial data infrastructure has been costly to run and hard to use. Cloud-native formats, AI and sovereignty change that. Introducing Portolan and CARTO SDI.
用户体验也未见明显改善。地理信息门户(geoportals)和通用数据门户往往响应迟缓、导航困难。找到合适的数据集仅是第一步:用户仍需理解相关服务、解读元数据、下载数据、进行重投影或清洗,并设法将其与其他数据源整合。实际上,尽管投入了大量成本与精力构建此类基础设施,其主要使用者仍局限于GIS专业人员。 传统数据服务将成本与扩展负担转嫁给了数据发布方:每次查询均需经过发布方自行运维的基础设施。而云原生格式则逆转了这一模式——用户可直接从对象存储中的文件读取所需字节,并利用自有计算资源执行查询;发布方只需一次性存储数据,无需为每种潜在使用方式单独部署专用服务;流量增长亦不会迫使发布方持续扩展数据库或API层。 然而,仅提供一个文件并不意味着其内容易于使用。用户仍需发现该文件、理解其含义、判断其是否适用于自身问题,并掌握查询方法。这标志着可访问性的一次重大转变:用户可直接从待解答的问题出发,而非被迫学习多个门户、API、文件格式及GIS工具。经验丰富的实践者得以处理远超人工审阅能力范围的数据集。 更大的机遇在于联邦式(federated)架构。智能代理(agent)无需将所有相关数据集集中存于单一门户或数据仓库中;它可发现不同层级的目录,理解各目录所含内容,并在提出问题时动态联结数据。这种联邦并非通过集中复制全部数据集实现,而是由各发布方持续负责自身目录;代理按需挂载相关数据源,实时联结并返回结果——每个数值均保留其原始发布方信息,且附带可重复执行的查询语句。由此,单个数据集无法回答的问题,借助跨数据集协同分析成为可能。 权威数据(authoritative data)与模型知识(model knowledge)之间的区分至关重要。AI系统不应依赖模型记忆即兴生成答案,而应主动发现恰当数据源、查询已发布数据、保留数据来源信息,并确保结果可验证。 越来越多的组织开始关注:数据实际存放于何处?查询在何处执行?涉及哪些AI模型?系统任一环节的变更难度如何?对于承担关键基础设施职责的政府及机构而言,这些问题绝非次要的采购考量,而是空间数据基础设施(SDI)本身的核心组成部分。 上述三项变革已在推进之中。当前难点在于,如何将它们整合为一套SDI,使真实的数据发布方可自主创建、运行与维护,而无需从零设计整套架构。正因如此,我们非常荣幸能与地理空间社区其他成员共同参与Portolan的建设。 Portolan注册中心(registry)提供了实现联邦所需的发现层:它是一个对独立托管目录进行编目的目录,而非其数据的中央存储库。随着更多发布方以统一方式描述并注册其目录,智能代理便能更清晰地掌握现有数据源、覆盖范围及其组合方式。Portolan发布说明文稿进一步详述了其规范、命令行工具(CLI)、校验器(validator)及注册中心。 CARTO创立之初即致力于推动空间分析的普及化。在其发展历程中,这主要体现为向更广泛用户开放空间分析与可视化能力,同时摆脱传统桌面GIS工作流的束缚。而权威数据正是令此类智能代理具备严肃应用价值的关键所在。若代理可直接发现并查询可信数据集,则用户只需提出问题,即可获得基于这些数据源的答案,无须专家手动协调每一次请求。 这并非消除对地理空间专家的需求,而是重新定义其专业价值的发力点:专家不再耗费时间于系统间数据搬运或重复应答同类请求,转而聚焦于甄选权威数据源、记录数据局限性、定义可靠的分析方法,并管控这些方法的应用方式。智能代理则使此类专业知识惠及更广泛的用户群体。 Portolan定义了一套开放的数据发布基础框架。面向全国范围构建SDI,或以商业数据产品形式运营目录,则需在此基础之上叠加额外能力。CARTO SDI即为Portolan的商业化实现,包含三大核心组件:一是供用户及现有软件发现与访问数据的目录;二是支持用户以自然语言提问的AI接口;三是供发布方统筹管理整个系统的控制平面。 该目录是数据的入口门户,为发布方提供品牌化空间,用以组织公开、私有及授权数据集,展示元数据与文档,并通过下载、现代分析引擎或既定GIS接口提供数据访问。所发布数据本身仍以标准Portolan目录格式保留在发布方自有存储中。 AI接口支持用户通过CARTO平台或发布方自有网站(采用组织批准的模型),以自然语言与目录交互。该代理可发现相关数据集,与其他目录实现联邦式联结,执行空间分析,并返回附带数据来源与可复现查询语句的结果。 对许多发布方而言,商业化亦是保障该模式可持续性的必要环节。部分数据应保持开放;另一些数据集则需通过许可、订阅或按用量计费等方式获取,以支撑其持续生产与维护。CARTO SDI通过同一目录支持上述两种模式,避免发布方被迫另行构建独立的商业交付系统。 这些均为基于真实数据集构建的可用目录。您既可直接查询,亦可将智能代理指向该注册中心,要求其查找相关数据。
The experience for users isn’t much better. Geoportals and general-purpose data portals can be slow and hard to navigate. Finding the right dataset is only the beginning: users still have to understand the service, interpret the metadata, download the data, reproject or clean it, and work out how to combine it with other sources. In practice, much of this infrastructure remains accessible mainly to GIS specialists, despite the cost and effort that went into building it. Traditional data services put the cost and scaling burden on the publisher. Every query passes through infrastructure that the publisher has to operate. Cloud-native formats invert that model. A user can read only the bytes needed from a file in object storage and run the query with their own compute. The publisher stores the data once instead of running a dedicated service for every way it might be used. Traffic can grow without requiring the publisher to keep scaling a database or API in front of it. Making a file available doesn’t make its contents easy to use. People still need to find it, understand it, decide whether it is appropriate for their question, and know how to query it. This is a major change in accessibility. A user can begin with the question they need answered instead of learning several portals, APIs, file formats, and GIS tools. Experienced practitioners can work across far more datasets than they could reasonably inspect by hand. The bigger opportunity is federation. An agent doesn’t need every relevant dataset to live in one portal or warehouse. It can discover catalogs at different levels, understand what each one contains, and join their data when the question is asked. This isn’t federation through a central copy of every dataset. Each publisher remains responsible for its own catalog. The agent attaches the relevant sources, joins them live, and returns an answer in which every figure keeps its publisher and a query that can be run again. A question that no single dataset can answer becomes possible because the agent can work across all of them. The distinction between authoritative data and model knowledge is critical. An AI system shouldn’t improvise an answer from what the model remembers. It should find the right sources, query the published data, preserve their provenance, and make the result possible to check. More organizations are asking where their data lives, where queries run, which AI models are involved, and how easily any part of the system can be changed. For governments and organizations responsible for critical infrastructure, these aren’t secondary procurement questions. They are part of the SDI itself. These three changes are already underway. What is still difficult is assembling them into an SDI that a real data publisher can create, operate, and maintain without having to design the whole architecture from scratch. That is why we’ve been very happy to join other members of the geospatial community in building Portolan. The Portolan registry adds the discovery layer needed for federation. It is a catalog of independently hosted catalogs, not a central repository of their data. As more publishers describe and register their catalogs in the same way, an agent gets a better view of which sources exist, what they cover, and how they can be combined. The Portolan launch post explains the specification, CLI, validator, and registry in more detail. CARTO was created to democratize access to spatial analytics. For much of our history, that meant making spatial analysis and visualization available to more people without requiring traditional desktop GIS workflows. Authoritative data is what makes those agents useful for serious work. If the agent can discover and query trusted datasets directly, a user can ask a question and receive an answer grounded in those sources without needing an expert to manually broker every request. That doesn’t remove the need for geospatial experts. It changes where their expertise has the most leverage. Instead of spending their time moving data between systems or answering the same requests repeatedly, they can select authoritative sources, document limitations, define reliable analytical methods, and govern how those methods are used. Agents make that expertise available to many more people. Portolan defines an open foundation for publishing data. Running an SDI for an entire country, or operating a catalog as a commercial data product, requires additional capabilities around that foundation. CARTO SDI is our commercial implementation of Portolan. It has three main parts: a catalog through which people and existing software can find and access data, an AI interface through which users can ask questions, and a control plane through which the publisher operates the whole system. The catalog is the front door to the data. It gives publishers a branded place to organize public, private, and licensed datasets; expose their metadata and documentation; and make them available through downloads, modern analytical engines, or established GIS interfaces. The published data remains a standard Portolan catalog in the publisher’s storage. The AI interface lets users interact with that catalog in natural language, either through CARTO or from the publisher’s own website using models the organization approves. The agent can discover relevant datasets, federate them with other catalogs, run spatial analysis, and return an answer with sources and a reproducible query. For many publishers, monetization is also part of making the model sustainable. Some data should remain open; other datasets require licenses, subscriptions, or usage-based access to fund their continued production and maintenance. CARTO SDI supports both through the same catalog rather than forcing publishers to build a separate commercial delivery system. These are working catalogs built from real datasets. You can query them directly or point an agent at the registry and ask it to find relevant data across several publishers. Each result carries its producer, the licence, and a link back to the authoritative source. For commercial and licensed data, that increased usage can also support the economics of maintaining the data. A publisher that can see demand, manage access, and charge where appropriate has a clearer path to sustaining the catalog over time. Portolan is still early. The specification and tooling will continue to change as more publishers test them against different datasets and operating environments. The next useful milestones are more catalogs, wider format coverage, native support in more clients, and implementations from organizations other than CARTO. Let us work with you to show what this new way of publishing data could look like for your organization. For publishers that already expose data in cloud-native formats, the first step may be mostly about metadata: organizing the catalog, adding the documentation agents need, and making the existing files discoverable through Portolan. There may be no reason to move or convert the underlying data. Where the current infrastructure relies on traditional services or formats, we can build a proof of concept from a representative group of datasets. The aim isn’t to replace an entire national SDI in a few weeks. It is to make the architecture concrete enough to evaluate: the catalog in infrastructure you control, an AI interface on top of it, and a real question answered by discovering and federating the necessary sources. The serverless model makes that experiment much less expensive than building another portal or standing up a parallel serving stack. It gives publishers a way to test the approach, understand what would need to change, and plan a broader modernization at the right pace. If you’d rather start with the open project, the specification, tools, and registry are available at portolan-sdi.org, and the community meets in the open every week