OpenDome is an open-core data platform for teams that need to put their
structured and unstructured data behind a single, governed, queryable
interface — without handing it to a third party.
You run a cell: a self-contained, per-tenant Kubernetes namespace that holds
your lakehouse (Iceberg + Lance), your pipelines, and the consumption engine.
The name of the project comes from what the cell implements: a dome is a
tenant’s governed enclosure — one policy boundary around all of that tenant’s
data — and the cell is its runtime. Everything that reads data goes through one
surface — the OpenDome Semantic Layer (OSL).
Structured data, unstructured data, metrics, entities and natural language are
consumed exclusively through OSL. Trino, Lance and dbt are internal
engines: they never expose a public endpoint. No consumer talks to Trino
directly.
OSL describes a semantic model that spans two physical substrates with a single
declarative YAML language:
Structured plane
Iceberg tables via Trino. Entities, metrics, dimensions, data contracts.
Here OSL’s contract target is a superset of dbt MetricFlow — your
MetricFlow YAML stays valid, and the engine documents honestly which
subset it resolves today.
Unstructured plane
Lance collections via LanceDB. Chunk schema, embeddings, governed vector
retrieval with metadata prefilters and row-level ACL. This is the OSL-only
extension.
The bridge between them is the JointEntity: a logical object (Customer,
Policy, Order) whose primary key comes from a structured model and which
exposes attributes drawn from both planes. One call, one policy decision, one
audit trail. See Concepts for the full model.
Excellent systems already govern individual legs of this problem. The gap
OpenDome targets is not any one mechanism — it is a common semantic model and a
bounded query language whose structured and content operations are governed
by one compiled policy.
The clearest illustration of that gap is vector serving. Databricks Unity
Catalog secures vector-index objects, while its own documentation notes that
table row filters and column masks do not govern vector serving. Snowflake
Cortex Search serves under the owner’s rights rather than the consuming
user’s row policy. In both, the relational policy you wrote stops at the edge
of retrieval — which is exactly where analysts and AI agents now start reading.
OpenDome compiles the same consumer policy into the retrieval leg.
How the neighbouring system families relate, factually:
Relational semantic models and governed business names
Structured plane only — no governed content retrieval, no cross-plane policy
Semantic-SQL / NL-over-data — SUQL, LOTUS
Structured + unstructured operations in one SQL, with LLM operators for semantic filtering and joining
No relational or cross-plane policy is compiled into the query
Foundry Ontology
The closest broad entity abstraction — granular policies, object-scoped embeddings
Cross-plane policy is partial; OSL’s modeling delta is an open, additive MetricFlow-compatibility target with explicit content facets, honest authoritative bridges and a consumable domain graph
Enterprise search — Glean, Kendra, Azure AI Search, Elastic DLS
Permission-aware document retrieval at scale
Document-ACL filtering — it does not normally compile the same relational policy the warehouse enforces
Mature relational rewriting and compliance mechanisms
Structured plane only
OpenDome’s contribution is the composition: one delegation-chain policy
compiled onto SQL row/column rewriting, governed vector retrieval, graph
hydration, exact redaction-version selection and cross-modal correlation — with
the same fail-closed default at every seam.
Two boundaries we state as plainly as the research paper does: OSL does not
advance ANN indexing (the reference engine currently evaluates FLAT retrieval),
and a relation-by-relation “AI join” is an explicit non-goal — the only
cross-modal join is a sealed-key equijoin on an authoritative entity key.
The Apache-2.0 distribution — the public
opendome.eu/platform repository —
is everything needed to run and consume a cell on your own infrastructure:
The cell data plane — Helm charts for a standalone, self-contained tenant
(standalone is the default profile; no control plane).
The OSL engine (tenant-semantic-api) serving POST /osl/query — the
single consumption endpoint — plus governed previews and metadata routes.
Connectors (Singer/Meltano taps: Postgres, MySQL, Salesforce, custom) and
the documents pipeline (extract → embed → Lance) with in-cell Dagster
orchestration.
Policy evaluation (SQL allowlist, row filters, projection, ACL’d
retrieval) — fail-closed, enforcement on by default.
Token minting and verification (ActorToken v2; BYO-OIDC upgrade path).
The OSL specification, JSON Schemas and the Segura worked example.
The embedder that vectorizes documents and queries is consumed as an
external endpoint you point the cell at — it is never an in-cluster
workload. See Ingest documents.
OSS vs EnterpriseWhat's Apache-2.0 today, and what the managed/Enterprise tier adds — the console, control-plane, vertical packs, and compliance.
Adopters (a customer data team) — model your domain in OSL, connect
sources, serve governed metrics and retrieval. Start with
Concepts then the Quickstart.
Agent authors — build agents against OSL’s MCP tools and REST. See the
API reference.
Spec implementers — build an OSL-conformant engine. See
Spec & governance.