Skip to content

What is OpenDome?

OpenDome is an open-core data platform for teams that need to put their structured and unstructured data behind a single, governed, queryable interface — without handing it to a third party.

You run a cell: a self-contained, per-tenant Kubernetes namespace that holds your lakehouse (Iceberg + Lance), your pipelines, and the consumption engine. The name of the project comes from what the cell implements: a dome is a tenant’s governed enclosure — one policy boundary around all of that tenant’s data — and the cell is its runtime. Everything that reads data goes through one surface — the OpenDome Semantic Layer (OSL).

The one rule: OSL is the sole consumption layer

Section titled “The one rule: OSL is the sole consumption layer”

Structured data, unstructured data, metrics, entities and natural language are consumed exclusively through OSL. Trino, Lance and dbt are internal engines: they never expose a public endpoint. No consumer talks to Trino directly.

OSL describes a semantic model that spans two physical substrates with a single declarative YAML language:

Structured plane

Iceberg tables via Trino. Entities, metrics, dimensions, data contracts. Here OSL’s contract target is a superset of dbt MetricFlow — your MetricFlow YAML stays valid, and the engine documents honestly which subset it resolves today.

Unstructured plane

Lance collections via LanceDB. Chunk schema, embeddings, governed vector retrieval with metadata prefilters and row-level ACL. This is the OSL-only extension.

The bridge between them is the JointEntity: a logical object (Customer, Policy, Order) whose primary key comes from a structured model and which exposes attributes drawn from both planes. One call, one policy decision, one audit trail. See Concepts for the full model.

Excellent systems already govern individual legs of this problem. The gap OpenDome targets is not any one mechanism — it is a common semantic model and a bounded query language whose structured and content operations are governed by one compiled policy.

The clearest illustration of that gap is vector serving. Databricks Unity Catalog secures vector-index objects, while its own documentation notes that table row filters and column masks do not govern vector serving. Snowflake Cortex Search serves under the owner’s rights rather than the consuming user’s row policy. In both, the relational policy you wrote stops at the edge of retrieval — which is exactly where analysts and AI agents now start reading. OpenDome compiles the same consumer policy into the retrieval leg.

How the neighbouring system families relate, factually:

System family Strength Boundary
Semantic layers — Looker, Cube (and MetricFlow itself) Relational semantic models and governed business names Structured plane only — no governed content retrieval, no cross-plane policy
Semantic-SQL / NL-over-data — SUQL, LOTUS Structured + unstructured operations in one SQL, with LLM operators for semantic filtering and joining No relational or cross-plane policy is compiled into the query
Foundry Ontology The closest broad entity abstraction — granular policies, object-scoped embeddings Cross-plane policy is partial; OSL’s modeling delta is an open, additive MetricFlow-compatibility target with explicit content facets, honest authoritative bridges and a consumable domain graph
Enterprise search — Glean, Kendra, Azure AI Search, Elastic DLS Permission-aware document retrieval at scale Document-ACL filtering — it does not normally compile the same relational policy the warehouse enforces
Fine-grained query rewriting — Oracle VPD, Qapla, Sieve, Blockaid Mature relational rewriting and compliance mechanisms Structured plane only

OpenDome’s contribution is the composition: one delegation-chain policy compiled onto SQL row/column rewriting, governed vector retrieval, graph hydration, exact redaction-version selection and cross-modal correlation — with the same fail-closed default at every seam.

Two boundaries we state as plainly as the research paper does: OSL does not advance ANN indexing (the reference engine currently evaluates FLAT retrieval), and a relation-by-relation “AI join” is an explicit non-goal — the only cross-modal join is a sealed-key equijoin on an authoritative entity key.

The Apache-2.0 distribution — the public opendome.eu/platform repository — is everything needed to run and consume a cell on your own infrastructure:

  • The cell data plane — Helm charts for a standalone, self-contained tenant (standalone is the default profile; no control plane).
  • The OSL engine (tenant-semantic-api) serving POST /osl/query — the single consumption endpoint — plus governed previews and metadata routes.
  • Connectors (Singer/Meltano taps: Postgres, MySQL, Salesforce, custom) and the documents pipeline (extract → embed → Lance) with in-cell Dagster orchestration.
  • Policy evaluation (SQL allowlist, row filters, projection, ACL’d retrieval) — fail-closed, enforcement on by default.
  • Token minting and verification (ActorToken v2; BYO-OIDC upgrade path).
  • The OSL specification, JSON Schemas and the Segura worked example.

The embedder that vectorizes documents and queries is consumed as an external endpoint you point the cell at — it is never an in-cluster workload. See Ingest documents.

  • Adopters (a customer data team) — model your domain in OSL, connect sources, serve governed metrics and retrieval. Start with Concepts then the Quickstart.
  • Agent authors — build agents against OSL’s MCP tools and REST. See the API reference.
  • Spec implementers — build an OSL-conformant engine. See Spec & governance.