cd /news/artificial-intelligence/using-okf-with-knowledge-catalog-to-… · home topics artificial-intelligence article
[ARTICLE · art-112016] src=cloud.google.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Using OKF with Knowledge Catalog to serve context for agents

Google Cloud announced that OKF v0.2 bundles can now be published into Knowledge Catalog, Google Cloud's context engine for agents, enabling organization-wide discovery, governance, and secure access. The integration uses a one-time setup and a single push via sample code in the Knowledge Catalog repository, mapping bundle concepts to EntryType 'okf-bundle' and AspectType 'okf' with 13 fields. This allows agents to retrieve context from a governed index alongside existing technical metadata, with IAM-based access control.

read10 min views2 publishedAug 26, 2026

We continue to iterate on the Open Knowledge Format (OKF), an open specification that formalizes the LLM-wiki pattern into a portable, interoperable format. But a big question remains: How can you share and govern access to an OKF bundle across an organization?

OKF v0.1 established a portable format for the context agents need: markdown files with YAML frontmatter, one required field, and five conventions. Then, OKF v0.2 added the trust signals (provenance, verification, freshness, attestation) that a machine-authored bundle requires to be relied on, allowing a team to publish a trustworthy bundle for its own agents.

However, what OKF does not answer is how teams share their bundles across an organization. A git repo per bundle is portable, but it is not searchable alongside the data it describes, it cannot be secured and governed using the same organizational identity and compliance policies, and it does not sit next to the technical metadata (schemas, lineage, ownership) that data teams already work in. Every downstream agent must know where each bundle resides, and that does not scale beyond a small number of bundles.

To scale an OKF bundle across an organization, you can use Knowledge Catalog, Google Cloud's context engine for agents. By mapping the bundle onto Knowledge Catalog's existing types, every concept becomes discoverable, governed, and reachable by any agent already reading from the catalog.

Every agent that queries Knowledge Catalog reads from one governed index over what the organization already has in BigQuery, Cloud Storage, operational databases, and applications. Each entry carries schema, lineage, ownership, and tags, and can be extended with typed aspects that add domain-specific fields. The same catalog exposes search and cross-project lookup to retrieve optimized context for each agentic query. The context retrieval is secure and governed by IAM controls, so agents can only see the entries they have access to based on IAM identity.

Publishing an OKF bundle into Knowledge Catalog takes a one-time setup and a single push. Both use the OKF sample code in the Knowledge Catalog repository, whose wrappers call gcloud dataplex

for setup and delegate push to `kcmd`

(the Metadata-as-Code CLI in the same repository).

The setup registers three Knowledge Catalog resources: an EntryGroup to hold the bundle, an EntryType named okf-bundle

for its concepts, and an AspectType named okf that carries the OKF signal fields (from the okf-aspect.json

schema in the sample code). The push then creates one okf-bundle

Entry per concept, each with two Aspects: an overview

Aspect for the markdown body, and an okf

Aspect for the structured signals. Display name, description, and tags live on the Entry itself. The bundle's index.md

navigation files and its root log.md

are also published as Entries: index files carry only the overview

Aspect (no OKF frontmatter), and log.md

carries both Aspects with okf_type: Log

.

Everything Knowledge Catalog already does for technical metadata (search, IAM, lineage, cross-project discovery) applies equally to OKF bundles, alongside the data they describe.

The okf-aspect.json schema in the sample code defines the AspectType. It carries 13 fields covering the full

| | | | | |---|---|---|---| | 1 | | string | The OKF document type (freeform, e.g. | | 2 | | record | Actor and timestamp for the last meaningful change. | | 3 | | array of | Materials the concept derives from, with credibility signals. | | 4 | | array of | Verification events. A | | 5 | | string | Lifecycle state: | | 6 | | datetime | Absolute point in time (RFC3339 with an explicit offset) on or after which the content is stale. | | 7 | | record | Period the source usage counts were measured over. | | 8 | | string | How an Attested Computation runs (e.g., | | 9 | | array of | Typed named holes a caller may fill. The only surface a caller may vary. | | 10 | | string | Path to a file holding the computation body. | | 11 | | record | How the computation runs and what evidence it must return. | | 12 | | record | Deterministic code that takes a receipt and returns a verdict. | | 13 | | string | Producer-defined frontmatter the template does not model, as JSON |

Every field is annotated with a display name, a description, and a mandatory index. Any top-level scalar field in the okf

Aspect (okf_type

, status

, stale_after

, runtime

, computation

, extra

) can drive Knowledge Catalog search predicates directly, so aspect:acme-analytics.us-central1.okf.okf_type=Metric

returns every OKF Metric in scope. Scalar subfields of record fields (generated.by

, usage_window.from

, executor.resource

, attester.resource

) also drive predicates. The array fields (sources

, verified

, parameters

) are not server-side searchable on their subfields; agents narrow on them client-side after entries.get

with view=ALL

. One caveat for search predicates on datetime

-typed fields (stale_after , generated.at

, usage_window.from

/.to

), use a bare date (`stale_after=2026-12-31`

) or a range comparison (`stale_after>2026-01-01`

), not the full RFC3339 timestamp.

kcmd push

reads an OKF bundle from git and writes each concept as an Entry in the target Knowledge Catalog EntryGroup. index.md

files become Entries too, and each concept is parented to the index above it, so the bundle's directory structure survives as a browsable hierarchy.

kcmd

expects a bundle in the Documents Layout: markdown files under a catalog/

subdirectory, and a catalog.yaml

at the bundle root that lists the snapshot's entry and aspect types. The sample code's setup.ts

generates catalog.yaml

from its --entry-group flag (default okf_demo

), so a reader wiring the sample to a new bundle passes the flag rather than editing catalog.yaml

by hand.

Here is an end-to-end workflow for the Acme Retail bundle that we introduced in the OKF v0.2 blog:

To pick a different EntryGroup name or push a different bundle, pass --entry-group your-name

to setup.ts

and --bundle path/to/your/bundle to push.ts

. For example: bun run setup.ts --entry-group acme-bundle followed by bun run push.ts

. This regenerates the manifest, so subsequent push

, pull

, and cleanup

all target the new EG; delete earlier EGs manually with gcloud dataplex entry-groups delete <name> --project <your-project> --location <your-location> .

The Acme Retail bundle is a synthetic OKF bundle for a US retailer's BigQuery estate. It contains nine leaf concepts across six directories (attesters

, tables

, metrics

, computations

, policies

, skills

), each with its own index.md

, plus a bundle root with its own index.md

and log.md

. That's 17 pushed Entries in total; Dataplex auto-creates one <eg>_entry

alongside, so gcloud dataplex entries list

returns 18 rows.

After the push completes:

Every concept markdown file is a Knowledge Catalog Entry, discoverable by search across the whole project or organization, depending on IAM configuration.

The revenue-ytd

Attested Computation appears in the console with its sanctioned SQL, its executor, its attester, its verification history, and the full concept body.

An analyst searching Knowledge Catalog for "revenue" finds Acme Retail's business definition alongside the BigQuery table it computes from, both under one permission model.

A downstream agent that already calls LookupContext for BigQuery table Entries retrieves the bundle's context by adding the OKF entry names to its resources

list.

Further, metrics/revenue.md

becomes an Entry with two Aspects. The full entries.get

response (with `view=ALL`

) looks like:

The overview

Aspect holds the full body of revenue.md

. The okf

Aspect carries the structured signal fields, so agents get provenance, source, and OKF type in a form they can filter on directly instead of parsing markdown. Server-side searchEntries filters on the top-level scalar fields and on the scalar subfields of record fields; agents narrow further on the array-element subfields client-side after entries.get. (Aspects and EntryTypes are keyed by project number in real API responses and search predicates; the acme-analytics

project ID is shown throughout for readability.)

Once the bundle is in Knowledge Catalog, it provides two capabilities to any agent that reads from the catalog:

Discoverability across the organization. Agents find bundle concepts through the same searchEntries and LookupContext APIs they already use for cataloged data, so an OKF bundle appears alongside BigQuery tables and other resources in every query it matches.

Governance. Bundle Entries inherit IAM from the EntryGroup, so a single agent call returns exactly what the caller is permitted to read, with no parallel permission model to maintain.

Discoverability across the organization OKF bundle Entries appear in searchEntries results alongside BigQuery tables and other cataloged resources, so an agent already querying the catalog picks up new bundles automatically. To retrieve a concept's body, trust signals, or linked concepts from a match, the agent moves to LookupContext and

entries.get

.A LookupContext call looks like this:

The response is a single context

field containing a pre-formatted YAML block. The block carries the entry's catalogEntry

, its type, its description, its tags as labels, and its overview

: the full markdown body of the concept, including its trust and freshness section. LookupContext does not render custom Aspects, so an agent that needs the structured OKF signal fields (okf_type

, generated

, sources

, and the other ten) reads them with entries.get

and view=ALL

alongside the LookupContext call.

There is no repository clone, no manual Aspect merging, and no re-parse of frontmatter. The agent uses the same API call any Knowledge Catalog client already makes.

An agent traversing an OKF bundle typically follows a three-step flow. An agent that already knows the specific Entry names it needs skips step 1. An agent that already knows the target EntryGroup and wants to enumerate the bundle exhaustively substitutes entryGroups.entries.list

for step 1. searchEntries returns candidate Entry names and descriptions. Its scope

accepts a project or organization; narrowing within that scope happens through query terms, including aspect predicates like aspect:acme-analytics.us-central1.okf.okf_type=Metric

.

LookupContext on the top few Entry names (up to ten per call) returns the full concept body as pre-formatted YAML; context_budget

caps the response size.

entries.get

with view=ALL

on any Entry returns its structured OKF signals (okf_type

, generated

, sources

, and the other ten) directly, which the agent can then filter or attest on.

When a concept's sources[] references another concept by path, the agent calls LookupContext on that Entry name to walk the reference.

The full response for the Revenue Entry:

Governance Permissions on the EntryGroup use standard Knowledge Catalog IAM. An agent that names both a bundle concept and the BigQuery table it grounds against in one call receives both, each subject to its own existing access control list (ACL), so the response carries only what the caller is already permitted to read. There is no parallel permission model to maintain.

Reading agents use roles/dataplex.catalogViewer

, which grants the read paths: entries.get

, LookupContext, and searchEntries. The identity that runs kcmd push

uses roles/dataplex.catalogEditor

, which grants the write paths: entries.create

and entries.patch

. One EntryGroup per bundle-owning team is the multi-team pattern, and IAM on the EntryGroup cascades to its Entries.

LookupContext resolves the entry names it is given, up to ten per call, within a single location. It does not follow links out of a concept's body, so an agent that wants a referenced concept must name it explicitly. Place the bundle's EntryGroup in the same location as the data it describes to fetch both in one call.

kcmd push

is an idempotent upsert. Re-running is safe (no duplicates, no error), but every push writes every Entry. Concept deletes require an explicit kcmd delete

on the Entry, or cleanup.ts

to remove the whole EntryGroup at once; cleanup.ts

deletes only the EntryGroup and its Entries, so the shared okf

AspectType and okf-bundle

EntryType stay in place for other bundles that reference them. For continuous ingestion in production, wire a CI job to kcmd push

on every commit to the bundle repository, using a service-account credential with roles/dataplex.catalogEditor

on the target EntryGroup.

OKF defines what a trustworthy bundle looks like. Knowledge Catalog makes it reachable across the organization. To get started, check out the following resources:

Read the OKF v0.2 spec and browse the Acme Retail bundle.

Author a small bundle for one domain your team owns.

Sync it into your Knowledge Catalog project using the sample code's setup.ts

(which registers the resources) and push.ts

(which delegates to kcmd

).

Point your existing agents at Knowledge Catalog. New context becomes reachable through the same LookupContext and searchEntries calls they already use.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google cloud 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/using-okf-with-knowl…] indexed:0 read:10min 2026-08-26 ·