Light-Fabric
Light-Fabric is a high-performance, unified platform for managing the lifecycle, governance, and orchestration of enterprise AI services including agentic services, agents, tools, skills, memories, MCP servers, APIs, gateways and workflows.
Why Light-Fabric?
We chose the name Light-Fabric because it embodies the “Unified Governance” required for enterprise-grade AI:
- Unified Control Plane: Light-Fabric provides a single point of truth for discovering, governing, and auditing agents, MCP servers, and APIs via the
light-portal. - Enterprise Governance: It prioritizes security and policy enforcement (such as fine-grained authorization) over pure decentralized autonomy, making it safe for corporate environments.
- Integrated Ecosystem: It “weaves” together distributed components—from memory units (Hindsight) to centralized skills—into a cohesive, observable system.
- Durable Identity: The name emphasizes the platform’s role as the infrastructure foundation, remaining relevant regardless of the underlying implementation details.
Technical Advantages
By building Light-Fabric on a Rust foundation, we achieve:
- Performance: Built on top of
tokioandaxumfor maximum throughput and memory safety. - Native Intelligence: Specialized crates for Hindsight memory, tool calling, and workflow orchestration.
- Production Ready: Includes robust features like retries, failover, and observability out of the box.
Core Components
The Light-Fabric is composed of modular crates, infrastructure frameworks, and reference applications:
Crates
crates/model-provider: A unified interface for multiple LLM providers (Ollama, etc.).crates/hindsight-client: Client for the Hindsight biomimetic memory system.crates/mcp-client: Implementation of the Model Context Protocol (MCP) for tool discovery and execution.crates/portal-registry: Integration with the Light-Portal for service registration and discovery.crates/light-runtime: Core runtime foundation for building agentic and microservice components.crates/light-rule: High-performance rule engine for fine-grained authorization and data filtering.crates/workflow-core&workflow-builder: Core engine and builder for complex agentic workflows.crates/config-loader: Flexible configuration management for enterprise environments.crates/asymmetric-decryptor&symmetric-decryptor: Security utilities for sensitive data handling.
Frameworks
frameworks/light-axum: A specialized microservice & agentic framework built on top of the Axum web ecosystem.frameworks/light-pingora: High-performance proxy and gateway framework built on top of Cloudflare’s Pingora.
Applications
apps/light-agent: A managed AI agent capable of using tools, accessing memory, and executing complex tasks.apps/light-gateway: An enterprise-grade gateway for securing and governing API and agent traffic.apps/light-workflow: A service for orchestrating and executing long-running agentic workflows.
Getting Started with Light-Fabric
This guide will help you set up a local development environment for Light-Fabric, including the AI Gateway, Agent Engine, and the management Portal.
Prerequisites
- Rust: Latest stable version.
- Docker: For running database and backend services.
- Node.js: For running the
portal-viewUI. - Git: To clone the necessary repositories.
Local Development Setup
To run the entire ecosystem locally, use portal-config-loc. Its deployment
script downloads released service and UI archives from the configured CDN.
1. Initialize Workspace
Create a unified workspace directory (e.g., ~/lightapi) and clone the core management repositories:
cd ~
mkdir -p lightapi
cd lightapi
# Clone the local deployment configuration
git clone [email protected]:lightapi/portal-config-loc.git
2. Deploy Local Services
Light-Fabric services are orchestrated via Docker Compose scripts in portal-config-loc. The following command starts the PostgreSQL database and the core services (including the Rust-based components):
cd ~/lightapi/portal-config-loc
./scripts/deploy-local.sh pg rust
3. Import Initial Data
For a full deployment, deploy-local.sh downloads events.zip from the CDN
and defaults IMPORT_EVENTS to auto. It imports the extracted events.json
only when the event store is empty. Set IMPORT_EVENTS=force only when you
intentionally need to replay the released baseline into a non-empty store.
4. Update /etc/hosts
The platform uses virtual hosts for local routing. Add the following entry to your /etc/hosts file (replace with your actual local IP if necessary):
127.0.0.1 local.lightapi.net locsignin.lightapi.net
Running the Management Portal
The Light-Portal provides a unified UI for onboarding MCP servers, configuring AI Gateways, and interacting with agents.
cd ~/lightapi
git clone [email protected]:lightapi/portal-view.git
cd portal-view
npm install
npm run dev
Navigate to https://localhost:3000 and log in with your developer credentials.
Cloud Development (Coming Soon)
We are currently preparing a Cloud Development Server. This will allow developers to:
- Connect to a shared, high-performance AI Gateway.
- Onboard and test MCP servers without a full local installation.
- Collaborate on shared agentic workflows and Hindsight memory banks.
Stay tuned for the connection details and onboarding guide for the cloud environment.
Contributing to Light-Fabric
If you are developing for the Rust crates specifically:
cd ~/lightapi
git clone [email protected]:networknt/light-fabric.git
cd light-fabric
cargo build
Model Providers
Light-Fabric provides a unified, high-performance interface for interacting with diverse Large Language Model (LLM) providers. This abstraction is centered around the Provider trait, allowing applications to remain model-agnostic while leveraging advanced capabilities like native tool calling and prompt caching.
The Provider Trait
All model integrations implement the Provider trait, which supports:
- One-shot and Multi-turn Chat: Simplified APIs for simple prompts and full conversation histories.
- Structured Tool Calling: Native integration for function calling (OpenAI-style).
- Capabilities Detection: Programmatic checks for vision, native tool support, and prompt caching.
Supported Cloud Providers
Light-Fabric supports all major LLM providers. Because the Provider trait is model-agnostic, the framework is compatible with the latest flagship releases as soon as they are available.
- OpenAI: Native support for the GPT-5 series (5.4, mini, nano), the o4 reasoning models, and full legacy support for GPT-4o and GPT-4 Turbo.
- Anthropic: Support for the Claude 4 generation, including Opus 4.7, Sonnet, and Haiku.
- Google Gemini: Support for Gemini 3.1 Pro and Flash, leveraging Vertex AI or AI Studio for multi-modal and long-context tasks.
- Azure OpenAI: Enterprise-grade OpenAI deployments with support for the latest model deployments.
- AWS Bedrock: Access to the latest Claude and Titan models hosted on Amazon Web Services.
- OpenRouter: Access to hundreds of open-source and proprietary models via a single unified API.
- Telnyx: Support for models hosted on the Telnyx platform.
- GLM (Zhipu AI): Support for the ChatGLM/GLM-5 series of models.
Local & Specialized Providers
- Ollama: Seamless integration with local models running on your machine.
- OpenAI-Compatible: A generic
CompatibleProviderfor any service implementing the OpenAI REST API. - GitHub Copilot: Integration with GitHub Copilot Chat for developer-centric workflows.
Meta-Providers (Orchestration)
These providers wrap other providers to add resilient or intelligent behavior:
- ReliableProvider: Enhances any base provider with retries, exponential backoff, and automatic failover to fallback models.
- RouterProvider: Dynamically routes requests to different models based on hints or input complexity.
CLI & Tooling Integrations
Light-Fabric includes specialized integrations for developer tools and terminal environments:
- Claude Code CLI: Integration with Anthropic’s Claude Code environment.
- Gemini CLI: Terminal-based access to Google’s Gemini models.
- KiloCLI: Light-Fabric’s native CLI integration for rapid testing and automation.
Key Capabilities
Providers can be queried for their support of advanced features:
- Native Tool Calling: Efficiently generate structured function calls.
- Vision: Process images alongside text prompts.
- Prompt Caching: Leverage provider-side caching to reduce latency and costs for long contexts.
Why Rust
Light-Fabric is written in Rust. Not because Rust is fashionable, and not because the team enjoys fighting a borrow checker, but because the platform is built for a specific world: enterprise backend services and agentic workflows that are increasingly written by AI agents and verified by automated gates rather than read line-by-line by humans.
In that world the traditional trade-off — “Python is faster to write, Rust is faster to run” — no longer holds the way it did. The cost of writing code is collapsing. The cost of running it, trusting it, and operating it is not. This document explains the reasoning.
1. The premise has changed
The classic argument for Python and JavaScript on the backend was never really about the machine. It was about people:
- Fewer lines to type, so features ship faster.
- A larger hiring pool, so code is easier to staff and maintain.
- A shallower learning curve, so review and onboarding are cheap.
Every one of those advantages is a human throughput argument, and human throughput is exactly what stopped being the bottleneck. When a model writes the first draft, another model reviews it, and CI gates decide whether it merges, “how long does it take a developer to type this” is no longer the constraint worth optimizing. What remains is what the machine has to do afterwards, forever: execute it, pay for it, and not break in production at 3 a.m.
Rust wins on everything that is left.
2. The compiler is the cheapest reviewer we have
An AI agent is a probabilistic writer. It produces plausible code. The engineering question is not “will it make mistakes” — it will — but how early and how cheaply the mistakes are caught.
Light-Fabric’s quality pipeline has several gates: cargo build, clippy with
warnings denied, unit and integration tests, contract/conformance suites, and a
model-based review pass. They are not equally priced:
| Gate | Cost per run | Catches |
|---|---|---|
cargo check / cargo build | seconds, deterministic, no tokens | types, ownership, lifetimes, exhaustiveness, nullability, data races |
clippy -D warnings | seconds | misuse patterns, unchecked unwraps, sloppy idioms |
| Tests / conformance | minutes | behavior the tests anticipated |
| Model review | tokens and wall-clock, non-deterministic | intent, design, what the tests did not anticipate |
The compiler is the only gate that is free, total and deterministic. It does not sample, it does not get tired, and it does not need a test case to have been written in advance. In a dynamic language, every class of error the Rust compiler rejects outright must instead be caught by a test someone remembered to write or by a reviewer who happened to look — and in an agentic pipeline, both of those are paid for in tokens.
Rust’s type system also encodes decisions the platform actually cares about and that reviewers routinely miss:
Result<T, E>makes every failure path explicit; an ignored error is a compile-time warning, not a silent production incident.Option<T>eliminates the null-dereference class entirely.Send/Syncand the borrow checker make data races across the thousands of concurrent agent sessions Light-Fabric runs unrepresentable, not merely “unlikely if the code is careful”.- Exhaustive
matchmeans adding a new variant to an event, message or state enum forces every handler to be revisited. Across a workspace this large, that single property has caught more would-be regressions than any test suite.
This is the real answer to “if it compiles, it runs.” It is not that Rust code is bug-free. It is that a whole family of bugs cannot reach the tests, the reviewer, or production, so the expensive gates spend their budget on logic and design instead of on null checks and race conditions.
3. Verbosity is a one-time cost; runtime is a forever cost
The standard objection is the “token tax”: Rust is more verbose, so agents spend more tokens generating it. That was a real argument. It has weakened considerably, and it was always measured against the wrong denominator.
The generation cost is paid once per change. The runtime cost is paid on every request, in every environment, for the life of the service. Light-Fabric runs gateways, agent runtimes and workflow engines that sit on the hot path of every call in the fabric; the asymmetry is not close:
- Memory. Rust services hold a resident set measured in tens of megabytes, with no GC heap to size and no GC pause to tune. The equivalent JVM or Python service reserves an order of magnitude more before it serves a single request. On a per-pod basis in Kubernetes, that is the difference between packing dozens of services on a node and packing a handful.
- Latency tail. There is no stop-the-world collector, so p99 tracks p50 instead of spiking. For a gateway that fronts every LLM and MCP call, tail latency is the product.
- Startup. A statically linked binary is serving traffic in milliseconds. This is what makes scale-to-zero, fast rollouts, and short-lived sandboxed workers practical rather than theoretical.
- Footprint. A single self-contained binary with no runtime, no interpreter, and no dependency tree to install at deploy time. Containers are small, the attack surface is small, and the supply chain is one artifact.
There is also a second-order effect that matters specifically for an AI platform: compute spent on the runtime is compute not spent on inference. Every gigabyte and every core the control plane does not consume is budget available to the models the platform exists to serve.
4. Verbosity is also, increasingly, an illusion
Rust reads as verbose next to a Python one-liner, but the comparison usually
omits what the Python one-liner postponed: the type annotations added later, the
validation, the error handling, the test that exists only to catch a None, and
the runtime guard for the concurrency case. Light-Fabric’s Rust states those up
front, where the compiler can enforce them, instead of scattering them through
tests and incident reports.
The asymmetry is also shrinking on the model side. Frontier models generate
idiomatic Rust competently today, and the “compiler loop” failure — an agent
thrashing against lifetime errors — is now mostly a symptom of poor architecture
rather than of the language. In practice it is contained by the same design
rules that make the code good for humans: prefer owned data and Arc at
boundaries, keep lifetimes out of public APIs, keep functions small, and let
clippy steer idiom. Where a module does provoke thrashing, that is a signal
about the design, and a useful one.
5. Where Rust would be the wrong answer, and what we do instead
This is a considered choice, not a purity rule. Rust is the wrong tool in three places, and Light-Fabric does not pretend otherwise:
- The browser. The web is JavaScript and TypeScript. UIs for the portal and consoles belong there, and WASM does not change that.
- Glue, scripting and one-off automation. Build helpers, migration scripts, data wrangling and repo tooling are written in whatever is shortest — shell or Python. They are not on the hot path, they are not mission-critical, and compiling them buys nothing.
- The ML/data ecosystem. Where a mature Python library is the state of the art, the right move is to call it across a process or service boundary, not to reimplement it.
The rule is a boundary, not a language ban: anything on the request path, holding state, enforcing a security decision, or running unattended in production is Rust. Everything around it can be whatever is convenient. The platform is the compiled tier; the conveniences live outside it.
6. What this buys Light-Fabric concretely
The architecture depends on properties that are difficult or expensive to obtain elsewhere:
- Shared engines, many services.
light-agent,light-agent-worker,light-agent-channel,light-gatewayandlight-workfloware thin trust-boundary executables over shared domain crates. Cargo’s workspace and the type system make that sharing safe: a contract change in a shared crate fails to compile in every consumer that has not been updated. See Agent Engine Pattern. - Hot-reload without restarts. Configuration, rules and agent metadata swap
under live traffic via
arc-swap, with the type system guaranteeing readers never observe a torn state. - Gateway-grade proxying.
frameworks/light-pingorabuilds on Cloudflare’s Pingora, a proxy engine that exists in Rust because that class of workload cannot afford a garbage collector. - Massive concurrency per instance.
tokiolets one instance carry thousands of concurrent agent sessions and streaming LLM responses on a small footprint, with the borrow checker — not convention — preventing the races. - Enterprise security posture. Memory-safety defects are the dominant class of exploitable vulnerability in C and C++ infrastructure, and the class Rust removes by construction. For software that terminates TLS, holds credentials and enforces authorization, that is a compliance argument as much as an engineering one.
7. The longer arc
The direction of travel is toward specifications and exit criteria as the real source of truth, with the implementation language demoted to an artifact of the build. Light-Fabric is already organized that way in places: the metadata-driven agent engine, the rule specification, and the workflow specification all describe what should happen, while Rust implements the engine that makes it happen safely and quickly.
If that arc completes and the implementation tier eventually becomes something lower-level still, the property being preserved is not “we write Rust” — it is machine-checkable correctness at the boundary between a probabilistic author and a deterministic machine. Today, Rust is the most practical, widely supported, production-proven form of that guarantee. It also happens to be the fastest and leanest option on the table, which makes the choice easy rather than merely defensible.
Summary
| Concern | Why it favours Rust |
|---|---|
| AI-authored code | The compiler is a free, deterministic, total reviewer; errors are caught before any expensive gate runs |
| Correctness | Null, data-race and unhandled-error classes are eliminated by construction |
| Cost | Generation cost is paid once; memory and CPU are paid forever, on every request |
| Latency | No GC, so tail latency stays flat — decisive for a gateway on every call path |
| Operations | One static binary, millisecond startup, tiny images, small attack surface |
| Evolution | Workspace-wide type checking turns contract changes into compile errors instead of production incidents |
| Security | Memory safety by construction for software that holds credentials and enforces authorization |
Human readability was the argument for dynamic languages, and human readability is the constraint that is going away. What remains is execution cost and verifiable correctness — and those are the two things Rust was built to win.
Agentic Workflow Design
Hybrid Agentic Workflow Specification
Agentic Workflow in Light-Fabric implements a hybrid orchestration model for enterprise business processes. The workflow is deterministic, auditable, and stateful, while selected steps can be executed by agents, API calls, rule engine checks, or humans.
The design goal is not to replace enterprise process control with an open-ended agent loop. The goal is to let agents work inside a managed process that has clear state, clear ownership, repeatable execution, and human approval where needed.
Enterprise Challenge
In regulated or operationally sensitive environments, a purely autonomous AI agent is not enough for long-running business work.
- Compliance requires deterministic process paths, approval records, and audit history.
- Reliability requires long-running state to survive process restarts, UI disconnects, and agent failures.
- Safety requires human-in-the-loop checkpoints for decisions with business, security, or financial impact.
- Coordination requires multiple humans and roles to participate in the same process.
- Testing requires the same workflow to run interactively with humans or headlessly with example data.
Light-Fabric solves this by separating orchestration from execution.
Hybrid Model
The workflow is the deterministic process manager. It defines the ordered steps, conditions, retries, error handling, human checkpoints, and outputs.
Agents are workers inside that process. They can reason, call tools, ask for missing data, and use skills, but they do not own the overall process state.
| Feature | Traditional Workflow | Pure Agent Loop | Light-Fabric Hybrid |
|---|---|---|---|
| Path | Fixed | Dynamic | Fixed path with flexible task execution |
| State | Durable | Often transient | Durable workflow and task state |
| Human input | Forms and approvals | Ad hoc chat | First-class waiting tasks |
| Audit | Strong | Weak | Step-level audit and agent trace |
| API calls | Built into code | Tool calls | Spec-described endpoint invocations |
| Testing | Separate test harness | Prompt replay | Same workflow can run live tests |
Core Separation
There are two related specifications:
-
Agentic Workflow Specification Describes orchestration: task order, branching, human input, assertions, API calls, retries, errors, exports, and state transitions.
-
LightAPI Description Specification Describes API capabilities at the endpoint level: how an endpoint is invoked, what inputs it accepts, what result shape it returns, examples, behavior notes, and result expectations.
This separation is important. The workflow should not duplicate every endpoint contract. It should reference endpoint descriptions and use them to invoke calls, guide agents, and verify results.
Endpoint-Level Consumption
Light-Portal manages API descriptions at the endpoint level, not only at the whole API level.
This is necessary because real workflows often combine one endpoint from one API with one endpoint from another API. For example, onboarding an API to an AI gateway may involve:
- register an API
- create an API version from a specification
- create a development API instance
- configure the API through config server
- link the API instance to a gateway instance
- select endpoints to expose as MCP tools
- create a gateway config snapshot
- reload the gateway through controller
- run MCP tests against the gateway
Each step may come from a different API surface. The workflow consumes only the endpoints it needs.
The recommended model is:
- API-level descriptions can be authored for convenience and consistency.
- Endpoint-level descriptions are published and consumed by agents and workflows.
- Endpoint descriptions inherit shared context such as authentication, environments, sources, and secrets from an API catalog.
- Agents progressively load endpoint information by disclosure level instead of receiving the entire catalog up front.
Progressive Disclosure
Endpoint descriptions should be disclosed to agents in layers:
- index: operation id, title, tags, visibility
- summary: purpose, capability group, lifecycle
- invocation: input shape, request mapping, auth, examples
- behavior: result cases, errors, edge cases, assertions
- full: complete description for debugging or generation
This allows the agent to discover capabilities cheaply, load invocation details only for selected endpoints, and load behavior details only when verification or failure analysis needs it.
Workflow Task Types
The updated workflow specification adds first-class support for the task types needed by agentic API workflows.
Ask Task
ask pauses the workflow and waits for human input. It supports prompts, choices, validation, defaults, timeouts, and sensitive input.
The task returns the user’s answer as task output. The normal export block should move the answer into workflow context.
Example:
- ask-authz:
ask:
prompt: Do you want to configure endpoint authorization?
mode: choice
options:
- label: Configure authorization
value: configure
- label: Skip
value: skip
export:
as:
authzChoice: ${ .result }
Assert Task
assert validates workflow state or API results. It is used for both live tests and interactive workflows.
It supports simple comparisons, JSONPath-style checks, length checks, regex checks, and rule-engine-backed assertions for complex business logic.
Assertion failures should produce structured, catchable errors so workflows can route failures to remediation, task creation, or agent investigation. Complex business assertions can delegate to Light-Rule.
API Call Tasks
The workflow supports direct and description-backed API calls:
- HTTP / OpenAPI
- JSON-RPC
- OpenRPC
- gRPC
- MCP tool/resource/prompt calls
For direct internal calls, jsonrpc can be used with an endpoint, method, params, id, notification flag, and error policy.
For cataloged JSON-RPC, openrpc references an OpenRPC document and method.
For MCP, the workflow references a tool, resource, or prompt and passes arguments. MCP capability descriptions belong in the API description layer; the workflow only selects and invokes them.
Explanation Metadata
Tasks can include explain metadata to help an agent or UI explain what is happening.
Useful fields include:
- purpose
- visible
- before
- success
- failure
- requires
Example:
explain:
purpose: Link the API instance to the development gateway.
visible: true
requires:
- portal-command-token authentication
- apiInstanceId from prior step
Human Task State
Human-in-the-loop behavior must be represented as durable workflow state.
Recommended task states:
A = active
W = waiting for input
C = completed
F = failed
X = canceled
When an ask or approval task reaches W, the process remains active but the task is no longer picked up by the executor. A user, CLI, scheduler, or agent must complete the task through the workflow API.
Waiting tasks should carry:
- prompt
- input mode
- options
- validation rules
- default value
- sensitive flag
- assignment metadata
- explanation metadata
- timeout policy
Assignment And Worklist
Enterprise workflows need more than chat. Some tasks must be assigned to roles or users and coordinated across multiple humans.
Human tasks should support:
- assigned user
- assigned role
- candidate roles
- claimed by
- claimed timestamp
- due timestamp
- priority
- comments
- audit trail
A role-based task appears in the worklist for users with a matching role. Once claimed, it belongs to the claiming user until completed, released, delegated, or timed out.
Client Architecture
light-workflow should run as a containerized backend service alongside other portal services. It owns workflow execution and state. Portal chat, worklist, CLI, scheduler, and agents are all clients of the same workflow APIs.
The client surfaces are:
- Portal Chat: conversational guidance for a single user.
- Worklist: role-based task inbox for approvals, reviews, and coordination.
- CLI: developer, CI/CD, live test, and automation interface.
- Scheduler: periodic headless execution, such as hourly live integration tests.
- Agent: task executor that can call APIs, use skills, and report results back to the workflow.
See Workflow Client Architecture for the dedicated client design.
Workflow Service API
The workflow service should expose one stable API boundary for all clients.
Core operations:
workflow.start
workflow.getInstance
workflow.listInstances
workflow.getEvents
workflow.listTasks
workflow.getTask
workflow.claimTask
workflow.releaseTask
workflow.completeTask
workflow.delegateTask
workflow.cancelInstance
Streaming clients should subscribe to workflow events through Server-Sent Events, WebSocket, or another portal-standard event mechanism.
Important event types:
- workflow started
- task started
- task completed
- task failed
- task waiting for input
- task assigned
- task claimed
- task completed by human
- agent started
- agent completed
- workflow completed
- workflow failed
Live Testing
The same workflow runtime should support interactive runs and headless live tests.
Interactive workflows use ask tasks when decisions or missing values are needed.
Live tests should use example data from LightAPI endpoint descriptions and workflow input fixtures instead of asking the user. Assertions should verify results through assert tasks or rule-engine checks.
This lets the scheduler run workflows every hour against the latest deployed services. When a test fails, the workflow can create a task with the failure detail and assign an agent or human to investigate.
Example: API Onboarding To AI Gateway
An API onboarding workflow can guide a user through a complex multi-endpoint process without requiring a dedicated UI for every operation.
The workflow can:
- ask for or infer the API metadata
- call the register API endpoint
- create an API version from an OpenAPI specification
- create a development API instance
- configure the API
- ask whether fine-grained authorization should be configured
- route to create or select authorization rules
- link the API instance to the development AI gateway
- select endpoints to expose as MCP tools
- create a gateway config snapshot
- reload the gateway through controller
- run MCP tests through the gateway
- assert expected results
- report success or create remediation tasks
The same workflow can run interactively through portal chat, be managed through the worklist, or run headlessly with examples as a live test.
Technical Implementation
The Light-Fabric implementation is split across:
workflow-core: Rust models for the workflow specification.workflow-builder: fluent builders for programmatic workflow construction.light-workflow: runtime service and executor.light-agent: agent execution surface for delegated agent tasks.light-rule: rule engine used by workflow and assertion tasks. See Light-Rule Design.
Runtime responsibilities include:
- deserializing workflow definitions
- claiming active tasks
- executing supported task types
- storing task output
- applying exports into process context
- creating next tasks
- pausing waiting tasks
- resuming after human completion
- failing or completing process instances
- exposing workflow APIs to clients
The current executable slice supports API invocation and verification tasks such as HTTP, JSON-RPC, OpenRPC, MCP over enterprise HTTP transports, rules, assertions, and waiting human input. MCP stdio transport is intentionally not a priority for enterprise deployment.
Design Rule
There must be one workflow runtime and one task state model.
Chat, worklist, CLI, scheduler, and agents should never implement their own workflow execution. They should all use the same light-workflow service APIs.
This keeps enterprise workflow behavior auditable, testable, and consistent regardless of how a process is started, resumed, or observed.
Workflow Client Architecture
Light-Fabric workflow execution should run as a containerized backend service, not as logic embedded in a portal screen, CLI, scheduler, or agent. The workflow service owns process state, task state, audit records, API invocation, agent invocation, and human-in-the-loop transitions. Clients are thin interaction surfaces over the same service APIs.
This separation lets the same workflow instance be driven by a portal chat session, a worklist user, a CLI command, a scheduler, or an AI agent without creating multiple execution models.
Goals
- Provide one authoritative workflow runtime for long-running enterprise processes.
- Support human-in-the-loop tasks from both conversational and worklist interfaces.
- Support headless execution for live tests, scheduled runs, and CI/CD.
- Keep all clients stateless or lightly stateful; workflow state lives in
light-workflow. - Make role assignment, audit, and retry behavior consistent across UI, CLI, scheduler, and agent use.
Runtime Service
light-workflow should be deployed as a portal service in a container alongside the other portal services. It should expose APIs for workflow definitions, workflow instances, task claiming, task completion, event streaming, and operational control.
The service is responsible for:
- loading workflow definitions
- starting workflow instances
- persisting
process_info_tandtask_info_t - executing API calls and assertions
- invoking agents for agent-owned tasks
- pausing on
askand approval tasks - assigning human tasks to users or roles
- resuming workflows when a human answer is submitted
- emitting workflow and task events
- recording audit history
Clients should never execute workflow steps themselves. They should only start workflows, inspect workflow state, and complete assigned tasks.
Client Surfaces
Portal Chat
The portal chat client is the guided conversational interface for a single user working through a process. It is useful when the workflow needs to ask clarifying questions, explain the next action, or guide a user through a complex multi-endpoint operation.
Typical uses:
- API onboarding
- API endpoint publication to an AI gateway
- guided configuration
- troubleshooting and remediation workflows
- interactive approval with explanation
The chat client should call the workflow service for current state and submit answers to waiting tasks. It may stream workflow events and render agent explanations, but it should not own workflow state.
Worklist
The worklist is the enterprise task inbox. It is the right interface for multi-user coordination, role-based assignment, approvals, escalations, and audit-sensitive operations.
Typical uses:
- approval tasks
- compliance review
- operations handoff
- role-based queue processing
- task claim and release
- delegated work
- due-date and priority management
The worklist should be built around waiting human tasks. A task may have:
- assigned user
- candidate roles
- assigned role
- priority
- due time
- claim status
- comments
- completion payload
- audit trail
The worklist is especially important because many enterprise workflows are not purely conversational. They need accountable ownership and coordination between multiple humans.
CLI
The CLI is a developer and automation client. It should use the same workflow service APIs as portal-view and should not contain separate execution logic.
Typical uses:
- local workflow testing
- live parity tests
- CI/CD automation
- scheduled headless runs
- debugging stuck workflow instances
- submitting test data
- completing simple waiting tasks from scripts
Example commands:
light-workflow start portal.onboard-api --input input.yaml
light-workflow status <instance-id>
light-workflow tasks --role portal-admin
light-workflow claim <task-id>
light-workflow answer <task-id> --value approve
light-workflow logs <instance-id>
light-workflow cancel <instance-id>
The CLI should be added after the workflow APIs stabilize. It will be valuable for developers and automation, but the worklist and portal chat should drive the primary enterprise UX.
API Boundary
The workflow service should expose a stable API boundary that all clients use. The API can be HTTP, JSON-RPC, or both, but the concepts should remain the same.
Core operations:
workflow.start
workflow.getInstance
workflow.listInstances
workflow.getEvents
workflow.listTasks
workflow.getTask
workflow.claimTask
workflow.releaseTask
workflow.completeTask
workflow.delegateTask
workflow.cancelInstance
For streaming clients, the service should expose workflow events through Server-Sent Events, WebSocket, or another portal-standard event mechanism.
Important event types:
- workflow started
- task started
- task completed
- task failed
- task waiting for input
- task assigned
- task claimed
- task completed by human
- agent started
- agent completed
- workflow completed
- workflow failed
Human Task State
ask and approval-style tasks should enter a waiting state. While waiting, the workflow instance remains active, but the task is no longer executable by the worker loop until a human answer is submitted.
Recommended states:
A = active
W = waiting for input
C = completed
F = failed
X = canceled
The waiting task should include enough metadata for all clients:
- prompt
- input mode
- options
- validation rules
- default value
- sensitivity flag
- assignment metadata
- explanation metadata
- timeout policy
The completion API should validate submitted input against the task definition before resuming the workflow.
Assignment Model
Human tasks should support both direct assignment and role-based queues.
Recommended fields:
assigned_user
assigned_role
candidate_roles
claimed_by
claimed_ts
due_ts
priority
comments
A role-based task can appear in the worklist for all users with a matching role. Once a user claims it, the task becomes owned by that user until completed, released, delegated, or timed out.
Recommended Build Order
- Implement stable workflow service APIs for start, status, events, task list, task claim, and task completion.
- Harden the
askresume path and waiting task state machine. - Build the worklist because it forces the assignment, audit, and state model to be correct.
- Build the portal chat workflow interaction on top of the same task APIs.
- Add the CLI after the API shape stabilizes.
- Add scheduler integration for hourly live tests and headless workflow runs.
Design Rule
There must be one workflow runtime and one task state model. Chat, worklist, CLI, scheduler, and agents are only clients of that runtime.
This keeps enterprise workflow behavior auditable, testable, and consistent regardless of how a workflow is started or resumed.
LightAPI Description Design
lightapi-description-specification
LightAPI Description is the endpoint capability specification used by Light-Fabric agents, workflows, live tests, and portal API administration.
It describes how an API endpoint is discovered, invoked, explained, and verified. It is intentionally separate from the Agentic Workflow Specification. Workflow describes process orchestration. LightAPI describes endpoint capability.
Why LightAPI
OpenAPI is useful for REST APIs, and OpenRPC is useful for JSON-RPC APIs, but Light-Fabric needs a common description model across multiple enterprise protocols:
- REST / HTTP
- OpenAPI-described HTTP
- JSON-RPC 2.0
- OpenRPC-described JSON-RPC
- gRPC
- MCP tools, resources, and prompts
LightAPI provides a single agent-facing and workflow-facing description layer over these protocols.
The goal is not to replace OpenAPI or OpenRPC. The goal is to reference them where they exist and add the missing information needed by agents and workflow live tests.
API-Level Authoring, Endpoint-Level Consumption
Light-Portal may let teams author descriptions at the API level for convenience. However, workflows and agents consume descriptions at the endpoint level.
This distinction is important because real workflow processes rarely use a whole API. They usually combine selected endpoints from multiple APIs.
For example, onboarding an API to an AI gateway may consume:
- one endpoint from API registration
- one endpoint from API version management
- one endpoint from API instance management
- one endpoint from config server
- one endpoint from gateway linking
- one endpoint from controller reload
- one or more MCP tools exposed through the gateway
Each consumed operation should have an endpoint-level description with a stable endpointId.
API-level descriptions are still useful as catalogs. Endpoint-level descriptions may inherit shared API context such as:
- environments
- authentication
- secrets
- sources
- common tags
- lifecycle metadata
Relationship To Agentic Workflow
Agentic Workflow and LightAPI have different responsibilities.
| Concern | Agentic Workflow | LightAPI Description |
|---|---|---|
| Process order | Yes | No |
| Branching and retries | Yes | No |
| Human-in-the-loop | Yes | No |
| Endpoint invocation contract | Reference only | Yes |
| Input and result examples | Optional workflow fixtures | Yes |
| Result verification expectations | Calls assert | Describes expected result cases |
| Agent progressive disclosure | Uses selected endpoints | Defines disclosure levels |
| Live testing | Orchestrates execution | Supplies examples and expected results |
In live tests, the workflow should use example data from LightAPI descriptions and workflow fixtures instead of asking for user input.
In interactive runs, the workflow may ask the user for missing values, then invoke endpoints described by LightAPI.
Relationship To Centralized Agent Skills
LightAPI endpoint descriptions are a source of agent skills.
The centralized skill registry should not require every API operation to be manually rewritten as a separate skill. Instead, Light-Portal can publish selected LightAPI endpoint descriptions into the skill registry as invokable capabilities.
The skill registry adds:
- permission-aware discovery
- semantic search
- skill grouping
- agent persona scoping
- audit around skill disclosure and execution
LightAPI provides:
- endpoint identity
- protocol details
- input schema
- request mapping
- result shape
- examples
- behavior notes
- result cases
Together, they allow an agent to discover a capability as a skill, progressively load only the endpoint details it needs, and execute through the workflow or controller runtime.
See Centralized Agentic Skill Registry for the skill registry design.
Core Document Concepts
A LightAPI document should support both API-level catalogs and endpoint-level documents.
Important top-level concepts:
lightapi: specification versionprofile:apiorendpointinfo: name, title, version, namespace, owner, contactcontext: inherited catalog context for endpoint-level documentssources: OpenAPI, OpenRPC, protobuf, MCP, or raw protocol referencesenvironments: environment-specific server detailssecrets: required secret namesauthentications: reusable authentication policiesoperations: endpoint operation descriptionstestSequences: linear endpoint test sequencesagent: progressive disclosure and skill metadata
For profile: endpoint, the document should describe at most one operation.
Operation Model
Each operation represents one endpoint-level capability.
Common fields include:
operationId: local operation identifierendpointId: globally stable endpoint identifiertitlesummarydescriptionvisibilitylifecycletagscapabilityagentinputrequestresultexamples
The input section describes the logical interface the agent or workflow sees.
The request section describes how logical input maps to the wire protocol.
The result section describes expected output, result cases, and failure shapes.
Protocol Coverage
HTTP And OpenAPI
For raw HTTP, the operation describes method, endpoint, headers, query, path, and body mappings.
For OpenAPI, LightAPI references the OpenAPI document and operation, then adds agent-oriented behavior, examples, and result expectations.
JSON-RPC And OpenRPC
For direct JSON-RPC, the operation describes endpoint, method, params, id behavior, notification behavior, and error policy.
For OpenRPC, LightAPI references the OpenRPC document and method. The workflow runtime can use the OpenRPC document to validate that the method exists and that required params are present before calling it.
gRPC
For gRPC, the operation describes service, method, protobuf source, transport, metadata, request mapping, and result mapping.
For browser or gateway-mediated enterprise deployments, gRPC over WebSocket can be represented as a transport on the structured protocol operation.
MCP
For MCP, the operation describes tool, resource, or prompt invocation.
Tool listing alone is not enough. The description must also include:
- input schema
- result shape
- examples
- behavior differences for important input cases
- error cases
- verification expectations
MCP stdio is not a priority for enterprise portal deployment. HTTP and streamable HTTP transports should be the main runtime targets.
Result Cases And Verification
LightAPI should describe expected result behavior, but Agentic Workflow should execute the actual assertions.
This keeps verification orchestration in one place.
Recommended model:
- LightAPI operation result cases describe expected outputs, failure shapes, and examples.
- Workflow test steps invoke the operation.
- Workflow
asserttasks verify actual output against expected result cases. - Complex business checks can call the rule engine.
This allows the same endpoint description to support:
- agent skill usage
- workflow execution
- live integration testing
- failure diagnosis
Progressive Disclosure For Agents
A LightAPI document should support progressive disclosure so an agent can load only the information needed at each stage.
Recommended levels:
index: endpoint id, title, tags, visibilitysummary: purpose, capability group, lifecycleinvocation: input schema, request mapping, authentication, examplesbehavior: result cases, edge cases, errors, assertionsfull: complete endpoint description
The portal can expose query APIs such as:
lightapi.listOperations
lightapi.getOperation
lightapi.getCapabilityGroup
Agents should start with index or summary data, load invocation details only for selected endpoints, and load behavior details only for testing, troubleshooting, or failure repair.
Portal Publishing Flow
Light-Portal should manage endpoint descriptions as part of API endpoint administration.
Recommended flow:
- API owner creates or imports API metadata.
- Portal extracts initial endpoint descriptions from OpenAPI, OpenRPC, protobuf, MCP, or raw endpoint configuration.
- API owner enriches endpoint descriptions with examples, behavior notes, result cases, and visibility.
- Portal stores endpoint-level LightAPI descriptions.
- Authorized agents and workflows query descriptions by endpoint, tag, lifecycle, visibility, or capability.
- Selected endpoints can be published into the centralized skill registry.
- Workflow instances reference endpoint descriptions during execution and live testing.
Live Test Use
Live tests should be workflow-driven.
LightAPI supplies:
- example input data
- expected result cases
- protocol invocation details
- error behavior
Agentic Workflow supplies:
- sequence
- fixtures
- environment selection
- endpoint invocation
- assertions
- failure routing
- task creation
- agent assignment
This avoids building a second test runner model outside the workflow engine.
Design Rule
LightAPI describes endpoint capability. Agentic Workflow orchestrates endpoint use. Centralized Skills expose selected capabilities to agents.
Keeping these responsibilities separate lets Light-Fabric support API administration, agent skill discovery, workflow execution, and live integration testing without duplicating endpoint definitions across multiple systems.
Light-Rule Design
Light-Rule is the local YAML rule engine used by Light-Fabric services and workflows for deterministic business checks, transformations, authorization decisions, and workflow assertions.
It complements agentic workflow by keeping critical decisions explicit, repeatable, and auditable. Agents can propose or select rules, but the rule engine executes the deterministic logic.
Purpose
Light-Rule is designed for enterprise services that need fast local policy and transformation logic without a database call on every request.
Primary uses:
- fine-grained authorization
- request transformation
- response transformation
- response row and column filtering
- workflow assertions
- business validation
- permission and filter injection
- reusable rule templates selected from Light-Portal
The rule configuration is loaded locally by the target service. When permissions or rule mappings change, the controller can trigger a config reload so the service swaps to the latest rules.
Relationship To Agentic Workflow
Agentic Workflow orchestrates process steps. Light-Rule evaluates deterministic logic inside those steps.
Workflow uses Light-Rule in two main ways:
-
Rule call task A workflow task can call a named rule to validate or mutate workflow context.
-
Assert task extension Simple checks can be handled directly by
assert, while complex business checks can delegate to Light-Rule.
This separation keeps workflows readable. The workflow says when a check happens; Light-Rule defines the reusable business logic for the check.
Example workflow responsibilities:
- decide when authorization configuration is needed
- select or create a rule
- invoke a rule during live testing
- route failures to a human or agent
Example Light-Rule responsibilities:
- evaluate role, group, position, or attribute checks
- inject endpoint permissions into the context
- compute row or column filters
- execute transformation plugins
- return pass/fail for business assertions
See Agentic Workflow Design for the workflow orchestration model.
Relationship To LightAPI
LightAPI endpoint descriptions describe endpoint invocation and expected result behavior. Light-Rule can implement complex result checks that are too business-specific for simple schema assertions.
Recommended model:
- LightAPI describes endpoint result cases and expected behavior.
- Agentic Workflow invokes the endpoint and runs
asserttasks. asserthandles simple checks directly.- Light-Rule handles complex checks, authorization logic, row filters, column filters, and reusable business policies.
See LightAPI Description Design for endpoint capability descriptions.
Rule Specification
Rules are described by the rule specification in rule-specification/schema/rule.yaml.
The top-level configuration contains:
ruleBodies: named rule definitionsendpointRules: endpoint-to-rule mappings
Each rule can contain:
ruleIdruleDescruleTypeversionauthorupdatedAtconditionsconditionLanguageconditionSecurityProfileexpressionactions
Each endpoint mapping can contain:
req-tra: request transformation rulesres-tra: response transformation rulesreq-acc: request access rulesres-fil: response filter rulesaccess-control: legacy or compatibility access control rulespermission: permission values injected into contextx-*: extension rule phases
Rule Conditions
Conditions evaluate fields in the input context.
Light-Fabric supports CEL rule conditions only. A Light-Fabric rule uses
conditionLanguage: cel and a single boolean expression, with an optional
conditionSecurityProfile.
The native condition-row format with operand, operator, expected, and
joinCode is a legacy format from the Java yaml-rule implementation. It can be
documented for migration and compatibility, but it is not supported by
Light-Fabric rule execution. Keep the detailed CEL contract in
CEL Rule Conditions; this page documents how CEL fits into the
broader Light-Rule model.
Legacy Native Conditions
Supported operand forms:
- direct field:
role - dotted path:
user.role - JSON Pointer:
/user/role - JSONPath-like path:
$.user.roles[0]
Supported operators:
==
!=
>
<
>=
<=
eq
ne
contains
matches
startsWith
endsWith
exists
notExists
expected is typed and may be a string, number, boolean, array, object, or null.
Flat condition arrays are evaluated left-to-right. joinCode combines the current condition with the previous result.
A AND B OR C
is evaluated as:
(A AND B) OR C
If explicit grouping is required, split logic into multiple rules and combine them through endpoint mapping or workflow orchestration.
These native condition rows are included here only to document the legacy Java yaml-rule shape. New Light-Fabric rules must use CEL.
CEL Conditions
CEL rules use conditionLanguage: cel and store the predicate in expression.
The expression must evaluate to a boolean. If it evaluates to false, the rule
actions do not run.
ruleBodies:
allowOfferSearch:
ruleId: allowOfferSearch
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
auditInfo.subject_claims.ClaimsMap.role != null
&& toolArguments.category == "travel"
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
CEL expressions should be used for rule eligibility and business predicates.
Endpoint-specific role lists, row filters, column filters, and similar policy
values should still live under permission so API owners can change policy data
without editing the reusable rule body.
CEL must not be used as a general JSON mutation language. A CEL expression can
answer whether a rule applies, or whether a specific row should be kept when an
action provides a row-scoped CEL context. It should not directly rewrite
responseBody, add or remove fields, or execute side effects.
Rule Actions
Actions execute plugin logic after conditions pass.
An action contains:
actionIdactionClassNameactionValues
actionClassName identifies the registered plugin. actionValues carries plugin-specific configuration.
Typical action plugins:
- add values to request context
- inject permission attributes
- compute filters
- transform request body
- transform response body
- call a local business function
Actions are intentionally plugin-based so the schema remains stable while implementation logic can evolve.
For response filtering, standard actions remain the transformation boundary.
ResponseRowFilterAction and ResponseColumnFilterAction own row and column
mutation, shape-specific handling, failure behavior, and audit logging. If a use
case needs richer row predicates, add a CEL-aware action such as
ResponseCelRowFilterAction that evaluates a CEL predicate for each row. Do not
expose raw response-body mutation functions to CEL.
Response filtering should parse the response JSON once for the whole res-fil
phase, pass the mutable JSON value through the ordered action pipeline, and
serialize once at the end. Individual actions should not repeatedly parse and
serialize responseBody when multiple filters are configured on the same
endpoint.
Endpoint Rule Phases
Endpoint mappings define when rules run.
Request Transformation
req-tra rules run before the service handles the request. They can enrich or transform request context.
Response Transformation
res-tra rules run after the service produces a response. They can filter, redact, or reshape response data.
Request Access
req-acc rules validate whether a request is allowed before the target handler
or backend service runs. These rules normally run in parallel because they
should not mutate shared state.
access-control can be accepted as a compatibility phase by runtimes that need
to load older configuration, but new endpoint mappings should use req-acc.
Response Filtering
res-fil rules run after the target service produces a response and before the
caller receives it. They are used for row filters, column filters, response
masking, and other response-reduction policies that should not be implemented
inside the business API.
res-fil rules always execute as a sequential pipeline in the order listed on
the endpoint. accessRuleLogic applies only to req-acc; it does not change
response-filter execution semantics.
Permission Injection
permission values are injected into the evaluation context before rule execution. This lets API owners configure roles, groups, attributes, row filters, or column filters without editing the technical rule body.
Extension Phases
Custom phases must use the x-* prefix. This avoids silent typos in standard phase names while preserving controlled extensibility.
Execution Model
The Rust implementation lives in crates/light-rule.
Core components:
RuleConfig: top-level config modelRule: rule definitionRuleCondition: condition modelRuleAction: action modelRuleEngine: evaluates one ruleActionRegistry: maps action class names to pluginsMultiThreadRuleExecutor: executes rule lists and endpoint phase mappings
Sequential phases such as req-tra and res-tra should run with all semantics so transformations happen in order.
Access control can run in parallel because it should be a validation step rather than a mutation step.
Response filtering should run sequentially when multiple filters can depend on
the same fields. For example, a row filter that checks active == true must run
before a column filter that removes the active field from the final response.
Why Not Replace With Cedar Or Casbin
Cedar and Casbin are strong policy engines, but Light-Rule has a different role in this platform.
Light-Rule supports:
- local YAML configuration
- request and response transformation
- permission injection
- row and column filters
- endpoint-specific rule selection
- technical-team-authored reusable rules
- API-owner-selected rule parameters
- config reload through controller
Cedar is excellent for authorization policy, but it does not naturally cover transformation, row filter, and column filter use cases. Casbin is strong for policy enforcement, but it introduces a different policy storage and matching model.
Light-Rule should remain the built-in rule engine for Light-Fabric service configuration and workflow assertions. External policy engines can still be integrated as action plugins if needed.
Governance
Rule bodies should be authored and reviewed like code or controlled configuration.
Recommended governance metadata:
versionauthorupdatedAtruleDesc
Recommended operational controls:
- validate rule YAML against the schema before publishing
- reject endpoint phase typos
- keep
ruleIdequal to theruleBodiesmap key - audit rule publication and reload events
- test rules with representative input contexts
- use workflow live tests to verify rules in integrated environments
Workflow Live Testing
Light-Rule is useful in live tests because it can express business checks that are more specific than generic JSON assertions.
Example flow:
- Workflow invokes an endpoint using LightAPI description.
- Workflow captures the endpoint response.
assertverifies simple fields.- A rule task validates business-specific behavior.
- On failure, workflow creates a task for a human or agent to investigate.
This keeps live test orchestration in workflow while preserving reusable business rules in Light-Rule.
Design Rule
Use workflow for process control. Use LightAPI for endpoint capability. Use Light-Rule for deterministic business logic.
Agents may select, explain, or help author rules, but the rule engine should execute the final deterministic decision.
CEL Rule Conditions
Light-Fabric supports CEL rule conditions only. A Light-Fabric rule uses
conditionLanguage: cel and one rule-level CEL boolean expression.
The old native condition schema with condition rows, operators, and joinCode
is a legacy Java yaml-rule format. It can still be documented for migration and
Java compatibility, but it is not a supported Light-Fabric runtime condition
format.
Each Light-Fabric rule should therefore use CEL. Mixing native condition rows and CEL expressions inside the same rule is not a canonical model because it makes portal authoring, validation, and runtime dispatch harder to reason about.
Goals
- support CEL expressions as the Light-Fabric rule-level condition language
- reuse the existing rule context for gateway, workflow, and test execution
- preserve existing
actions,endpointRules, and rule phase semantics - let Light-Portal choose the correct editor from rule metadata without parsing arbitrary rule bodies
- validate CEL before publishing or reloading rules where possible
- keep CEL execution deterministic and side-effect free
Non-Goals
- replacing actions with CEL
- allowing CEL expressions to perform I/O, network calls, mutation, or service lookups
- allowing CEL expressions to directly mutate
responseBodyor perform general JSON transformations - supporting the legacy Java yaml-rule native condition-row format in Light-Fabric
- supporting mixed native and CEL condition blocks in the canonical portal authoring flow
Current Model
The legacy Java yaml-rule model contains an optional flat list of native conditions:
ruleBodies:
allowMcpReader:
common: Y
ruleId: allowMcpReader
ruleName: Allow MCP reader
ruleType: req-acc
conditions:
- operatorCode: isNotNull
propertyPath: auditInfo.subject_claims.ClaimsMap.role
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
Each legacy native condition contains:
operatoroperandexpectedjoinCode
The Java yaml-rule engine evaluates conditions left-to-right. joinCode
combines each condition with the accumulated result. This format is shown here
only as migration context.
Portal persistence stores rule metadata in rule_t and the executable rule JSON
in rule_t.rule_body. Today there is no dedicated column that tells the portal
which condition editor to render, so the UI would have to inspect rule_body.
Proposed Rule Shape
Use a rule-level condition language flag. Light-Fabric accepts cel for a
single CEL expression. native is reserved for legacy Java yaml-rule data and
must not be emitted to Light-Fabric runtime configuration.
Persist the flag in both places:
rule_t.condition_language: indexed/listable portal metadataruleBody.conditionLanguage: self-contained exported runtime configuration
Recommended Light-Fabric value:
cel
Legacy Java yaml-rule native rule body:
ruleBodies:
allowMcpReader:
common: Y
ruleId: allowMcpReader
ruleName: Allow MCP reader
ruleType: req-acc
conditionLanguage: native
conditions:
- operatorCode: isNotNull
propertyPath: auditInfo.subject_claims.ClaimsMap.role
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
CEL rule body:
ruleBodies:
allowApprovedTransfer:
common: Y
ruleId: allowApprovedTransfer
ruleName: Allow approved transfer
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
auditInfo.subject_claims.ClaimsMap.role != null
&& 'roles' in permission
&& permission.roles != null
&& permission.roles.exists(r, r == auditInfo.subject_claims.ClaimsMap.role)
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
Recommended database shape:
ALTER TABLE rule_t
ADD COLUMN condition_language VARCHAR(16) DEFAULT 'cel' NOT NULL;
ALTER TABLE rule_t
ADD COLUMN condition_security_profile VARCHAR(32);
ALTER TABLE rule_t
ADD CONSTRAINT rule_t_condition_language_check
CHECK (condition_language IN ('cel'));
ALTER TABLE rule_t
ADD CONSTRAINT rule_t_condition_security_profile_check
CHECK (
condition_security_profile IS NULL
OR condition_security_profile IN ('strict', 'standard', 'internal-admin')
);
Recommended schema rules:
conditionLanguageis optional and defaults tocelconditionLanguage: celrequiresexpressionand rejectsconditionsconditionSecurityProfileis optional and names a runtime-defined profileconditionLanguage: nativeis rejected by Light-Fabric runtime config- unknown rule and condition fields should continue to be rejected by the schema
- command handlers should reject requests where the DB metadata and rule body condition language disagree
This can be represented with conditional validation in
rule-specification/schema/rule.yaml:
allOf:
- if:
properties:
conditionLanguage:
const: cel
required: [conditionLanguage]
then:
required: [expression]
not:
required: [conditions]
else:
properties:
conditionLanguage:
const: cel
The Rust model can add optional fields to Rule:
#![allow(unused)]
fn main() {
pub condition_language: Option<String>,
pub condition_security_profile: Option<String>,
pub expression: Option<String>,
}
This is less disruptive than changing RuleCondition into an enum and keeps old
rule bodies valid.
Cross-Repository Scope
This change crosses the rule specification, runtime engines, portal services, and
portal UI. The implementation should be tracked as a coordinated change rather
than a light-fabric-only feature.
| Area | Required work |
|---|---|
rule-specification | Add conditionLanguage, conditionSecurityProfile, expression, CEL rule schema validation, and explicit rejection of native condition rows for Light-Fabric runtime config. |
portal-db | Add rule_t.condition_language with default cel, optional rule_t.condition_security_profile, check constraints, and pending rule-change approval state if workflow task payloads are not sufficient. |
light-portal | Update persistence and projection code so rule create/update/read/export/import paths carry conditionLanguage and conditionSecurityProfile; ensure endpoint rule config generation emits only approved, self-contained rule bodies; integrate stronger-profile requests with worklist and assistant-task approval. |
rule-command | Accept conditionLanguage, conditionSecurityProfile, and expression, reject native condition-row payloads for Light-Fabric rules, validate mode/profile-specific shape, publish strict changes immediately, route stronger profile requests through approval, and write both DB metadata and rule body consistently after approval. |
rule-query | Return conditionLanguage, conditionSecurityProfile, and approval status for list/detail APIs, include selected/effective profiles in test-case execution payloads, and surface CEL parse/type/missing-field/profile errors from Java and Rust runners. |
portal-view | Render the CEL expression editor for Light-Fabric rules; keep any native condition builder scoped to legacy Java yaml-rule authoring; show a controlled profile selector for CEL rules; submit strict directly and route standard or internal-admin to worklist approval; do not require the UI to infer mode from ruleBody. |
| workflow and assistant task | Use the existing human-in-the-loop worklist flow for stronger profile approval, route tasks to admin and rule-admin, and attach an advisory assistant-task risk summary for the approver. |
light-fabric | Add conditionLanguage, conditionSecurityProfile, and expression to crates/light-rule, dispatch in RuleEngine, add policy-driven CEL evaluator/caching, and update gateway/workflow tests. |
yaml-rule | Add Java runtime parity for conditionLanguage: cel and named profile enforcement if Java services need to execute the same rules; otherwise reject CEL rules explicitly with a clear runtime-capability error. |
portal-db is listed even though it is not a rule engine because rule_t lives
there. Without the DB column, portal-view would need to parse the compact rule
body to choose the editor, which is the coupling this design is trying to avoid.
Operator Alias Alternative
Another possible shape is to add operatorCode: cel and store the CEL
expression in expected inside conditions:
conditions:
- operatorCode: cel
expected: >
context.toolArguments.amount < 1000
|| ('roles' in context.permission
&& context.permission.roles != null
&& context.permission.roles.exists(r, r == "approver"))
This has one advantage for legacy Java yaml-rule imports: operator, operand,
and expected already exist. It is not useful as the Light-Fabric runtime
contract because Light-Fabric does not support native condition rows.
It should not be the canonical schema because:
- CEL is a full boolean expression, not a comparison operator
- overloading
expectedmakes validation and portal rendering less clear operandbecomes ignored or artificial- the UI still has to draw a condition-row editor even though the rule is really a single expression
- future expression languages would continue overloading legacy native condition fields
The recommended contract is therefore:
- canonical form:
conditionLanguage: celplus rule-levelexpression - reject
operatorCode: celfor Light-Fabric runtime config - normalize any legacy import to the canonical rule-level CEL model before persistence or runtime export
Mixed Conditions Alternative
Another possible shape is to allow native and CEL conditions in the same
conditions array. Light-Fabric should not support this. Native condition rows
belong to the legacy Java yaml-rule model only.
Reasons to avoid canonical mixed rules:
- Light-Portal would need a hybrid editor that switches row-by-row
- validation errors become harder to explain to non-technical users
joinCodesemantics across native and CEL expressions are correct but subtle- users may expect CEL operator precedence inside the whole rule even though
native
joinCoderemains left-to-right - runtime dispatch is simpler and faster when the rule selects one evaluator
If mixed rules are accepted from an import path, they must be normalized to a single rule-level CEL expression before they are persisted or exported to Light-Fabric runtime configuration.
Execution Model
Rule execution should dispatch by conditionLanguage once per rule:
RuleEngine::execute_rule
-> conditionLanguage == cel
-> evaluate rule expression
-> execute actions when conditions pass
The outer behavior stays unchanged:
- rules with no conditions continue to run actions
- CEL rules without an expression fail validation before runtime
- failed conditions skip actions
- failed action execution fails the rule
- endpoint rule ordering and access-control logic stay unchanged
req-traandres-tracontinue to run sequentially- access-control rules can still be evaluated independently
Runtime should treat a missing conditionLanguage as cel only when an
expression is present. Native condition-row payloads must be rejected by
Light-Fabric config validation.
CEL And Response Filtering
CEL is the rule predicate language, not the response mutation engine. A rule-level
CEL expression decides whether a res-fil rule applies. The response body is
then transformed by a standard action.
This keeps the contract narrow:
- CEL expressions are side-effect free and return booleans.
- Actions own mutation of
responseBody. - Actions decide how JSON arrays, JSON objects, and malformed payloads are handled.
- Actions provide stable audit and failure behavior.
- The runtime can compile and cache rule-level CEL independently from response-body parsing.
- The response-filter pipeline parses JSON once, lets actions mutate the same
in-memory value, and serializes once after all
res-filactions complete. - The parsed mutable response value is action-owned state and must not be exposed as a rule-level CEL mutation target.
The default response-filter actions should stay declarative:
ResponseRowFilterAction: applies permission-defined row filters.ResponseColumnFilterAction: applies permission-defined column keep or remove lists.
Rule bodies should keep using the actions[].actionClassName field even when
the runtime is Rust. In Rust this value is not a Java class name. It is a stable
action registry key that selects the Rust action implementation. The Rust
gateway registers both the Java-compatible fully qualified names and short
aliases:
com.networknt.rule.ResponseRowFilterActionResponseRowFilterActioncom.networknt.rule.ResponseColumnFilterActionResponseColumnFilterActioncom.networknt.rule.ResponseCelRowFilterActionResponseCelRowFilterAction
For portal-authored and exported rules, prefer the fully qualified Java-compatible names. They preserve compatibility with existing yaml-rule configuration, schemas, import/export flows, and any Java runtime that reads the same rule bodies:
ruleBodies:
rowFilterByJwtClaims:
common: Y
ruleId: rowFilterByJwtClaims
ruleName: Row filter by JWT claims
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
row != null
&& (
("role" in row && "role" in auditInfo.subject_claims.ClaimsMap)
|| ("group" in row
&& ("grp" in auditInfo.subject_claims.ClaimsMap
|| "group" in auditInfo.subject_claims.ClaimsMap))
|| ("position" in row
&& ("pos" in auditInfo.subject_claims.ClaimsMap
|| "position" in auditInfo.subject_claims.ClaimsMap))
|| ("attribute" in row
&& ("att" in auditInfo.subject_claims.ClaimsMap
|| "attribute" in auditInfo.subject_claims.ClaimsMap))
)
actions:
- actionClassName: com.networknt.rule.ResponseRowFilterAction
colFilterByJwtClaims:
common: Y
ruleId: colFilterByJwtClaims
ruleName: Column filter by JWT claims
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
col != null
&& (
("role" in col && "role" in auditInfo.subject_claims.ClaimsMap)
|| ("group" in col
&& ("grp" in auditInfo.subject_claims.ClaimsMap
|| "group" in auditInfo.subject_claims.ClaimsMap))
|| ("position" in col
&& ("pos" in auditInfo.subject_claims.ClaimsMap
|| "position" in auditInfo.subject_claims.ClaimsMap))
|| ("attribute" in col
&& ("att" in auditInfo.subject_claims.ClaimsMap
|| "attribute" in auditInfo.subject_claims.ClaimsMap))
)
actions:
- actionClassName: com.networknt.rule.ResponseColumnFilterAction
The rule-level CEL expression only decides whether the response-filter action
runs. The action reads the endpoint permission.row or permission.col
configuration and matches it against JWT claims. The response-filter action
understands these permission dimensions:
role: matched against the JWTroleclaimgroup: matched againstgrporgroupposition: matched againstposorpositionattribute: matched againstattorattributeuser: matched againstuid,user_id, orsub
For example:
endpointRules:
/v1/accounts@get:
res-fil:
- rowFilterByJwtClaims
- colFilterByJwtClaims
permission:
row:
role:
manager:
- colName: status
operator: =
colValue: ACTIVE
group:
finance:
- colName: department
operator: =
colValue: FIN
position:
director:
- colName: level
operator: ">="
colValue: "5"
attribute:
region-east:
- colName: region
operator: =
colValue: EAST
col:
role:
manager: id,name,status,department
group:
finance: id,name,balance,status
position:
director: id,name,balance,status,level
attribute:
region-east: id,name,region,status
If a row predicate needs CEL, add an explicit CEL-aware action rather than turning rule-level CEL into a JSON transformation DSL:
ruleBodies:
filterOfferRows:
ruleId: filterOfferRows
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200 && responseBody != ""
actions:
- actionClassName: com.networknt.rule.ResponseCelRowFilterAction
actionValues:
rowExpression: >
auditInfo.subject_claims.ClaimsMap.role == "offer-admin"
|| (row.priority < 50 && row.active == true)
ResponseCelRowFilterAction would compile rowExpression at rule-load time and
evaluate it once per candidate row with a curated context containing row,
auditInfo, headers, endpoint, permission, and request metadata. The
implementation should avoid deep-cloning the full base context for every row.
Use a child CEL context that shadows row, or reuse one mutable evaluation
context and replace only the row variable before each evaluation.
Row-level CEL evaluation errors should be row-local by default. If a row is
missing a referenced field or its predicate evaluation returns an error, drop
that row and emit debug or trace diagnostics with the rule id and field error.
Only configuration errors, such as an invalid rowExpression that fails to
compile, should fail the entire action closed.
Column filtering should remain declarative unless there is a proven use case for dynamic column predicates. Role, group, attribute, user, and position based field lists are easier to review, safer to render in Portal, and cheaper to execute than arbitrary column-level CEL.
Column filtering must also support top-level JSON objects, not only arrays or
objects containing items. Single-object responses such as GET /offers/123
still need field hiding. Row filtering also treats a top-level JSON object as a
single candidate row. HTTP filtering replaces a denied object with an empty
object, while MCP filtering returns a tool error with isError: true.
Rule Context
CEL should evaluate against the rule engine JSON context. For gateway access-control and response filtering, this includes fields such as:
auditInfoheadersendpointtoolNametoolArgumentscorrelationIdresponseBodystatusCode
Endpoint permission values are merged into the root context as their configured
keys. For example, permission.roles in endpointRules is available to
conditions as roles, response row filters are available as row, and column
filters are available as col. A future runtime can also expose a namespaced
permission object as an additive convenience, but CEL support should not
require that shape to preserve compatibility with existing actions.
For standard and internal-admin profiles, the CEL environment can expose
variables in two ways:
- top-level context fields as direct CEL variables, such as
auditInfo,toolArguments, androles - the full root object as
context, so expressions can use explicit paths such ascontext.toolArguments.amount
Direct variables keep expressions concise. The context variable is safer for
generated expressions, collision avoidance, and future fields that are not valid
CEL identifiers.
For the strict profile, the runtime should expose only curated root variables
such as auditInfo, headers, toolArguments, endpoint metadata, and
permission values needed by the rule phase. It should not expose the full
context object by default. This prevents future internal runtime metadata from
becoming visible to tenant-authored CEL just because it was appended to the root
request context.
The context contract should be documented as part of Light-Rule because CEL expressions depend on stable field names. Adding fields is compatible. Renaming or changing field shapes is a breaking change for CEL rules.
Type Mapping
The CEL evaluator should receive deterministic values converted from
serde_json::Value:
- JSON object to CEL map
- JSON array to CEL list
- JSON string to CEL string
- JSON number to CEL integer or double
- JSON boolean to CEL bool
- JSON null to CEL null
Missing fields should evaluate according to the chosen CEL implementation’s standard behavior. The rule test API should expose these failures clearly so authors can distinguish “expression false” from “expression invalid”.
Authors should guard optional fields explicitly. Depending on the selected CEL
runtime and the field shape, this can use presence checks such as has(...) or
map membership checks such as:
"role" in auditInfo.subject_claims.ClaimsMap
&& auditInfo.subject_claims.ClaimsMap.role == "admin"
The portal rule tester should surface missing-field evaluation errors and suggest guarded expressions instead of letting these failures look like ordinary denied rules.
Context Injection Performance
CEL expressions run on request paths, so context conversion must be controlled. The implementation should not recursively deep-clone and convert large JSON payloads separately for every CEL rule evaluation.
Recommended approach:
- compile expressions once at rule load
- build the rule context once per request or response phase
- reuse converted CEL variables across evaluations in the same request or response phase when possible
- prefer lazy or reference-backed variable resolution if the selected CEL crate supports it
- if eager conversion is required, convert only the variables exposed to CEL and
avoid parsing large string fields such as
responseBodyunless an expression explicitly needs structured access to them - for per-row CEL, avoid cloning the full base context for each row; use a child
context or reusable mutable context that changes only the
rowbinding - benchmark access-control and response-filter scenarios before enabling CEL by default in high-throughput paths
The initial implementation can be pragmatic, but performance tests should guard against accidentally making CEL expression evaluation proportional to the full response body size when the expression only needs claims or endpoint metadata.
Validation
CEL should be validated earlier than request execution.
Recommended validation points:
- portal rule editor
- rule command create/update handler
- rule test API
- runtime config reload
Validation must enforce the Light-Fabric rule shape:
cel:expressionis required,conditionsis rejectednative: rejected by Light-Fabric runtime config- persisted
rule_t.condition_languagemust matchruleBody.conditionLanguage - persisted
rule_t.condition_security_profilemust matchruleBody.conditionSecurityProfilewhen either side is present
Runtime reload should reject invalid CEL when strict validation is enabled. If a service must preserve availability, it can keep the last known-good rule set and report the new config as rejected.
Approval workflow should not bypass validation. For profile escalation requests, the command path should validate the submitted rule shape and expression before creating the approval task. Final approval should revalidate the exact submitted rule body before emitting the active rule event.
Validation output should include:
- rule id
- condition language
- parse or type error
- source offset when provided by the CEL implementation
Compilation And Caching
Do not compile CEL on every request. Compile once per rule load and cache the compiled program with the loaded rule set.
Recommended cache key:
ruleId + expression hash + effective profile
The compiled expression cache should be replaced atomically when the rule config reloads. It should not outlive the rule version it was compiled from. Old compiled entries must be evicted during reload so repeated rule updates cannot leak memory through stale expression hashes.
Rust CEL Library
Light-Rule uses the cel crate for the Rust implementation. It provides
Program::compile(...), Program::execute(...), a Context for variables and
functions, and compiled Program values that are Send + Sync.
Implementation should still be isolated behind a small internal trait:
CelEvaluator
-> compile(ruleId, expression) -> compiled expression
-> evaluate(compiled expression, serde_json::Value context) -> bool
This keeps Light-Rule from leaking third-party crate types through its public model and allows the implementation to change if CEL crate maturity, feature flags, or Java parity requirements change.
Legacy Operator Migration
Legacy Java yaml-rule native conditions include operators that may not map one-to-one to the selected CEL runtime. Examples include:
containsIgnoreCasematchesandnotMatchinListandnotInListcontainsAny,containsAll, andcontainsNone- date-style comparisons such as
before,after, andon
Before importing legacy native rules into Light-Fabric, the implementation should define a small compatibility function registry for any gaps and convert the rule to CEL. Candidate pure helper functions include:
contains_ignore_case(value, substring)
matches(value, pattern)
in_list(value, values)
contains_any(value, values)
contains_all(value, values)
These functions must be deterministic, side-effect free, and shared by the rule tester and runtime evaluator. If Java parity is required, the same function names and edge-case behavior should be implemented in the Java runtime.
Safety
CEL support should be deterministic and sandboxed.
The evaluator does not need an operating-system sandbox for normal trusted/admin-authored rule configuration. CEL is an interpreted expression language, not arbitrary Rust or JavaScript execution, and expressions can only resolve variables and functions registered in the CEL context. The CEL context is therefore the primary sandbox boundary.
For the Rust cel integration, context construction should be explicit.
Context::default() exposes standard pure CEL functions such as
size, contains, string helpers, type conversions, regex matches, and time
parsing helpers depending on enabled crate features. If a service accepts
tenant-authored or otherwise untrusted CEL, prefer Context::empty() and add
only platform-approved helper functions.
Security policy should be engine-owned. A rule may request a named condition security profile, but it must not define its own function allowlist, size limits, resource limits, or isolation mode. If a rule author controls the rule body, then inline security settings are also attacker-controlled.
Recommended policy model:
runtime config defines profiles:
strict
standard
internal-admin
rule optionally requests:
conditionSecurityProfile: strict
effective policy:
runtime maximum profile intersected with requested profile
If a rule omits conditionSecurityProfile, the runtime default applies. If a
rule requests a profile that the service, tenant, or rule phase does not allow,
the rule config should be rejected during validation or runtime reload. The
engine may choose a stricter profile than requested, but it must never choose a
weaker one because the rule requested it.
Recommended profiles:
strict: default for tenant-authored, portal self-service, imported, or marketplace-style CEL. Use an empty CEL context, expose only approved variables, add only pure helper functions, and enforce tight size and expression-shape limits. Do not expose the fullcontextroot, and disable regex until both Java and Rust provide matching bounded or linear-time behavior.standard: default for internal business rules. Keep allowlists and resource limits, but permit common pure helpers such assize,contains,startsWith,endsWith,contains_ignore_case, and bounded regex support if needed.internal-admin: limited to trusted operator-maintained rules. This may be closer to the selected CEL runtime’s default behavior, but should still compile during rule load, validate references, enforce maximum input size, and protect reloads with the last known-good rule set.
Allowed:
- boolean logic
- comparisons
- arithmetic supported by the CEL implementation
- string operations
- list and map predicates
- approved pure helper functions
Not allowed:
- file access
- network access
- database access
- current time unless explicitly added as an input field
- random values
- mutation of the rule context
- action execution from inside CEL
- response-body mutation or field removal from CEL
Custom functions should be added conservatively. Standard Light-Rule actions remain the extension point for side effects and transformations.
The core runtime object should be a policy-driven condition evaluator rather
than ad hoc logic embedded directly in RuleEngine:
RuleEngineOptions
-> ConditionExecutionPolicy
-> defaultCelProfile
-> allowRuleProfileSelection
-> profiles[name] = CelSecurityProfile
CelSecurityProfile
-> allowedFunctions
-> allowedRootVariables
-> exposeContextRoot
-> exposeTopLevelAliases
-> maxExpressionBytes
-> maxContextBytes
-> maxStringBytes
-> maxCollectionItems
-> allowRegex
-> allowTimeParsing
-> allowComprehensions
-> maxComprehensionNesting
CEL still needs resource and robustness controls because expressions run on request paths and can iterate over input data. Runtime and publish-time validation should:
- allow-list functions and variables, using compiled expression references where available
- reject functions that perform I/O, mutation, service lookup, action execution, random generation, or implicit current-time access
- cap expression length and input context size
- reject or limit expensive access to large request or response bodies
- compile during rule load and fail invalid expression shapes before request execution
- keep the last known-good rule set if reload validation fails
Phase ceilings should be enforced by runtime policy. Response phases such as
res-tra and res-fil should default to a strict ceiling or tight
maxContextBytes limits because they can include large response payloads.
Access-control phases may allow standard only when the exposed context is
small and bounded. A rule request for a stronger profile than the phase ceiling
must be rejected or downgraded to the stricter effective profile.
For fully untrusted public input, evaluate CEL in a separate worker, process, or another resource-isolated execution path with CPU and memory limits. A Tokio timeout alone is not a complete guard for synchronous CPU-bound expression evaluation.
Portal Experience
Light-Portal should use conditionLanguage to choose the rule editor. For
Light-Fabric, the only supported editor is the CEL editor. Any native condition
builder must be scoped to legacy Java yaml-rule authoring and must not export
native condition rows to Light-Fabric runtime config.
Recommended authoring modes:
CEL: advanced text area for one rule-level CEL expression.Builder: legacy Java yaml-rule-only condition rows with operand, operator, expected, and join controls.
Recommended behavior:
- default new Light-Fabric rules to
cel - render a CEL expression text area only for
conditionLanguage: cel - reject
conditionLanguage: nativefor Light-Fabric rule publishing - require confirmation when switching modes if the existing mode has content
- do not try to round-trip arbitrary CEL into native builder rows
- store the selected mode in
rule_t.condition_languageand in the JSON rule body asconditionLanguage - for CEL rules, store only the selected profile name in
rule_t.condition_security_profileand in the JSON rule body asconditionSecurityProfile; do not expose raw policy limits in the form - do not show
internal-adminin standard self-service forms; allow it only through checked-in runtime configuration or an explicitly authorized internal admin JWT/role path
The CEL editor should provide:
- syntax validation
- test context input
- expression result preview
- visible context field reference
- selected and effective security profile display
- rule test execution against the same backend evaluator used by runtime
Profile Approval Workflow
Light-Portal may allow a user to select a CEL security profile, but the selected profile is only a request. Runtime policy still computes the effective profile from the requested profile, the service maximum, the tenant maximum, and the rule phase ceiling.
Recommended publish behavior:
strict: direct publish. If schema, CEL validation, and command authorization pass, create or update the rule immediately.standard: approval required. Submit the proposed rule change, create a worklist task forrule-adminandadmin, and keep the change pending until approval.internal-admin: hidden from standard self-service authoring. If exposed to an operator-only flow, require stronger approval and never allow ordinary self-service users to request it.
For approval-required changes, the command side should not emit the final active
RuleCreated or RuleUpdated event at submission time. It should emit a
submission event such as RuleChangeSubmittedEvent or
RuleApprovalRequestedEvent, store the proposed rule body and requested profile,
and create the human-in-the-loop worklist task. Only approval should emit the
active rule event. Rejection should emit a rejection event and leave the active
rule unchanged.
Assistant tasks can help the approver by summarizing the CEL expression, rule phase, requested profile, referenced context roots, use of response body fields, regex usage, and any runtime ceiling that would downgrade the effective profile. The assistant output is advisory only; the human approver remains responsible for the approval decision.
Recommended approval rules:
- changing the expression, action list, rule phase, requested profile, or exposed context assumptions invalidates prior approval
- downgrading from
standardtostrictcan publish directly after validation - upgrading from
stricttostandardorinternal-adminrequires approval - requester and approver should be different users except for an explicit break-glass workflow
- approval audit should record requested profile, effective profile, requester, approver, approval time, assistant-task summary id, and approval comments
- pending rules must not be exported to runtime endpoint rule config until approved
Compatibility
Existing Light-Fabric rule YAML must use CEL conditions.
Rules without conditionLanguage can be treated as cel only when a valid
expression is present. Rules containing legacy native conditions must be
rejected by Light-Fabric runtime config validation or converted to CEL before
publish. The database migration should add rule_t.condition_language with
default cel.
Rules without conditionSecurityProfile use the runtime default CEL profile.
The field is meaningful only for CEL rules.
Native condition aliases are legacy Java yaml-rule import details only:
operatorCodeas alias foroperatorpropertyPathas alias foroperandactionClassNameas alias foractionRef
CEL is the Light-Fabric capability. If the Java yaml-rule runtime needs to
execute the same rules, it must implement the same CEL rule shape. Until then,
Java runtimes must fail closed with a clear capability error, such as
UnsupportedConditionLanguageException, when loading or executing a rule with
conditionLanguage: cel. A runtime must not silently ignore a CEL rule because
that can fail open for access-control rules.
Java parity is feasible because Google maintains CEL-Java under the dev.cel
Maven group, including the dev.cel:cel artifact with compiler and runtime
APIs. The compatibility requirement is therefore mostly about aligning the rule
schema, context shape, custom functions, and error handling across the Rust and
Java runtimes.
Example: Access Control
ruleBodies:
allowEndpointClaims:
common: Y
ruleId: allowEndpointClaims
ruleName: Allow request when endpoint permission matches JWT claims
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
(
!("role" in permission)
|| (
("roles" in auditInfo.subject_claims.ClaimsMap
&& permission.role in auditInfo.subject_claims.ClaimsMap.roles)
|| ("role" in auditInfo.subject_claims.ClaimsMap
&& permission.role == auditInfo.subject_claims.ClaimsMap.role)
)
)
&& (
!("group" in permission)
|| (
("groups" in auditInfo.subject_claims.ClaimsMap
&& permission.group in auditInfo.subject_claims.ClaimsMap.groups)
|| ("scp" in auditInfo.subject_claims.ClaimsMap
&& permission.group in auditInfo.subject_claims.ClaimsMap.scp)
|| ("group" in auditInfo.subject_claims.ClaimsMap
&& permission.group == auditInfo.subject_claims.ClaimsMap.group)
|| ("grp" in auditInfo.subject_claims.ClaimsMap
&& permission.group == auditInfo.subject_claims.ClaimsMap.grp)
)
)
&& (
!("position" in permission)
|| (
("positions" in auditInfo.subject_claims.ClaimsMap
&& permission.position in auditInfo.subject_claims.ClaimsMap.positions)
|| ("position" in auditInfo.subject_claims.ClaimsMap
&& permission.position == auditInfo.subject_claims.ClaimsMap.position)
|| ("pos" in auditInfo.subject_claims.ClaimsMap
&& permission.position == auditInfo.subject_claims.ClaimsMap.pos)
)
)
&& (
!("attribute" in permission)
|| (
("attributes" in auditInfo.subject_claims.ClaimsMap
&& permission.attribute.key
in auditInfo.subject_claims.ClaimsMap.attributes
&& auditInfo.subject_claims.ClaimsMap.attributes[
permission.attribute.key
] == permission.attribute.value)
|| ("attribute" in auditInfo.subject_claims.ClaimsMap
&& permission.attribute.value
== auditInfo.subject_claims.ClaimsMap.attribute)
|| ("att" in auditInfo.subject_claims.ClaimsMap
&& permission.attribute.value == auditInfo.subject_claims.ClaimsMap.att)
)
)
actions: []
endpointRules:
/v1/claims/{claimId}@post:
req-acc:
- allowEndpointClaims
permission:
role: claims-approver
group: claims.write
position: adjuster
attribute:
key: region
value: east
The endpoint above matches a caller JWT with claims like:
{
"roles": ["claims-approver"],
"groups": ["claims.write"],
"positions": ["adjuster"],
"attributes": {
"region": "east"
}
}
With this reusable req-acc rule, the technical rule body stays stable and API
owners define the required authorization dimensions at the endpoint. The example
above allows the request only when all configured endpoint permissions match
claims from the caller JWT.
Example: Response Filter Guard
ruleBodies:
filterAccountsForPortalUsers:
common: Y
ruleId: filterAccountsForPortalUsers
ruleName: Filter accounts for portal users
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200
&& responseBody != ""
&& auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.ResponseRowFilterAction
Rollout Plan
- Add
rule_t.condition_languagewith defaultcel, optionalrule_t.condition_security_profile, and check constraints. - Extend the rule specification with CEL rule validation plus optional
conditionSecurityProfile, and reject native condition rows for Light-Fabric runtime config. - Add
conditionLanguage,conditionSecurityProfile, andexpressionfields to the RustRulemodel. - Update command/query APIs so the portal can persist and read the condition
language, security profile, and approval state without parsing
ruleBody. - Reject
operatorCode: celin runtime config and normalize any legacy import to the rule-level CEL shape before publishing. - Choose and pin the Rust CEL crate behind an internal evaluator abstraction.
- Add runtime-owned CEL security profiles and policy-driven context building.
- Add approval workflow integration for
standardandinternal-adminprofile requests, including worklist and assistant-task support. - Dispatch inside
RuleEngine::execute_rulebased onconditionLanguage. - Compile and cache CEL expressions during rule config load.
- Add unit tests for CEL true, CEL false, invalid expression, mode validation, and missing-field behavior.
- Add tests for custom legacy-operator compatibility helper functions.
- Add performance tests for context conversion with large
toolArgumentsand response payloads. - Add gateway integration tests using the existing rule context and the
contextroot variable. - Add rule test API support so Light-Portal can validate CEL before publish.
- Add CEL rule editing, a controlled CEL profile selector, and approval UX for stronger profile requests.
- Document runtime compatibility and Java parity requirements.
Decision
Support CEL conditions as the only Light-Fabric rule condition language. Native
condition rows remain a legacy Java yaml-rule format and must not be emitted to
Light-Fabric runtime configuration. A Light-Fabric rule should use
conditionLanguage: cel; mixed native/CEL condition arrays are not a supported
authoring or runtime model.
Debugging CEL Rules
Status
The short-term referenced-context trace logging described here is implemented. Structured decision outcomes and the rule-test API remain proposed long-term work.
Problem
The Light-Gateway access-control handler and MCP router both evaluate CEL rules through the shared access-control runtime. When a request is denied, a rule author usually sees only a generic access-denied response. That response does not distinguish among these cases:
- the CEL expression evaluated successfully and returned
false - CEL compilation or evaluation failed
- the expression returned a non-boolean value
- the endpoint referenced a missing rule
- an action rejected the rule after its CEL condition matched
accessRuleLogic: alloraccessRuleLogic: anyproduced the final denialdefaultDenyapplied because no endpoint orreq-accrule matched- MCP
tools/listhid a tool because of policy, an unknown rule, or themaxCelEvaluationslimit
These cases must remain fail-closed, but they should not be indistinguishable to an authorized operator or rule author.
Printing the complete request on every denial is not an acceptable solution. Access control runs after security, so its CEL context should contain normalized identity claims and policy inputs rather than raw authentication credentials. Diagnostics can then project only the properties referenced by the expression instead of copying unrelated claims, headers, tool arguments, request data, or response data.
Current Behavior
The current implementation already provides a useful starting point:
RuleEnginecatches CEL execution errors and interpreter panics.- Failed CEL evaluations and
falseresults can emit a separateTRACEevent containing only statically referenced context properties. Metadata is logged by default;logFullCelContext: truelogs bounded values. - Context diagnostics bound depth, node count, collection size, string length, key length, and null-path traversal.
- Access-control context construction excludes authorization, proxy authorization, cookie, set-cookie, and API-key headers.
- Successful CEL evaluation returns only a boolean.
- The shared access-control runtime converts a missing rule or any rule-engine
error into
falsefor request authorization. - The final HTTP and MCP denial responses deliberately avoid exposing internal policy details.
The main gap is therefore not only context visibility. It is loss of structured decision information between CEL execution and the final access decision.
Goals
- Explain why an HTTP request, MCP tool call, or MCP tool-list entry was allowed, denied, hidden, or filtered.
- Distinguish a valid
falseresult from a malformed rule or runtime error. - Show the effective CEL context during local and development rule testing.
- Use the exact runtime evaluator for pre-deployment rule testing.
- Correlate diagnostics with the request, policy revision, endpoint, tool, and rule version that produced the decision.
- Preserve current fail-closed authorization behavior.
- Keep diagnostic overhead negligible when debugging is disabled.
- Share the implementation between HTTP access control and MCP routing.
- Keep raw credentials and other security-handler inputs out of the CEL context.
Non-Goals
- Return policy internals or request context to ordinary API or MCP callers.
- Log complete request, response, token, or tool-argument payloads by default.
- Rewrite CEL expressions into simpler expressions for diagnostic evaluation. Rewriting could change short-circuiting, macros, presence behavior, or types.
- Turn CEL into a mutation or general scripting language.
- Guarantee a natural-language proof for every
falseresult. The first implementation should report facts and outcomes rather than speculate. - Weaken
strictprofile field exposure to make debugging easier. - Support CEL rules that inspect raw authorization headers, cookies, API keys, tokens, or other authentication credentials.
Design Principles
Separate enforcement from explanation
Authorization still maps every non-matching or erroneous outcome to deny when the policy requires fail-closed behavior. A separate diagnostic result retains the reason for authorized consumers.
Prefer pre-deployment testing
The best production diagnostic is a rule that was tested before publication. Runtime diagnostics remain necessary because live tokens, headers, endpoint resolution, and tool arguments can differ from test fixtures.
Report facts, not invented explanations
For a false result, report the expression, statically referenced paths,
structural metadata or full values according to logFullCelContext, profile,
and rule-combination behavior. Do not claim that a particular clause caused the
result because the current CEL evaluator has no execution observer that can
prove it.
Outcome Model
The boolean returned by the current rule path should be replaced internally by a structured outcome. The exact Rust types can evolve, but the semantic model should be stable:
#![allow(unused)]
fn main() {
enum RuleConditionOutcome {
Matched,
NotMatched,
CompileError { message: String, source: Option<SourceLocation> },
EvaluationError { message: String },
NonBoolean { actual_type: String },
SecurityProfileRejected { message: String },
}
enum RuleExecutionOutcome {
Matched,
ConditionNotMatched,
ConditionError(RuleConditionOutcome),
ActionRejected { action_ref: String },
ActionError { action_ref: String, message: String },
ActionNotFound { action_ref: String },
RuleNotFound,
}
}
RuleEngine::execute_rule should return a structured result. Compatibility
wrappers can continue returning bool where callers do not need diagnostics.
The shared access-control runtime should then build an aggregate decision:
#![allow(unused)]
fn main() {
struct AccessEvaluation {
decision: AccessDecision,
reason: AccessDecisionReason,
rules: Vec<RuleEvaluation>,
skipped_rule_ids: Vec<String>,
}
}
AccessDecision remains the enforcement projection. AccessEvaluation is the
diagnostic projection.
Rule Aggregation Trace
The trace must preserve accessRuleLogic behavior:
- With
all, evaluation stops at the first rule that does not match or errors. Remaining rule IDs are recorded as skipped because of short-circuiting. - With
any, evaluation stops at the first matching rule. Remaining rule IDs are recorded as skipped because of short-circuiting. - A rule-engine error is recorded as an error outcome even when the enforcement
projection treats it like
false. - A missing rule body is recorded as
rule_not_found, notnot_matched. defaultDenydecisions are recorded without fabricating a rule evaluation.
Actions can mutate the rule context, so the existing candidate-context behavior
for any must remain unchanged. Diagnostic collection must observe the same
execution and must not evaluate a rule a second time.
Decision Trace
A decision trace should use a stable structured shape suitable for JSON logs and the future rule-test API:
{
"timestamp": "2026-07-24T15:42:11.184Z",
"correlationId": "request-123",
"serviceId": "com.networknt.gateway-1.0.0",
"policyRevision": "sha256...",
"surface": "mcp-tools-call",
"endpoint": "/config/query@post",
"toolName": "queryConfig",
"ruleType": "req-acc",
"ruleLogic": "all",
"decision": "denied",
"reason": "condition_not_matched",
"rules": [
{
"ruleId": "allow-config-read",
"expressionHash": "sha256...",
"requestedProfile": "strict",
"effectiveProfile": "strict",
"outcome": "condition_not_matched",
"contextMode": "full",
"referencedPaths": [
"permission.roles",
"auditInfo.subject_claims.ClaimsMap.roles"
],
"referencedValues": {
"permission.roles": ["config-admin"],
"auditInfo.subject_claims.ClaimsMap.roles": ["developer"]
},
"contextTruncated": false
}
],
"skippedRuleIds": [],
"referenceAnalysisIncomplete": false,
"traceTruncated": false
}
Trace logs can include the expression text because this feature is intended for
local and development use. They should also include ruleId, expression hash,
and policy revision so that a diagnostic can be tied to the exact loaded policy.
Diagnostic Context Modes
The runtime has two context projections:
| Mode | Contents | Intended Use |
|---|---|---|
| metadata | referenced paths plus presence, JSON type, null state, and collection or string size without property values | default trace behavior |
| full | actual values for statically referenced CEL properties | local and development environments only |
Full context must never mean an unbounded raw dump. Both modes use the same CEL-profile projection and diagnostic budgets as evaluator-error diagnostics.
CEL Reference Discovery
light-rule currently pins cel 0.14.0. That crate exposes two relevant APIs:
Program::references()returns the root variables and functions referenced by the compiled expression. ForauditInfo.subject_claims.ClaimsMap.roles, it reports the root variableauditInfo.Program::expression()exposes the public parsed AST.Expr::Selectnodes contain their operand and selected field, so Light-Fabric can walk the AST and recover the complete static member pathauditInfo.subject_claims.ClaimsMap.roles.
The crate does not expose an evaluation observer or a list of properties actually read at runtime. Reference discovery is therefore static: it includes properties in branches that short-circuit evaluation and cannot always resolve computed map keys.
The compiled-program cache should store a reference projection alongside each program:
#![allow(unused)]
fn main() {
struct CelProgramEntry {
program: Arc<CelProgram>,
referenced_roots: Vec<String>,
referenced_paths: Vec<String>,
reference_analysis_incomplete: bool,
}
}
Static dot selections and indexes with literal string keys should produce exact
paths. For dynamic indexing, the projection should fall back to the smallest
known root or static prefix and set referenceAnalysisIncomplete: true. Macro
and comprehension-local variables must not be mistaken for root context
variables.
This fallback deliberately broadens full mode. For example,
ClaimsMap[claimName] cannot identify the selected claim statically, so full
mode emits the bounded ClaimsMap parent object as well as claimName.
Metadata mode emits only the parent’s type and size. Rule authors should prefer
literal indexes or dot selections when they want the narrowest diagnostic
projection.
The reference walker is coupled to the public AST and operator names in the
pinned cel 0.14.0 crate. A CEL dependency upgrade must revalidate the walker
and its literal-index, dynamic-index, and comprehension tests.
Access-Control Context Boundary
Security authenticates the request before access control runs. The security
handler should expose normalized identity and authorization facts through
auditInfo.subject_claims.ClaimsMap; it should not forward the credential used
to establish those facts into CEL.
The access-control CEL context must therefore exclude raw values such as:
AuthorizationandProxy-Authorizationheaders- cookies and session tokens
- API keys and client secrets
- private keys or credential material owned by an earlier handler
If a rule needs an identity fact derived from one of these inputs, the security
handler should expose the normalized claim instead. For example, a rule should
read a roles claim from auditInfo, not parse the bearer token.
Other CEL inputs should be policy-oriented: endpoint and tool identity, permissions, selected non-sensitive headers, correlation metadata, referenced tool arguments, and the request or response properties required by the rule phase.
Referenced Context Projection
Diagnostics start from the variables exposed by the effective CEL security profile and keep only the statically referenced properties. A diagnostic cannot include an unrelated property merely because it exists in the root context.
The projection then keeps only the statically referenced properties. Most
current request-access rules reference JWT claims below
auditInfo.subject_claims.ClaimsMap, so a rule that reads only the caller’s
roles should not cause unrelated headers, claims, or tool arguments to be
logged.
The diagnostic path does not add header or JSON-path masking. Sensitive
credentials are excluded when the access-control context is constructed, and
reference projection removes unrelated properties. logFullCelContext: true
can therefore emit the actual values of referenced policy properties. It is
intentionally a local and development-only setting and must emit a startup
warning.
Values that require special handling
- JWT claims and
toolArgumentsinclude only statically referenced properties. responseBodyandresponseBodyJsoninclude only statically referenced properties.- A row-level CEL filter captures at most the current bounded row and should not repeat the shared context for every rejected row.
- Binary data is represented by type and length, not encoded into the trace.
Bounds
Reuse the existing limits for diagnostic context depth, nodes, collection items, string characters, key characters, null paths, and null-path traversal. Add an overall serialized trace byte limit. Every limit must have a corresponding truncation field so an operator can distinguish absent data from omitted data.
Runtime Configuration
The short-term configuration should be one root-level property in
access-control.yml:
logFullCelContext: false
This property does not enable trace logging. The logging filter still controls whether CEL trace events are emitted, for example:
RUST_LOG=light_rule::cel=trace,info
The property controls only the context projection used by those events:
falseor absent: emit referenced paths and structural metadata without property values atTRACEtrue: emit actual values for the referenced CEL properties atTRACEfor local or development use
No mode performs diagnostic masking. The serialized trace remains size-limited and reports truncation, but full-mode values are otherwise emitted as they appear in the credential-free access-control context.
The runtime should emit a prominent startup warning when
logFullCelContext: true is loaded:
Full CEL context logging is enabled. This setting is intended only for local or development environments.
Trace events should cover CEL evaluation errors and successful evaluations that
return false. A per-rule false event must be labeled as a rule outcome, not
as a final access denial, because accessRuleLogic: any can allow the request
through a later rule.
This changes earlier evaluator-error logging behavior: bounded context and
candidate-null-path diagnostics previously appeared with the WARN or ERROR
event. They now appear only in the separate light_rule::cel TRACE event;
the warning or error retains the expression and failure details without request
context. Operators who relied on warning-level context must enable the trace
target while diagnosing CEL failures.
Rule-Test API
Rule authors should be able to test before publishing. The test path must use
the same light-rule evaluator, profile enforcement, context conversion, and
action registry as the target runtime.
Suggested request:
{
"rule": {
"ruleId": "allow-config-read",
"ruleType": "req-acc",
"conditionLanguage": "cel",
"conditionSecurityProfile": "strict",
"expression": "permission.roles.exists(r, r in auditInfo.subject_claims.ClaimsMap.roles)"
},
"context": {
"auditInfo": {
"subject_claims": {
"ClaimsMap": {
"roles": ["developer"]
}
}
},
"permission": {
"roles": ["config-admin"]
},
"endpoint": "/config/query@post",
"toolName": "queryConfig",
"toolArguments": {}
}
}
Suggested response:
{
"outcome": "condition_not_matched",
"requestedProfile": "strict",
"effectiveProfile": "strict",
"referencedPaths": [
"permission.roles",
"auditInfo.subject_claims.ClaimsMap.roles"
],
"referencedValues": {
"permission.roles": ["config-admin"],
"auditInfo.subject_claims.ClaimsMap.roles": ["developer"]
},
"warnings": [],
"contextTruncated": false
}
The API should support two modes:
- isolated rule evaluation for the editor
- endpoint-policy evaluation using a selected configuration snapshot, including
rule ordering,
allorany, permissions, and default-deny behavior
The endpoint-policy mode is essential because a rule can work in isolation but still not be selected for the deployed endpoint.
The test API must not accept arbitrary production credentials or fetch live requests. Test contexts are explicit input, access-controlled, size-limited, and excluded from ordinary logs.
Portal Experience
The CEL editor should present:
- the documented context schema for the selected rule phase
- requested and effective security profiles
- syntax and profile validation before publication
- an editable constructed test context
- matched, not-matched, or error as distinct states
- per-rule endpoint-policy results and short-circuiting
- referenced paths beside their constructed test values
- missing-field, null-receiver, type, and non-boolean errors
A production denial shown to an ordinary caller remains generic. Rule testing uses a constructed context in the Portal rather than retrieving a live request context from the gateway.
MCP-Specific Behavior
The same trace model applies to MCP with a surface field that identifies:
mcp-tools-callmcp-tools-listmcp-response-filter
For tools/call, include the resolved configured tool name and endpoint, but
include only tool-argument properties referenced by the CEL expression.
For CEL-based tools/list, one request can evaluate many tools. Do not emit one
large context event per hidden tool by default. Emit a bounded aggregate summary
with counts and keep any per-tool trace detail bounded. Distinguish:
- hidden by a CEL
falseresult - hidden by CEL error
- hidden by unknown-rule fallback
- skipped after
maxCelEvaluations - served from the tools-list visibility cache
If a result came from cache, include the policy revision and cache outcome. Do not claim that CEL was evaluated for that request.
Response-Filter Behavior
Response filtering needs separate outcomes because false can mean different
things at different layers:
- rule-level CEL condition did not select the filter
- a response-filter action rejected execution
- a row-level CEL expression excluded a row
- a row-level expression failed and the row was excluded fail-closed
- the complete top-level object was denied
Row filtering can evaluate the same expression hundreds or thousands of times. Capture the first bounded failure sample, total matched/excluded/error counts, and the number of suppressed diagnostics. Never emit the entire response body.
Logs, Metrics, and Audit
Use stable tracing targets and event names rather than prose-only messages:
target: light_rule::decision
event: cel_rule_decision
Summary events should contain scalar fields that remain useful in text and JSON logging. Detailed context can be a bounded JSON field.
Recommended counters:
- CEL evaluations by outcome, rule type, and security profile
- access decisions by reason and surface
- diagnostic traces captured, truncated, sampled out, or rate-limited
- rule-test requests by outcome
- tools-list evaluations skipped by
maxCelEvaluations
Do not use rule IDs, endpoints, tool names, correlation IDs, or user identities as unbounded metric labels. Those belong in logs or traces.
Individual rule evaluations normally remain operational tracing events, not durable security audit records, unless deployment policy requires otherwise.
Performance and Abuse Controls
- When
TRACEis disabled for the CEL tracing target, avoid cloning or serializing context solely for diagnostics. - Build a context projection only after confirming that the trace event is enabled.
- Reuse the compiled CEL program and any compile-time reference analysis.
- Never reevaluate a CEL expression to explain its first result.
- Bound trace size, row samples, and tools-list samples.
- Rate-limit rule-test execution.
- Apply existing CEL profile and expression-complexity limits to test requests.
- Record when sampling or limits omitted diagnostic data.
Failure Behavior
Diagnostics must never change the access decision. If reference analysis, serialization, or log emission fails:
- preserve the original allow or deny result
- emit a bounded diagnostic-system error without request context
- increment a diagnostic failure counter
- do not retry on the request path
The diagnostic system itself should be panic-contained where it processes untrusted context values.
Implementation Phases
Phase 1: Preserve outcomes
- Introduce structured condition and rule outcomes in
light-rule. - Stop collapsing every
RuleEngineerror into an unexplainedfalseinside the shared access-control runtime. - Preserve boolean compatibility wrappers for unrelated callers.
- Add aggregate outcomes for
all,any, missing rules, and default deny. - Emit safe summary tracing events for errors and denials.
Phase 2: Referenced diagnostic projection
Implemented for the short-term runtime diagnostic path:
- Exclude raw credentials when constructing the access-control CEL context.
- Centralize context reference analysis, projection, and bounds.
- Apply the same safe projection to existing CEL error and panic logs.
- Use
Program::references()and the public AST to extract referenced roots and static member paths. - Add the root-level
logFullCelContextconfiguration toaccess-control.yml.
Remaining enhancements:
- Add expression hashes, policy revisions, and stable reason codes.
- Add MCP tools-list aggregation and response-row sampling.
Phase 3: Authoring workflow
- Add isolated-rule and endpoint-policy test APIs.
- Integrate the APIs into the Portal CEL editor.
- Validate rules with the target runtime evaluator before publication.
Testing Strategy
Outcome tests
- CEL
trueandfalse - compile error, evaluation error, panic, and non-boolean result
- missing rule and missing action
- action rejection and action error
allandanyshort-circuit traces- default allow and default deny
- HTTP, MCP
tools/call, MCPtools/list, and response-filter surfaces
Security tests
- raw authorization headers, cookies, API keys, and credentials never enter the access-control CEL context
- strict diagnostics never expose roots unavailable to strict CEL
- unrelated headers, claims, tool arguments, and response fields are absent
- metadata mode reports structure without property values
- full mode reports values only for referenced properties
- request input cannot enable full context logging
- truncation flags are correct
Parity tests
- rule-test and runtime execution return the same outcome for the same rule, context, profile, and policy revision
- diagnostic collection does not change action mutation or short-circuiting
- cached MCP tools-list decisions are labeled as cached and are not reported as fresh evaluations
Performance tests
- disabled trace logging has no material request-path allocation regression
- metadata and full trace logging remain within defined latency budgets
- large contexts, large rows, and tools-list fan-out remain bounded
Resolved Decisions
- The pinned
cel 0.14.0crate exposes referenced root variables and a public AST, so Light-Fabric will statically extract related context paths and log only those properties. It does not expose actual runtime property reads. - The diagnostic path will not add masking. Access-control context construction excludes raw credentials, and full mode logs the actual bounded values of referenced policy properties.
- The design does not add log retention or audit requirements. Local logging and its existing rotation policy own trace retention.
Recommended Default
The referenced-context trace logging and root-level logFullCelContext switch
are the short-term implementation. Next, add structured outcomes, then the
rule-test API and Portal editor support.
This sequence fixes the information-loss problem, gives authors a safe way to inspect rule context during local development and test rules before deployment, and can adopt runtime property-read tracing later if the CEL evaluator adds a trustworthy observer API.
Access Control Handler Design
The access-control handler enforces fine-grained authorization for normal HTTP API endpoints. It should reuse the same Light-Rule policy model that the MCP router uses for tool authorization:
access-control.ymlcontrols whether policy is enabled, whether missing endpoint rules deny by default, how multiple request access rules combine, and which endpoint prefixes are skipped.rule.ymlcontains CEL rule bodies and endpoint mappings.req-accrules run before the upstream API endpoint is called.res-filrules run after the upstream API endpoint responds and before the response is returned to the caller.
The access-control handler and the MCP router should share the same
frameworks/light-pingora/src/access_control.rs runtime. The difference is the
boundary where that runtime is applied. The MCP router protects MCP tools. The
access-control handler protects API endpoints in the normal handler chain.
Goals
- Enforce fine-grained access control for REST or HTTP API endpoints.
- Reuse the existing
req-accandres-filrule phases. - Reuse the built-in action classes:
RoleBasedAccessControlAction,ResponseRowFilterAction, andResponseColumnFilterAction. - Keep rule definitions portable between gateway products when the endpoint key and context fields are equivalent.
- Support exact endpoint rules, Java-style path templates, and parent path entries.
- Keep the business API unaware of caller-specific row and column filtering.
Non-Goals
- Do not create a second rule engine for HTTP APIs.
- Do not support the legacy native condition-row format in Light-Fabric.
Light-Fabric rules must use
conditionLanguage: cel. - Do not replace base authentication. The access-control handler assumes an earlier security handler has already built the caller principal.
- Do not push row or column filtering into business handlers.
HTTP routing and policy identity
Handler selection and HTTP ACL resolution are independent. A handler route such
as /github/repos/* selects a shared chain; it does not become an ACL grant.
The Gateway retains its routing endpoint for router rewrites, metrics, LLM
routing policy, and evidence. The access-control handler uses the original URI
path without the query and the original HTTP method, captured before header,
path, or upstream method rewriting. Verified caller claims remain the rule input.
Ordinary HTTP policies resolve in this order:
- An exact path for that method wins, including over templates.
- A unique matching full-path
{parameter}template wins over literal prefixes. Multiple matching templates deny, even if one has longer parameter names or more literals. Resolve overlap with an exact policy or disjoint templates. - Existing literal parent-path policies retain longest-prefix semantics, with
a
/boundary and the same method. They explicitly authorize descendants; default-deny cannot make an endpoint unknown beneath such a configured grant. - With no match, retain the concrete
path@methodidentity and apply configureddefaultDeny. No wildcard handler fallback supplies permission.
Methods compare case-insensitively; policies differing only in method case are
ambiguous and deny. Ambiguity never falls back to a broader prefix or to
defaultDeny: false. An exact policy resolves otherwise overlapping templates.
Paths remain case-sensitive. Gateway preserves the received URI: routing uses
legacy raw matching and forwarding retains punctuation, escapes, repeated and
trailing slashes, apart from existing configured routing rewrites. There is no ingress rewrite or new path-signing requirement.
Ordinary HTTP ACL lookup uses a separate policy identity: decode one level,
validate UTF-8, retain URI segment characters, and use uppercase escapes for
space/UTF-8 data. Encoded unreserved aliases select their stricter exact policy.
A single trailing slash is equivalent for policy lookup only; the wire slash
remains, preserving directory redirects and signed paths.
On a chain containing access-control, ambiguous separators, dot segments,
malformed escapes, ASCII controls, invalid UTF-8, decoded percent signs and
repeated slashes reject. Literal or encoded semicolons also reject on these
chains: some upstream routers strip segment path parameters, which could let
/api/private;x=1 miss an exact policy and select a broader prefix. This
restriction includes semicolons intended as literal segment data.
If raw routing would select an unprotected chain but
an alias would select an ACL chain, a guard rejects before dispatch. Its
conservative alias probe is never used for forwarding or handler selection;
it honors route order and recognizes separators, double encoding, segment
parameters and dot collapse. Parameter stripping is only a denial probe; raw
unprotected paths such as /x/;jsessionid=... remain unchanged unless their
alias crosses an ACL boundary. Protected route patterns and base paths are validated and compiled
once at load/reload; invalid candidates fail with diagnostics and retain the
active runtime. Unprotected patterns are not canonicalized or silently dropped.
An explicitly spelled unprotected alias route (for example /open%2Fprivate)
fails load/reload if its alias overlaps a protected route or default chain
without earlier complete unprotected coverage. Validation respects methods,
route order, base paths, segment templates and terminal wildcards; partial
unprotected coverage is conservatively insufficient. Use an unambiguous route
spelling or explicitly protect the route, retaining endpoint-specific ACL
policies. Do not broaden grants to silence this diagnostic. Rejected reloads
retain the active handler set and pinned requests. Dynamic aliases in ordinary
template/wildcard captures remain guarded at request time.
Normal route matching performs no per-pattern canonicalization.
Spaces and UTF-8 remain supported in ACL-protected segment data. Queries are
unchanged. Direct /branches/feature%2Ftopic remains restricted on an ACL chain,
because upstream separator/data interpretation is not uniform. An applicable
query ref is unchanged; it does not make the direct branch endpoint supported.
Unprotected %2F, %25 and repeated slashes retain legacy behavior unless an
alias would cross into an ACL chain. HMAC sees the original request path.
The legacy /@method policy matches the root itself; it does not implicitly
grant all descendants. A non-root trailing slash selects its canonical policy.
skipPathPrefixes is checked against the canonical concrete identity and the
selected policy, and disabled/skipped ACLs retain their shared request/response
gates.
The selected endpoint key and immutable ACL runtime generation are retained for request authorization and response filtering. Ordinary body-bearing proxy/router requests are bounded and authorized before upstream dispatch, then replayed through Pingora’s existing prebuffered-body hook. Earlier tokenization is applied once before authorization, as in the former body-filter path. Portal command/query envelope decoding, MCP tool identities, A2A policy identities, and the MCP/LLM-owned body authorization paths keep their protocol-specific behavior.
Handler Placement
The access-control handler should run after authentication and before routing to the upstream API service:
request
-> TLS / CORS / rate-limit / header handlers
-> security or unified-security handler
-> access-control req-acc
-> proxy or route handler
-> access-control res-fil
-> response
If no authenticated principal is available, a req-acc rule can still evaluate
headers and endpoint metadata, but role, group, user, and claim-based rules will
normally fail closed.
Shared Runtime
The existing runtime already models the common policy engine:
#![allow(unused)]
fn main() {
AccessControlRuntime
-> authorize_tool(...)
-> filter_mcp_response(...)
}
For API endpoints, these functions should be generalized rather than duplicated. The MCP-specific names can remain as compatibility wrappers, but the shared runtime should expose endpoint-neutral operations:
#![allow(unused)]
fn main() {
authorize_request(
endpoint,
headers,
auth,
request_context,
correlation_id
)
filter_response(
endpoint,
headers,
auth,
request_context,
response_status,
response_body,
correlation_id
)
}
The MCP router can keep passing toolName and toolArguments. The API handler
should pass API-oriented values such as path parameters, query parameters,
request method, and request body metadata.
Endpoint Keys
Endpoint rule keys should use the same stable format as the MCP router:
{path}@{method}
Examples:
/offers@get
/v1/accounts/{accountId}@get
/v1/accounts@post
The query string must not be part of the endpoint key. Query parameters belong in the rule context so CEL can inspect them without multiplying endpoint rule entries.
Endpoint matching order should remain:
- Exact endpoint key.
- Java-style path template match, such as
/v1/accounts/{id}@get. - Parent path entry, such as
/v1/accounts@getfor/v1/accounts/123@get.
Configuration
access-control.yml is the handler-level switch:
enabled: true
accessRuleLogic: any
defaultDeny: true
defaultInclude: false
skipPathPrefixes:
- /health
- /adm
claimMappings: {}
Fields:
enabled: when false, the handler allows requests and does not filter responses.accessRuleLogic:anyallows a request if anyreq-accrule passes;allrequires every listedreq-accrule to pass. This setting applies only toreq-acc; it does not apply tores-fil.defaultDeny: when true, a request with no matching endpoint rule or noreq-accrule is denied.defaultInclude: controls response row-filter behavior when a row filter is configured but no caller claim matches any configured row-filter entry. Whenfalse, the row filter returns no rows. Whentrue, the row filter preserves the legacy include-all behavior.skipPathPrefixes: endpoint prefixes that bypass access-control entirely.claimMappings: maps permission dimensions to JWT claim names for built-in request-access, row-filter, column-filter, and MCP tool-visibility behavior. Standard keys areroles,groups,positions,attributes, andusers. Custom row or column dimensions use the dimension name as the mapping key.
For example, a deployment with custom role and tenant claims can use:
claimMappings:
roles:
- custom_roles
tenant:
- tenant_id
Standard aliases remain active when a dimension has no configured mapping.
Existing toolsListAccessControl.claimMappings configuration remains supported
as a compatibility fallback, but top-level claimMappings takes precedence and
applies consistently to authorization and response filtering.
rule.yml contains the reusable rules and endpoint policy:
ruleBodies:
allowOfferRead:
common: Y
ruleId: allowOfferRead
ruleName: Allow offer read
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
filterOfferRows:
common: Y
ruleId: filterOfferRows
ruleName: Filter offer rows
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200
&& responseBody != ""
&& auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.ResponseRowFilterAction
filterOfferColumns:
common: Y
ruleId: filterOfferColumns
ruleName: Filter offer columns
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200
&& responseBody != ""
&& auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.ResponseColumnFilterAction
endpointRules:
/offers@get:
req-acc:
- allowOfferRead
res-fil:
- filterOfferRows
- filterOfferColumns
permission:
roles: offer-viewer offer-admin
row:
role:
offer-viewer:
- colName: priority
operator: "<"
colValue: 50
- colName: active
operator: "="
colValue: true
col:
role:
offer-viewer: offerId,title,segment,state,category,priority
In this example, offer-viewer can call GET /offers but only receives active
offers with priority below 50, and the active field is removed from the final
payload. With defaultInclude: false, offer-admin can call the same endpoint
only if a matching row-filter entry exists or the endpoint omits row filtering.
If an endpoint has a row block but no row entry matches the caller’s role,
group, position, attribute, or user claim, the filtered result is empty. Set
defaultInclude: true only when a deployment intentionally wants the legacy
include-all behavior for unmatched row-filter claims.
Rule Context
The access-control handler should build the same core context shape as the MCP router so existing CEL rules and actions stay reusable:
| Field | Description |
|---|---|
auditInfo | Normalized authenticated principal claims and correlation id. |
headers | Lower-cased request headers. |
endpoint | Stable endpoint key, such as /offers@get. |
permission | The endpoint permission object. |
correlationId | Correlation id when present. |
statusCode | Response status code during res-fil. |
responseBody | Response body string during res-fil. |
For API endpoints, add API-specific fields:
| Field | Description |
|---|---|
requestMethod | HTTP method. |
requestPath | Path without query string. |
queryParameters | Parsed query parameter map. |
pathParameters | Values captured from a path template when available. |
requestBody | Parsed JSON request body when available and within size limits. |
requestBodyText | Raw request body string when parsing is not enabled. |
The existing MCP fields can remain optional:
| Field | Usage |
|---|---|
toolName | Present for MCP router calls, absent or empty for API endpoints. |
toolArguments | Present for MCP router calls. API endpoints should prefer queryParameters, pathParameters, and requestBody. |
Permission values should continue to be injected twice:
- as the namespaced
permissionobject - as top-level convenience fields, such as
roles,row, andcol
This preserves compatibility with existing rule bodies and built-in action
classes. Runtime-owned context fields are the exception: permission keys named
auditInfo, headers, endpoint, toolName, toolArguments,
correlationId, permission, responseBody, responseBodyJson, statusCode,
or accessControl remain available under permission but are not promoted to
the top level. This prevents endpoint configuration from replacing verified
identity, request, or response context.
Request Access
req-acc runs before the upstream API call.
The handler should:
- Build the endpoint key from request path and method.
- Skip the request if the endpoint matches
skipPathPrefixes. - Find endpoint rules by exact, template, or parent match.
- Deny when
defaultDeny: trueand no matchingreq-accrule exists. - Build the rule context from auth, headers, endpoint, request fields, and endpoint permissions.
- Execute the listed
req-accrules withaccessRuleLogic. - Return
403when access is denied.
When accessRuleLogic: any, each candidate rule should receive a cloned
context, and the first passing rule should win. When accessRuleLogic: all,
rules should run sequentially against the same context and all must pass.
Response Filtering
res-fil runs after the upstream API response returns.
The handler should:
- Only filter response payloads that are safe and useful to parse, starting
with JSON arrays, JSON objects containing an
itemsarray, and single JSON objects for column filtering. - Buffer the full response body before filtering.
- Decode or avoid upstream compression before JSON parsing.
- Add
statusCode,responseBody, and the parsed mutable JSON value to the same rule context shape. - Execute
res-filrules sequentially in the order listed on the endpoint. - Serialize the filtered JSON once after all
res-filactions complete. - Replace the response body with the final filtered JSON.
- Recompute response headers that depend on body size, such as
content-length.
Ordering matters. Row filters must run before column filters when the row
predicate depends on a field that should be hidden in the final response. For
example, a row filter can use active == true, and the later column filter can
remove active from the returned rows.
res-fil is always a sequential all pipeline. accessRuleLogic: any applies
only to req-acc; response filters never use any semantics.
Response filtering requires a full payload. It is not compatible with streaming
or indefinite responses unless the gateway buffers the entire response first.
For Transfer-Encoding: chunked, the gateway must buffer and then emit a normal
filtered response. Server-Sent Events and other long-lived streaming responses
should bypass res-fil or be rejected when an endpoint requires response
filtering.
Compressed upstream responses need explicit handling. The gateway should either
strip or normalize Accept-Encoding on the upstream request so the backend
returns plaintext JSON, or it must decompress before filtering and recompress
afterward. Filtering compressed gzip, br, or deflate bytes as JSON must
fail closed.
If a res-fil rule is missing, fails, or returns false, the handler should fail
closed for protected API endpoints. For early rollout, a deployment can choose a
fail-open compatibility mode only if it is explicit in configuration and emits a
high-severity log or module-registry status.
CEL must not directly rewrite the HTTP response body. A res-fil rule-level CEL
expression decides whether the filter action should run. The response-filter
pipeline owns JSON parsing, final serialization, response body replacement, and
header updates such as content-length. Actions own row or column mutation of
the parsed response value.
The default model should remain declarative:
ResponseRowFilterActionapplies permission-defined row filters.ResponseColumnFilterActionapplies permission-defined field keep or remove lists.
Row Filter Default Behavior
Row filtering must fail closed by default. If ResponseRowFilterAction runs for
an endpoint with a configured permission.row block, but the caller has no
matching entry under any supported dimension (role, group, position,
attribute, or user), the action must return an empty row set when
defaultInclude: false.
This prevents a common policy gap:
permission:
row:
role:
teller:
- colName: accountType
operator: "="
colValue: C
In the legacy include-all behavior, a caller without the teller role would
match no row-filter entry and receive every row. With defaultInclude: false,
the same caller receives no rows. A caller with the teller role receives only
rows where accountType == "C".
defaultInclude applies only to row-filter miss behavior:
false: unmatched row-filter dimensions retain no rows. This is the secure default and should be used for new deployments.true: unmatched row-filter dimensions retain all rows. This is a compatibility mode for deployments that relied on the old behavior.
If a row-filter entry matches the caller, normal row predicate evaluation still
applies. If multiple dimensions match, the configured filter groups are combined
with the existing sequential all behavior so a row must satisfy every matched
group.
If an API needs a richer row predicate, add a CEL-aware action rather than making rule-level CEL mutate JSON:
ruleBodies:
filterOfferRowsWithCel:
ruleId: filterOfferRowsWithCel
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200 && responseBody != ""
actions:
- actionClassName: com.networknt.rule.ResponseCelRowFilterAction
actionValues:
rowExpression: >
auditInfo.subject_claims.ClaimsMap.role == "offer-admin"
|| (row.priority < 50 && row.active == true)
ResponseCelRowFilterAction should compile rowExpression during rule load and
evaluate it once per row with a curated context containing row, auditInfo,
headers, endpoint, permission, and API request metadata. It owns row
retention and failure handling for the parsed response value. It should not
deep-clone the full base context for every row; use a child context that shadows
row, or reuse one mutable context and update only the row binding.
If a row-level CEL evaluation fails for one row, for example because priority
is missing, the action should drop that row and continue. Compile errors,
invalid actionValues, or other configuration errors should fail the whole
action closed.
MCP Router Comparison
The MCP router and access-control handler should share configuration and runtime semantics:
| Concern | MCP router | Access-control handler |
|---|---|---|
| Protected target | MCP tools | HTTP API endpoints |
| Endpoint key | Tool endpoint or derived {path}@{method} | Request {path}@{method} |
| Request input | toolArguments | query parameters, path parameters, request body |
| Request access phase | req-acc before backend tool call | req-acc before upstream API call |
| Response filter phase | res-fil before JSON-RPC result | res-fil before HTTP response |
| Response body target | MCP structuredContent or text content | HTTP response body |
| Rule language | CEL only | CEL only |
This split keeps MCP behavior specialized for JSON-RPC tool calls while letting API endpoint authorization use the same policy and action implementation.
Reload And Observability
The loader should continue to support both standalone files and values.yml
projection:
access-control.ymloraccess-control.yamlrule.ymlorrule.yamlaccess-control.*values invalues.ymlrule.ruleBodiesandrule.endpointRulesvalues invalues.yml
The root-level access-control setting below controls CEL context trace detail:
logFullCelContext: false
At the default false, failed or rejected CEL expressions log referenced paths
and structural metadata at TRACE. When set to true, the same event includes
bounded values for only the referenced properties. Full mode is intended only
for local or development environments. Access-control context construction
excludes credential headers in both modes.
The module registry should report:
- whether access control is enabled
- whether rule config is loaded
- number of rule bodies
- number of endpoint mappings
- last reload status
- validation errors for rejected CEL, missing rule ids, or invalid endpoint keys
Config reload should build a new immutable runtime and swap it atomically after validation succeeds. If validation fails, the handler should keep the last known-good runtime.
Implementation Notes
The current AccessControlRuntime is already close to the shared runtime. The
main implementation work is to remove MCP-specific naming from the reusable API
and add an HTTP response-body adapter:
- Keep
authorize_toolandfilter_mcp_responseas wrappers for the MCP router. - Add endpoint-neutral authorization and response-filter methods.
- Add API-specific request context fields without changing the existing
auditInfo,headers,endpoint,permission,responseBody, andstatusCodefields. - Reuse
find_service_entry,rule_ids_for,permission_for, and the default action registry. - Reuse
ResponseRowFilterActionfor JSON arrays, object payloads with anitemsarray, and single top-level JSON objects. A denied single object is replaced with an empty object for HTTP responses; the MCP adapter returns a tool error withisError: true. - Reuse
ResponseColumnFilterActionfor JSON arrays, object payloads with anitemsarray, and single top-level JSON objects. - Add handler-level tests for exact endpoint, path template endpoint, parent path endpoint, default deny, skip prefixes, row filtering, and column filtering.
Ordinary HTTP capture and chain compatibility
Wildcard handler routes (including /github/repos/* for every method) are
independent of ACL keys and remain supported. Wildcard ordinary HTTP ACL keys
containing * are unsupported: startup/reload rejects them with a migration
diagnostic under both default-deny modes. Migrate ACL rules to exact paths,
full-path templates, or deliberate literal-prefix policies; do not change
wildcard handler routing or broaden grants. A rejected reload retains the valid
active ACL runtime. Already admitted requests retain their selected key and
runtime for both authorization and response filtering across a valid reload.
access-control.yml has two explicit capture settings, independent of HMAC,
LLM, and tokenizer configuration:
bodyReadTimeoutMillis: 10000
maxBufferedBodyBytes: 268435456
Both must be positive. Ordinary HTTP ACL retains its 10 MiB per-body limit.
Declared lengths reject before allocation/interim responses; streamed bytes
remain bounded. The deadline covers the entire body read, including writing
100 Continue. An aggregate permit accounts for retained payload bytes and is
released with request ownership on completion, denial, failure, timeout, or
cancellation. HMAC-captured originals keep their existing HMAC permit; shared
Bytes avoid a second original-body allocation. ACL-transform output reserves
its configured tokenizer maximum (capped by the ACL limit) before tokenization,
then reduces its charge to actual retained output. Budget exhaustion returns
503, size excess 413, and read timeout 408. These payload quotas are not a
process-RSS bound: protocol chunks, allocator slack, JSON parsing/serialization,
and token-vault operations have separate resource costs. The capture deadline
does not bound token-vault database operations or rule evaluation.
Supported body chains authenticate first, run tokenize before access-control,
and dispatch only after authorization. HMAC (standalone or through
unified-security/unified) authenticates original bytes before ACL capture;
ACL observes tokenized bytes and upstream receives the cached transformation
once. Downstream framing is preserved until capture; transformed upstream
framing is updated independently. Detokenization requests uncompressed upstream
responses even when ACL has already authorized the body. Empty-body DELETE
remains valid. HMAC checks signature/selector/delivery-header syntax before body
capture but validates the MAC only over original captured bytes.
Startup and handler/security reload reject tokenizer-after-body-capturing-ACL,
recognized-security-gate-after-ACL, or tokenizer-before-HMAC/unified orders with precise
order diagnostics. jwt is a security alias. Existing HMAC validation also
rejects duplicate entry points or simultaneous standalone HMAC and unified
security on the same effective chain. The recognized gates are hmac,
security, jwt, unified-security and unified; this is not a validator of
all authentication handlers (such as stateless, token, Basic or API-key).
GET/HEAD ACL evaluates query/header data and does not require tokenizer-before-ACL;
other gates and original-byte HMAC order still apply. Wildcard-method routes
remain unsupported by handler configuration loading; the body-order helper
treats a wildcard conservatively when called directly. HMAC terminal chains
remain router-only. The shared MCP/A2A matcher is unchanged.
Deferred tokenization uses chunked HTTP/1.1 upstream framing or HTTP/2 DATA frames, selected from the actual negotiated protocol, not its configuration preference. Known empty bodies retain length zero; preauthorized transformed bodies retain their exact length. Deferred transformation over HTTP/1.0 is unsupported and fails explicitly.
The 10-second capture deadline is total elapsed read time, unlike an idle timeout that resets after every chunk. Slow uploads that previously progressed indefinitely can now receive 408. The ACL body ceiling remains 10 MiB; the tokenizer default is 1 MiB. Owners should explicitly select a total deadline for supported upload sizes and minimum acceptable throughput before rollout, without changing unrelated proxy idle settings. Reserving a full tokenizer output maximum under the default 256 MiB budget reduces concurrent upload capacity: each original plus its maximum output is charged until transformation completes. This conservative reservation also accounts for originals shared from HMAC through its independent budget; the combined budgets are not one process-wide memory ceiling.
Local regression commands (no live requests):
rtk cargo test -p light-gateway http_acl -- --test-threads=2
rtk cargo test -p light-pingora --lib http_policy -- --test-threads=2
rtk cargo test -p light-pingora --lib capture_headers -- --test-threads=2
rtk cargo test -p light-gateway -- --test-threads=2
rtk cargo test -p light-pingora --lib -- --test-threads=2
The cache-only composition fixture is enabled by a development-only
test-support dependency. It uses seeded token-cache entries and a disconnected
pool, exercises production Gateway/tokenizer methods, and does not qualify a
PostgreSQL vault or any deployed runtime. Recorder finish requires stopped
producers, drains a nonblocking OS listener through WouldBlock, and joins every
accepted worker. A pending asynchronous accept is never treated as queue proof.
Design Document: Centralized Agentic Skill Registry
Subject: Transitioning from File-Based Markdown Skills to a Database-Backed Skill Registry
1. Executive Summary
Currently, most AI agent frameworks rely on localized Markdown (.md) files to define agent “skills.” While Markdown is highly LLM-native and human-readable, it creates significant bottlenecks at an enterprise scale regarding strict typing, API integration, and context window limits.
This document proposes transitioning to an Agentic Control Plane (Centralized Skill Registry) backed by a database. By decoupling skill metadata, schemas, and instructions, and by utilizing dynamic routing, we will achieve hierarchical structuring, strict schema enforcement, and progressive disclosure of tools to agents.
The registry serves enterprise business agents, native workflow agents, coding agents, and personal assistants. It stores governed content and immutable package references; profile-specific runtime hosts materialize the selected skill. See Light-Agent Execution for those runtime and isolation boundaries.
2. Problem Statement
Managing agent skills as flat Markdown files introduces several scaling challenges:
- Lack of Strict Typing: Markdown cannot enforce data types (e.g., ensuring a parameter is an integer vs. string), leading to hallucinated or malformed tool inputs.
- Context Window Exhaustion: Loading dozens or hundreds of skill definitions at startup overwhelms the LLM context window, increasing latency, token costs, and tool-misuse.
- Static Deployments: Updating a skill or changing access permissions requires a full application redeploy.
- Poor Discoverability: Flat file structures offer no native mechanism for progressive disclosure or tool search.
3. Data Models & Formats
To solve the limitations of purely text-based skills, we will adopt a hybrid, structured format stored within a database (e.g., PostgreSQL/MongoDB). The architecture uses the right format for the right job:
- JSON Schema: Used strictly for defining parameters, inputs, and tool shapes. Natively supported by OpenAI/Anthropic/Google tool-calling APIs.
- LightAPI Description (YAML/JSON): Used to map endpoint-level API capabilities to skills across REST, JSON-RPC, gRPC, and MCP.
- OpenAPI / OpenRPC / Protobuf: Referenced by LightAPI where protocol-native specifications already exist.
- Immutable Execution Artifact / Endpoint Reference: API skills reference a governed endpoint. Scripted skills reference a signed, content-addressed package and entrypoint; mutable source code is not executed directly from a database row.
- Markdown: Retained only for the
instructionsorpromptfields, as LLMs excel at parsing markdown headers and lists for constraints and persona instructions.
LightAPI is the preferred source format for API-backed skills because it describes endpoint identity, protocol invocation, input schema, request mapping, result shape, examples, and behavior notes in one agent-oriented document. See LightAPI Description Design for the endpoint description model.
YAML and JSON are the external skill document formats. In the portal database,
they should not replace the Markdown instruction field. The normalized model is
structured columns and relationships for identity, versioning, taxonomy, tools,
and execution metadata, plus content_markdown for the LLM-facing instruction
body. If the portal later needs to persist a full structured skill document,
add a nullable JSONB skill-spec column beside content_markdown and normalize
YAML imports to JSON.
3.1 Proposed Database Schema Structure
Light Portal stores skills in structured catalog tables. Below is a representation of the skill payload:
{
"skill_id": "sk_finance_001",
"name": "generate_financial_report",
"version": "1.2.0",
"tags": ["finance", "reporting"],
"tool_schema": {
"type": "function",
"function": {
"name": "generate_financial_report",
"description": "Generates a Q3 report based on ticker symbol.",
"parameters": {
"type": "object",
"properties": {
"ticker": {"type": "string", "description": "The stock ticker"}
},
"required": ["ticker"]
},
"response_schema": {
"type": "object",
"properties": {
"report_url": {"type": "string"},
"status": {"type": "string"}
}
}
}
},
"execution": {
"type": "rest_api",
"endpoint_id": "ep_finance_report_001",
"endpoint": "https://internal-api.company.com/v1/finance/report",
"method": "POST"
},
"instructions": "## Role\nYou are a financial analyst.\n## Constraints\n- Never hallucinate financial data.\n- Always return exact numbers."
}
3.2 Skill Authority And Executable Packages
A skill is discovery and guidance content. Assignment of a skill never grants tools, credentials, network, filesystem, workflow, or model-provider access. The effective capability is always the intersection of caller authority, agent-definition policy, skill policy, live gateway/controller policy, and the selected execution profile.
Each tool link resolves to a stable internal tool reference with a server-owned
execution placement (gateway, runner, workflow, or fixed-service),
model-facing alias, and schema digest. Materialization cannot change that
placement. Gateway tools intersect live gateway tools/list; runner tools
intersect execution policy, lease allowedTools, the approved runtime-tool
manifest, and live local availability. The independently authorized sets may be
combined only after alias collisions are rejected or resolved by deterministic
server-owned aliases. One placement never grants another placement’s tool by
name coincidence.
Keep skill_t.content_markdown as the instruction source. Add a separate
skill_package_t only for skills that require scripts, binaries, templates, or
other runtime assets. A package record should contain:
- host, skill, semantic version, package ID, and immutable artifact URI;
- SHA-256 digest, media type, size, and entrypoint;
- supported runtime profiles and required capability names;
- minimum sandbox boundary, network, workspace, and credential requirements;
- provenance/attestation reference, signer, scanner result, and review state;
- created, deprecated, revoked, and retention state.
The package bytes belong in immutable artifact storage, not a TEXT column.
The portal may store authoring source separately, but only a reviewed, signed,
scanned, active package can be materialized for execution. The runtime host
verifies the package, lease, policy digest, and entrypoint before mounting it
read-only.
Existing tool_t.script_content is a legacy authoring/runtime shortcut. It
must not become the production execution path for untrusted Python or
JavaScript. Publication should compile or package that source into an immutable
artifact and require runner placement.
4. Hierarchical Structure & Progressive Disclosure
Dumping 500 JSON schemas into an LLM’s context window will cause system failure. The Centralized Controller will act as a mediator, enforcing hierarchy and progressive disclosure (giving the agent only the schemas it needs, exactly when it needs them).
4.1 Implementing Hierarchy & Tagging
Because JSON Schema does not have built-in folders, hierarchy and categorization are enforced via the platform’s global entity management system:
- Namespacing: Tool names follow a strict convention:
[domain]_[subdomain]_[action](e.g.,aws_rds_provision). - Tags & Categories: Instead of hardcoded columns, the registry utilizes the
entity_tag_tandentity_category_ttables (withentity_type = 'skill'). This allows for unlimited flat tagging and deep hierarchical folder structures that are consistent across the entire portal. - Discovery API: Portal-query filters by these tags/categories to scoped skill sets for specific agent personas. Agents cache the effective catalog locally and reload it when runtime cache-management invalidation is triggered.
4.2 Progressive Disclosure Patterns
Agents should not load every executable tool into the LLM context. Instead, they should load their assigned skill/tool catalog from the portal API, cache it locally, and use one of the following progressive disclosure patterns:
Phase 5 starts with the Rust light-agent. The agent loads
genai-query/getEffectiveAgentCatalog, keeps a local cache keyed by
hostId + agentDefId + serviceId + envTag, ranks cached skill/tool entries with
keyword and routing-field matching, and intersects the selected tool names with
the live gateway tools/list result before giving schemas to the model.
Execution remains gateway tools/call.
Pattern A: Meta-Tools (Dynamic Injection)
The agent is booted with only two “meta-tools” designed for discovery.
- Local catalog search: Agent searches its cached assigned skills. The cache contains lightweight summaries and mapped tool names.
- Schema loading: Once the agent identifies the correct tool, it loads the schema from the local catalog cache or refreshes the cache from portal-query.
Pattern B: Semantic Tool RAG (Zero-Shot Discovery)
For highly complex systems with thousands of skills:
- Tool descriptions are embedded into a Vector Database (e.g.,
pgvector). - When the user prompts the system (e.g., “Reset my AWS password”), portal-query or the agent’s local cache performs semantic search and retrieves the Top-3 most relevant JSON Schemas.
- The agent boots with only those 3 tools in its context.
Pattern C: Multi-Agent Orchestration (Supervisor / Worker)
Hierarchy is mapped to agent teams.
- A Supervisor Agent holds routing tools (e.g.,
delegate_to_finance,delegate_to_devops). - When
delegate_to_devopsis triggered, the supervisor routes to a DevOps Worker Agent, loading only the specific DevOps JSON schemas into its context.
4.3 Runtime Profiles And Materialization
One centrally assigned skill can serve several agent products without forcing every runtime to consume the same physical format.
| Runtime profile | Materialized skill input | Execution boundary |
|---|---|---|
| Enterprise business agent | Bounded Markdown instructions and selected API/MCP schemas | Long-lived light-agent plus light-gateway |
| Native workflow agent | Instructions, structured task input, and output schema | light-workflow, with no local tools |
| Coding agent | Read-only SKILL.md, references, and verified package assets | light-agent-worker inside a runner sandbox |
| Personal assistant | Instructions, connector mappings, schedule/notification policy, and optional reviewed package | light-agent plus gateway or personal edge runner |
| External agent adapter | Adapter-specific files generated from the immutable skill version | Same sandbox as the selected runtime adapter |
Materializers are deterministic and versioned. Their output digest becomes part of the turn/runtime policy snapshot. Runtime-specific rendering may adapt file names or metadata, but it cannot add a tool or capability absent from the effective catalog.
For sandboxed profiles, the materializer emits immutable skill-package
references, digests, sizes, and mount/entrypoint policy; it does not fetch from
inside the sandbox. Trusted light-workflow-runner code downloads each package
before sandbox creation, verifies digest/signature/provenance/scan bindings and
archive safety, and stages it as a read-only mount with nodev, nosuid, and
noexec unless a reviewed entrypoint requires execution. The worker may
revalidate the mounted manifest, but neither the worker nor generated code
receives artifact-store credentials or package-download egress. Verification
or staging failure prevents the sandbox from starting.
Use the following content precedence, from strongest to weakest:
- server and execution policy;
- signed platform/tenant skill versions assigned to the agent;
- reviewed user-specific skill configuration;
- repository or workspace-local instructions;
- prompts, retrieved data, messages, and tool output.
Repository-local and user-generated skills are useful context but untrusted. An agent-generated skill is stored as an inactive proposal. It becomes usable only after schema validation, security scanning, human or policy review, immutable packaging, and explicit assignment. Self-modification never hot-activates new authority in the current turn.
5. Example Flow: Dynamic Loading in Action
User: “I need to provision a new database for the marketing team.”
- Turn 1: Discovery
- Agent Context: Has a local cache of assigned skill summaries.
- Agent Action: Searches the local cache for
provision database.
- Turn 2: High-Level Awareness
- Local Cache Result: Returns token-efficient summaries from the portal catalog:
[{"name": "aws_rds_provision", "description": "Creates AWS RDS DB"}, {"name": "mongo_atlas_create", "description": "Creates Mongo cluster"}] - Agent Action: Decides AWS is needed and loads the cached schema for
aws_rds_provision.
- Local Cache Result: Returns token-efficient summaries from the portal catalog:
- Turn 3: Strict Execution
- Agent Catalog: Provides the full JSON schema (requiring
instance_type,storage_gb). - Agent Action: Understands parameters and safely executes
aws_rds_provisionthrough the gatewaytools/callpath.
- Agent Catalog: Provides the full JSON schema (requiring
6. Operational Benefits & Security
By centralizing skills in a database, the platform gains enterprise-grade operational capabilities:
- Dynamic Updates: API endpoints, instructions, and schemas can be updated in the database without restarting agents.
- Permission-Aware Discovery (RBAC): By linking skills to LightAPI endpoint descriptions and
api_endpoint_t, portal-query can limit catalog disclosure to the current agent or tenant, while runtime gateway policy still authorizes execution. - A/B Testing: Portal catalog metadata can route 50% of an agent’s requests to
skill_v1and 50% toskill_v2to measure prompt/tool efficacy. - Audit Logging: Catalog disclosure and gateway execution can be logged separately, preserving a compliance trail without moving tool execution into the registry.
- Distilled Memory RAG: Following the “Hindsight” pattern, raw conversation history (
agent_session_history_t) is separated from RAG-optimized memory (session_memory_t). This prevents the “noisy context” problem while maintaining a perfect audit trail. - Profile Reuse Without Privilege Reuse: The same logical skill can be rendered for enterprise, coding, workflow, or personal-assistant runtimes, while each runtime receives only its independently authorized tools and execution capabilities.
- Supply-Chain Controls: Executable packages are content-addressed, signed, scanned, reviewable, revocable, staged by the trusted runner before sandbox creation, mounted read-only, and always run through an approved sandbox profile.
7. LightAPI As Skill Source
API-backed skills should be generated from endpoint-level LightAPI descriptions whenever possible.
The skill registry should store skill metadata, access control, grouping, and agent-facing instructions. The LightAPI description should remain the source of truth for endpoint invocation and verification details.
Recommended flow:
- Light-Portal creates or imports endpoint-level LightAPI descriptions.
- API owners enrich endpoint descriptions with examples, behavior notes, result cases, and visibility.
- Approved endpoint descriptions are published as agent skills.
- The agent loads assigned skill summaries from portal-query and caches them locally.
- When the agent selects a skill, it loads the relevant LightAPI disclosure level from the local cache or refreshes from portal-query.
- Execution goes through the gateway
tools/callpath, preserving runtime policy and downstream authorization.
This avoids manually duplicating every API endpoint as a separate hand-written skill while still giving agents strict schemas and progressive disclosure.
8. Workflow-Backed Skills
Some skills need more than instructions and a curated tool set. A skill that
must orchestrate several tools, wait for human approval, retry failed steps,
run assertions, or preserve a durable audit trail should be backed by
light-workflow.
The boundary should stay clear:
| Layer | Responsibility |
|---|---|
| Skill | Discovery metadata, taxonomy, instructions, allowed tools, and agent guidance. |
| Workflow | Ordered execution, branching, retries, assertions, human tasks, durable state, and audit events. |
| Gateway | Runtime tool execution through tools/list and tools/call. |
Workflow backing should be optional. Simple skills can stay as instructions plus
tool mappings. Durable or regulated processes should link to workflow
definitions and let light-workflow own execution.
Recommended storage:
- Keep
wf_definition_t.definitionas the canonical workflow YAML. - Keep
skill_t.content_markdownas the LLM-facing skill instruction body. - Add
skill_workflow_tto link skills to workflow definitions with a role such asprimary,validation,remediation, ortest. - Treat
skill_tool_tas the allowed tool set for a workflow-backed skill. Validation should flag workflow tool-call steps that are not linked to the skill.
The Portal Skill Workspace should embed a generic Workflow Editor instead of creating a skill-specific workflow runtime. The editor provides YAML editing, step preview, reference lookup, validation, and test runs. Skill authoring provides the surrounding context: skill metadata, taxonomy, allowed tools, effective prompt preview, and workflow link configuration.
9. Next Steps
- Complete phase 3 by adding category and tag assignment to existing skill create/update forms, backed by
entity_category_tandentity_tag_twithentity_type = 'skill'. - Save skill taxonomy through a composite skill command so the skill row and selected taxonomy associations are emitted from the same user action.
- Move the richer authoring workspace, effective prompt preview,
skill_tool_t.configformalization, workflow-backed skills, and “create skill from LightAPI/tool” flows to phase 3.5. - Build the generic Workflow Editor for YAML editing, parsed step preview, catalog references, validation, and workflow test runs.
- Complete phase 4 agent assignment by improving the
agent_skill_tUI, adding an Agent Definition assignment context, and adding a batch assignment composite command that emits oneAgentSkillCreatedEventper selected skill. - Enforce phase 4 assignment validation in command handlers and UI preflight: assigned skills must be active and must have at least one active direct
skill_tool_tlink. Workflow-backed skills still rely onskill_tool_tas the allowed tool set. - Keep live gateway
tools/listruntime executability checks as a diagnostics or governance concern, not as phase 4 persistence validation. - Complete phase 5 for the Rust agent with the
genai-querygetEffectiveAgentCatalogendpoint, claim checks againsthost,sid, andenv, local catalog caching, keyword/routing search, gatewaytools/listintersection, and controller-driven cache invalidation. - Complete phase 6 governance for the Rust agent only: normalize sensitivity
tiers to
public,internal,confidential, andrestricted; filter blocked tools before catalog disclosure; compare the effective catalog with gatewaytools/listthrough/diagnostics/tools; and keep execution through gatewaytools/call. - Enforce destructive, approval-required, and sensitivity metadata at the
gateway with debug/auditInfo fields when a call is blocked. Do not use
workflow
audit_log_tfor catalog disclosure; use auditInfo/file logging until a generic governance audit table is introduced. - Keep current active row plus aggregate version as the approval/version boundary until workflow-owned approval state is implemented.
- Add publishing from LightAPI endpoint descriptions into the skill registry.
- Migrate existing file-based skills into structured catalog payloads, keeping instructions in Markdown and converting parameters to JSON Schema.
- Implement Pattern B (Semantic Tool RAG) after indexed catalog fields and embeddings are ready for production search.
- Add runtime-profile compatibility, stable tool references, server-owned
gateway/runner/workflow/fixed-service placement, schema/alias digests, and
deterministic materializer metadata without overloading
content_markdownor relying on a tool name as authority. - Add
skill_package_t, immutable artifact publication, signature and scan verification, revocation, trusted runner-side download/safe extraction, read-only staging, and runner-only package execution. Workers receive no artifact-store credential or package-download authority. - Add reviewed proposal lifecycle for repository-local and agent-generated skills; never activate generated content automatically.
- Add materializer conformance fixtures proving that enterprise, workflow, coding, personal-assistant, and external-adapter outputs preserve the same skill version and cannot widen its effective capability set.
Skill Workflow Orchestration
Status
Proposed demo design.
Executive Summary
This design describes a focused demo for agent-driven orchestration in Light-Fabric. The demo uses one agent with two skills:
- A skill that starts a workflow which calls two REST APIs directly.
- A skill that starts a workflow which calls the same two REST APIs through the MCP router.
Both paths solve the same business use case and return the same output. The visible difference is the execution trace:
- The REST workflow shows
light-workflowinvoking HTTP endpoints directly. - The MCP workflow shows
light-workflowinvoking MCPtools/call, withlight-gatewayrouting each tool call to the same backend REST APIs.
This demonstrates that skills provide agent-facing guidance and discovery, workflows provide durable orchestration, and the gateway provides the MCP data plane for tool execution.
Goals
- Show one agent selecting between two assigned skills.
- Show a workflow that orchestrates multiple REST APIs directly.
- Show a second workflow that orchestrates the same APIs through MCP tools.
- Keep the input and output contract identical across both workflows.
- Keep the demo small enough to explain in a few minutes.
- Preserve the runtime boundary: skills guide, workflows orchestrate, gateway executes MCP tool calls.
Non-Goals
- Do not benchmark REST versus MCP latency.
- Do not claim that MCP replaces REST. The demo shows two supported access patterns over the same backend capabilities.
- Do not require every skill to be workflow-backed. Simple skills can remain instructions plus allowed tools.
- Do not move MCP tool execution into the portal registry or agent catalog.
Runtime tool execution stays on the gateway
tools/callpath. - Do not make the demo depend on a large endpoint catalog.
Recommendation
Use two APIs, not one.
A one-API demo can show sequencing, but it does not clearly prove cross-service orchestration. Two APIs show a more realistic enterprise shape: the workflow has to collect data from one business capability and make a decision through another capability.
Use four endpoints for the base demo.
| Demo size | Endpoint count | Recommendation | Why |
|---|---|---|---|
| Smoke test | 2 | Optional only | Shows a happy path, but not enough variation. |
| Base demo | 4 | Recommended | Covers path parameters, query parameters, arrays, request bodies, branching, and transformation. |
| Advanced demo | 6 | Later phase | Adds parallel enrichment, compensation, or audit callbacks. |
The base demo should be small enough to run repeatedly while still proving meaningful orchestration behavior.
General Agent And Workflow Boundary
The demo illustrates one direction of a bidirectional integration. The same boundary applies to enterprise, coding, and personal-assistant profiles:
| Handoff | Owner after handoff | Intended use |
|---|---|---|
| Agent starts workflow | light-workflow | Durable branching, retries, assertions, human tasks, long waits, regulated business processing |
| Workflow performs native agent call | light-workflow | Bounded model reasoning with structured input/output and no interactive session or local tools |
| Workflow submits agent-service job | light-agent | Interactive or tool-using work, coding/research jobs, memory-aware work, or runner-agent placement |
The existing call.agent behavior remains the backward-compatible
native-workflow mode. It runs inside light-workflow and validates
schema-bound JSON. A future explicit agent-service mode submits a typed job
to light-agent. The selected agent definition and policy—not the workflow
prompt—decide whether light-agent handles the job in its service or through a
runner sandbox.
The handoff includes authenticated caller and tenant context, correlation ID, input and output schemas, deadline, idempotency key, cost/action budget, cancellation behavior, and bounded delegation depth. light-workflow never spawns Codex, Pi, Claude Code, Hermes, OpenClaw, or another external agent binary directly. Those products can only run behind a registered light-agent runtime adapter in an approved execution profile.
Do not convert every agent turn into a workflow. Conversation history, streaming, interruptions, tool correction, personal channel delivery, and workspace-aware model loops remain agent-domain concerns. Conversely, an agent that starts a workflow stores the workflow reference and observes its public result; it does not reproduce the workflow state machine in its own context.
Reject cyclic or unbounded agent/workflow delegation. A child handoff inherits or narrows the initiating deadline, budget, data boundary, and authorization.
Demo Scenario
The demo domain is personalized offer recommendation.
The agent receives a prompt such as:
Recommend an offer for customer CUST-1001.
The agent can use either skill:
Personalized Offer via REST WorkflowPersonalized Offer via MCP Router
If the prompt does not specify REST or MCP, the demo agent should not pick a path at random. It should ask a short clarification question:
Do you want to run this through the direct REST workflow or through the MCP
router workflow?
Scripted demos can avoid the clarification by naming the path in the prompt.
Both skills start a workflow that:
- Loads the customer profile.
- Loads customer preferences and consent.
- Stops if the customer has not consented.
- Searches for eligible offers.
- Selects the best offer.
- Records the offer decision.
- Returns a normalized decision payload.
APIs And Endpoints
Customer Profile API
The Customer Profile API owns customer data and preferences.
| Endpoint | Shape | Purpose |
|---|---|---|
GET /customers/{customerId} | Path parameter, object response | Load customer identity, segment, region, and account status. |
GET /customers/{customerId}/preferences?channel=portal | Path parameter plus query parameter | Load consent, preferred categories, and contact channel rules. |
Offer Decision API
The Offer Decision API owns eligible offer lookup and decision recording.
| Endpoint | Shape | Purpose |
|---|---|---|
GET /offers?segment={segment}&state={state}&category={category} | Query parameters, array response | Search active offers matching the customer profile and preferences. |
POST /offer-decisions | JSON request body, object response | Persist the selected offer decision and return a decision id. |
Demo API Runtime Services
The two business APIs should be implemented as real Rust services using the
light-axum framework, not as ad hoc mocks. This keeps the demo aligned with
normal Light-Fabric service lifecycle behavior:
- load runtime configuration from config-server
- bind HTTP using configured server settings
- register with controller through
portal-registry - appear in the control panel service-discovery view
- support gateway service discovery by
serviceIdandenvTag
Recommended demo apps:
| App | Service id | Default HTTP port | Purpose |
|---|---|---|---|
demo-customer-profile-api | com.networknt.demo.customer-profile-1.0.0 | 8085 | Serves customer profile and preference data. |
demo-offer-decision-api | com.networknt.demo.offer-decision-1.0.0 | 8086 | Serves offer lookup and decision recording. |
The ports are config defaults only. They must be configurable through config-server values so local, Docker, Kubernetes, and shared demo environments can choose different ports without recompiling.
Both services should expose:
GET /health
The API endpoints should return deterministic demo data. A database is not required for the first demo; in-memory seed data is enough as long as the data is stable and documented. If later demos need persistence, keep it behind the same endpoint contract.
Light-Axum Bootstrap
Each demo API should follow the normal light-axum pattern: implement
AxumApp, return an axum::Router, and let LightRuntimeBuilder own binding,
configuration, shutdown, and controller registration.
The service should read config from the same runtime config files used by other Light-Fabric services:
startup.yml
server.yml
portal-registry.yml
Example config-server values for the Customer Profile API:
startup.host: dev.lightapi.net
startup.externalConfigDir: /var/lib/demo-customer-profile-api/config-cache
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.networknt.demo.customer-profile-1.0.0
server.environment: demo
server.ip: 0.0.0.0
server.advertisedAddress: demo-customer-profile-api
server.httpPort: 8085
server.enableHttp: true
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
Example config-server values for the Offer Decision API:
startup.host: dev.lightapi.net
startup.externalConfigDir: /var/lib/demo-offer-decision-api/config-cache
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.networknt.demo.offer-decision-1.0.0
server.environment: demo
server.ip: 0.0.0.0
server.advertisedAddress: demo-offer-decision-api
server.httpPort: 8086
server.enableHttp: true
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
server.advertisedAddress must be a reachable address, not 0.0.0.0. In
Kubernetes, use the Service DNS name. In local Docker Compose, use the Compose
service name. In a native VM demo, use the VM hostname or another reachable
address.
Controller Registration
The services should register with controller using the runtime’s
portal-registry integration. The controller registration payload must include
at least:
serviceIdenvTag- protocol
- advertised address
- port
- discovery token or portal registry token, according to environment policy
After startup, the control panel should show two registered service instances:
com.networknt.demo.customer-profile-1.0.0 / demo
com.networknt.demo.offer-decision-1.0.0 / demo
The MCP router configuration should prefer these service IDs over fixed
targetHost values where service discovery is available. Fixed targetHost
values are still useful for a minimal local smoke test.
Optional Advanced Endpoints
The base demo should start with four endpoints. If we later want to demonstrate more workflow shapes, add one or two optional endpoints:
| Endpoint | Shape Demonstrated | Use |
|---|---|---|
GET /customers/{customerId}/risk | Parallel enrichment | Run profile, preferences, and risk lookup before offer selection. |
POST /offer-decisions/{decisionId}/audit | Follow-up side effect | Record a compliance audit event after the decision is created. |
POST /offer-decisions/{decisionId}/cancel | Compensation | Cancel the decision if a later step fails. |
Agent, Skills, And Workflows
Use one agent so the demo highlights skill selection rather than agent handoff.
| Object | Name | Responsibility |
|---|---|---|
| Agent | Demo Orchestration Agent | Receives the user request and selects one of the assigned skills. |
| Skill | Personalized Offer via REST Workflow | Guides the agent to start the direct REST workflow. |
| Skill | Personalized Offer via MCP Router | Guides the agent to start the MCP-backed workflow. |
| Workflow | personalized-offer-rest-v1 | Orchestrates direct HTTP calls to the two REST APIs. |
| Workflow | personalized-offer-mcp-v1 | Orchestrates MCP tool calls through the gateway router. |
The skill registry should link each skill to its workflow definition through
skill_workflow_t. The workflow definition remains canonical in
wf_definition_t.definition. The skill content_markdown remains
agent-facing guidance, not the executable workflow source.
Execution Paths
Direct REST Workflow
User prompt
-> Demo Orchestration Agent
-> Personalized Offer via REST Workflow skill
-> light-workflow
-> Customer Profile API
-> Offer Decision API
-> normalized decision result
This path is useful for showing direct, durable API orchestration.
MCP Router Workflow
User prompt
-> Demo Orchestration Agent
-> Personalized Offer via MCP Router skill
-> light-workflow
-> MCP tools/call
-> light-gateway MCP router
-> Customer Profile API
-> Offer Decision API
-> normalized decision result
This path is useful for showing MCP protocol orchestration over the same backend API capabilities.
Common Workflow Contract
Both workflows should accept the same input:
{
"customerId": "CUST-1001",
"channel": "portal"
}
Both workflows should return the same successful output shape:
{
"status": "APPROVED",
"customerId": "CUST-1001",
"selectedOfferId": "OFFER-TRAVEL-01",
"decisionId": "DEC-1001"
}
Both workflows should return comparable business outcomes for known edge cases:
{
"status": "NO_CONSENT",
"customerId": "CUST-3003",
"reason": "Customer has not consented to personalized offers."
}
{
"status": "NO_ELIGIBLE_OFFER",
"customerId": "CUST-2002",
"reason": "No active offer matches the customer profile and preferences."
}
Workflow Shape
The REST and MCP workflows should have the same logical steps.
| Step | REST workflow action | MCP workflow action |
|---|---|---|
| Load profile | GET /customers/{customerId} | tools/call customer_get_profile |
| Load preferences | GET /customers/{customerId}/preferences | tools/call customer_get_preferences |
| Check consent | Workflow condition | Workflow condition |
| Search offers | GET /offers | tools/call offer_search |
| Select offer | Workflow expression or rule | Workflow expression or rule |
| Record decision | POST /offer-decisions | tools/call offer_record_decision |
| Return result | Workflow output mapping | Workflow output mapping |
The workflow should own branching, retries, and output normalization. The agent should not manually sequence each API call after the workflow starts.
Error Handling And Retries
Business outcomes and technical failures should be treated differently.
Business outcomes are expected workflow results and should not be retried:
NO_CONSENTNO_ELIGIBLE_OFFER
Technical failures should use bounded workflow retries:
| Failure | Recommended behavior |
|---|---|
| Customer Profile API timeout | Retry the profile step with exponential backoff. |
Offer Decision API returns 503 | Retry the affected offer step with exponential backoff. |
Gateway MCP tools/call timeout | Retry the MCP tool-call step with the same workflow policy. |
| Persistent downstream failure | End with a controlled technical failure result and preserve the workflow trace. |
Recommended transient retry status codes:
408, 429, 502, 503, 504
The POST /offer-decisions step should include an idempotency key derived from
the workflow instance id and selected offer id. This prevents duplicate
decisions when a retry happens after the backend processed the first request
but the response was lost.
For parity, the REST and MCP workflows should use the same retry policy. In the MCP path, the gateway should preserve enough error detail for the workflow trace to show the tool name, mapped backend endpoint, status code, and correlation id.
MCP Tool Mapping
The MCP workflow should use a small, explicit tool set.
| MCP tool | Backend endpoint | Arguments |
|---|---|---|
customer_get_profile | GET /customers/{customerId} | customerId |
customer_get_preferences | GET /customers/{customerId}/preferences | customerId, channel |
offer_search | GET /offers | segment, state, category |
offer_record_decision | POST /offer-decisions | customerId, offerId, channel, source, reason |
The MCP tool input schemas should be normalized JSON objects. The gateway router maps those objects to path parameters, query parameters, or request bodies for the backend REST APIs.
The MCP skill should list these tools in skill_tool_t as its allowed runtime
tool set. Workflow validation should flag an MCP tool-call step if it references
a tool that is not linked to the skill.
Gateway Tool Configuration Example
Current gateway HTTP tool execution maps GET arguments to query parameters and
sends non-GET arguments as JSON request bodies. To support endpoint shapes such
as GET /customers/{customerId} without changing the backend API, the demo
should add or configure explicit path-template substitution before the request
is sent.
Recommended minimal mapping shape:
mcp-router.tools:
- name: customer_get_profile
description: Get a customer profile by id.
protocol: http
serviceId: com.networknt.demo.customer-profile-1.0.0
envTag: demo
path: /customers/{customerId}
method: GET
apiType: http
inputSchema:
type: object
required:
- customerId
properties:
customerId:
type: string
toolMetadata:
pathParams:
- customerId
With this mapping, the MCP tool call:
{
"name": "customer_get_profile",
"arguments": {
"customerId": "CUST-1001"
}
}
should be routed to:
GET /customers/CUST-1001
The path parameter should not also be appended as a query parameter. Arguments
not listed under pathParams can still be appended as query parameters for GET
requests or sent as JSON body fields for POST requests.
Skill Content Markdown Guidance
The skill content_markdown should explain when and how the agent should use
the skill. It should not duplicate the workflow definition or the full API
contract.
Example REST skill content:
## Purpose
Use this skill when the user asks for a personalized offer decision through the
direct REST workflow.
## Inputs
- customerId: customer identifier, such as CUST-1001
- channel: request channel, default portal
## Behavior
- Start workflow personalized-offer-rest-v1.
- Return the workflow result as the answer.
- Do not manually call offer APIs outside the workflow.
- If the user does not specify REST or MCP, ask which execution path they want.
Example MCP skill content:
## Purpose
Use this skill when the user asks to demonstrate MCP router orchestration for a
personalized offer decision.
## Inputs
- customerId: customer identifier, such as CUST-1001
- channel: request channel, default portal
## Behavior
- Start workflow personalized-offer-mcp-v1.
- The workflow will call MCP tools through the gateway.
- Return the workflow result as the answer.
- If the user does not specify REST or MCP, ask which execution path they want.
Structured execution metadata belongs in registry rows and workflow definitions, not only in markdown. The markdown is the LLM-facing explanation.
Output Normalization
The workflows should not pass raw endpoint responses directly to the agent. They should normalize backend responses into a stable business result.
Example raw POST /offer-decisions response:
{
"decisionId": "DEC-1001",
"customerId": "CUST-1001",
"offerId": "OFFER-TRAVEL-01",
"decision": "approved",
"createdAt": "2026-05-25T14:12:00Z",
"auditRef": "AUD-7788"
}
Normalized workflow output:
{
"status": "APPROVED",
"customerId": "CUST-1001",
"selectedOfferId": "OFFER-TRAVEL-01",
"decisionId": "DEC-1001"
}
The workflow should own this transformation so the REST and MCP variants produce identical final results even if their intermediate transport envelopes are different.
Demo Data
Use deterministic seed data so the demo is repeatable.
| Customer | Profile | Preferences | Expected result |
|---|---|---|---|
CUST-1001 | Premium segment, active, Ontario | Consent true, travel preferred | APPROVED with OFFER-TRAVEL-01. |
CUST-2002 | Standard segment, active, Ontario | Consent true, travel preferred | NO_ELIGIBLE_OFFER. |
CUST-3003 | Premium segment, active, Ontario | Consent false | NO_CONSENT. |
Seed offers:
| Offer | Match condition | Result |
|---|---|---|
OFFER-TRAVEL-01 | segment=premium, state=ON, category=travel | Eligible for CUST-1001. |
OFFER-CASHBACK-01 | segment=premium, state=BC, category=shopping | Not eligible for Ontario travel scenario. |
Demo Script
Run the REST workflow path first:
Use the REST workflow skill to recommend an offer for CUST-1001.
Expected observation:
- The agent selects
Personalized Offer via REST Workflow. - The workflow trace shows direct HTTP calls to the Customer Profile API and Offer Decision API.
- The final response contains
status=APPROVEDand a decision id.
Run the MCP workflow path second:
Use the MCP router skill to recommend an offer for CUST-1001.
Expected observation:
- The agent selects
Personalized Offer via MCP Router. - The workflow trace shows MCP
tools/callinvocations. - The gateway trace shows those tool calls routed to the same backend REST endpoints.
- The final response uses the same output shape as the REST workflow.
Then run one edge case:
Use either skill to recommend an offer for CUST-3003.
Expected observation:
- The workflow stops after the consent check.
- No offer decision is recorded.
- The result is
NO_CONSENT.
Run one ambiguity case:
Recommend an offer for CUST-1001.
Expected observation:
- The agent asks whether to use the direct REST workflow or the MCP router workflow.
- After the user chooses, the agent starts the selected workflow.
Run one technical failure case:
Use the MCP router skill to recommend an offer for CUST-1001 while the Offer
Decision API returns one transient 503.
Expected observation:
- The workflow retries the failed tool-call step.
- The gateway trace records the failed
offer_record_decisioncall and the successful retry. - The final response still uses the normalized
APPROVEDoutput shape.
Portal Authoring Flow
The portal should make the demo visible from the existing GenAI and workflow surfaces:
- Create or import the two REST APIs and four endpoint descriptions.
- Implement the two APIs as
light-axumservices. - Add config-server values for both API services.
- Start both services and verify controller registration.
- Publish MCP router tools for the same four endpoints.
- Create
personalized-offer-rest-v1in the workflow catalog. - Create
personalized-offer-mcp-v1in the workflow catalog. - Create the two skills in the skill registry.
- Link each skill to its primary workflow through
skill_workflow_t. - Link the MCP skill to its allowed tool set through
skill_tool_t. - Assign both skills to
Demo Orchestration Agent. - Use Skill Workspace preview and test panels to validate the effective prompt, workflow link, allowed tools, and sample test input.
Validation Rules
The authoring experience should validate the following before the demo is considered complete:
- Each skill has exactly one primary workflow link.
- The REST workflow does not require MCP tools.
- The MCP workflow references only MCP tools linked through
skill_tool_t. - Both workflows declare the same input schema.
- Both workflows declare the same normalized output shape.
- The four backend endpoint descriptions are active.
- Both demo API services load config from config-server.
- Both demo API services register with controller and appear in the control panel service-discovery view.
- The MCP router
tools/listresult includes the four expected tool names. - MCP router tools resolve the demo APIs by
serviceIdandenvTagin the service-discovery environment. - MCP path-parameter mappings are validated before the workflow test run.
POST /offer-decisionsincludes an idempotency key for retry safety.- Test runs for
CUST-1001,CUST-2002, andCUST-3003produce the expected outcomes.
Observability
The demo should show three different traces:
- Agent trace: which skill the agent selected and what workflow it started.
- Workflow trace: step order, branches, retries, and final output.
- Gateway trace: MCP tool name, mapped backend endpoint, status, duration, and correlation id for the MCP path.
Use the same correlation id across the agent request, workflow instance, and gateway calls where possible. This makes the REST and MCP execution paths easy to compare.
Security And Authorization
Authorization should be enforced at each layer:
- The agent can discover only assigned skills.
- The workflow can start only definitions visible to the authenticated caller or service identity.
- The MCP skill can expose only tools linked to the skill and allowed for the agent.
- The gateway still performs runtime MCP access checks before executing
tools/call. - Backend REST APIs continue to enforce their own authorization policies.
The skill registry is not a runtime bypass. It narrows discovery and guidance, while the workflow and gateway remain responsible for execution-time controls.
Context And Auth Propagation
The demo should explicitly show that caller context is preserved.
For direct REST workflow steps:
- The workflow start request records the initiating user, host, tenant, correlation id, and authorization context.
- The workflow executor exchanges the initiating authorization for a short-lived audience- and operation-scoped delegation token, or uses a workload identity whose on-behalf-of claims preserve the initiating subject, workflow instance, task, policy digest, and data boundary. It does not forward the caller’s unrestricted bearer token.
- Backend APIs enforce their normal authorization policies.
For MCP workflow steps:
light-workflowcalls the gateway MCP endpoint with the same correlation, tenant, locale, and a short-lived workflow-task-scoped delegation token.light-gatewayvalidates the MCP request and runtime tool authorization.- The MCP router forwards only approved identity/delegation context to the
backend REST API while regenerating transport-specific headers such as
Host,Content-Length, and connection management headers. - Backend APIs see the same business identity context they would see on the direct REST path.
The trace should show this propagation without exposing sensitive token values.
Acceptance Criteria
- One demo agent has both skills assigned.
- The REST skill starts
personalized-offer-rest-v1. - The MCP skill starts
personalized-offer-mcp-v1. - Both workflows accept the same input JSON.
- Both workflows return the same normalized output shape.
- The REST workflow trace shows direct REST calls to two APIs.
- The MCP workflow trace shows MCP
tools/callrouted through the gateway to the same two APIs. - The two APIs run as
light-axumservices with config-server supplied HTTP ports. - The two APIs register with controller and are visible in the control panel service-discovery view.
- The demo succeeds for
CUST-1001. - The demo returns controlled business outcomes for
CUST-2002andCUST-3003. - Ambiguous user prompts trigger a clarification question instead of random skill selection.
- A transient
503from the Offer Decision API is retried and appears in the workflow trace. - The MCP path preserves caller context through workflow, gateway, and backend REST calls.
- Neither workflow path can directly launch an external agent binary; a future service-mode agent task must enter light-agent through the typed job contract.
- Agent/workflow delegation preserves correlation and cannot exceed the initiating deadline, budget, authorization, or maximum delegation depth.
Related Designs
- Agentic Workflow
- Workflow Client Architecture
- Centralized Skills
- Light-Agent Execution
- LightAPI Description
- MCP Router
Hindsight Memory
Hindsight Memory is the core memory system for light-rs, designed to move beyond simple chat logs. Instead of just remembering what was said, the agent learns and forms mental models over time.
This design is strongly inspired by the paper Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects and extends it with multi-tenant support.
1. Core Concepts
Hindsight memory organizes information into three distinct “Pathway” types:
- World Facts: Objective truths about the environment (e.g., “The production server is in US-East-1”).
- Experiences: The agent’s own history of actions and results (e.g., “I tried to deploy to US-East-1 and it failed due to a timeout”).
- Mental Models: Synthesized understandings formed by reflecting on facts and experiences (e.g., “Deployments to US-East-1 are unstable during peak hours”).
2. The Three Operations
Interaction with the memory system is standardized into three primary operations:
Retain (Storage)
The retain operation ingests information. Behind the scenes, the system:
- Extracts entities and relationships.
- Normalizes time and temporal data.
- Stores the data in
agent_memory_unit_t.
Recall (Retrieval)
The recall operation retrieves relevant context using a hybrid strategy:
- Semantic: Vector similarity using the
hnswindex. - Graph: Following links in
agent_memory_link_t(causes, enables, prevents). - Temporal: Time-series filtering.
Reflect (Synthesis)
The reflect operation performs “deep thinking.” It analyzes existing memories to generate new insights, which are stored in agent_memory_reflection_t.
3. Database Architecture
The current Hindsight tables are integrated into the Portal-era multi-tenant schema. In the target architecture, immutable memory policy and hard directives are published through Config Server, while concrete banks, memory content, session history, provenance, reflection, and erasure state are owned by a tenant operational Memory API and store. See Control Plane And Operational Data.
| Table Name | Description |
|---|---|
agent_memory_bank_t | The primary container. Defines personality and disposition (skepticism, empathy). |
agent_memory_doc_t | Source documents (logs, files, transcripts) that provide the raw text for memory units. |
agent_memory_unit_t | Sentence-level “atoms” of thought. Stores content, embeddings, and fact types (world, experience, etc.). |
agent_memory_entity_t | Resolved Knowledge Graph nodes, optionally linked to platform users (user_t). |
agent_memory_unit_entity_t | The join table linking individual memories to the entities they mention. |
agent_memory_entity_cooccur_t | Association graph tracking concept relationships and co-occurrence counts. |
agent_memory_link_t | Defines causal and semantic relationships between memories (causes, enables, etc.). |
agent_memory_directive_t | “Hard rules” that override probabilistic learning. Control-plane content in the target architecture: directives become versioned, reviewed, digest-bound authoring data compiled into the Agent projection and targeting a bank profile or scope, not operational rows bound to a concrete bank. This table is a semantic migration, not a move into the operational memory schema. |
agent_memory_reflection_t | Synthesized high-level insights generated during the “Reflect” phase. |
agent_session_history_t | The materialized conversation context for active sessions, linked to a specific bank. Effectful action attempts and append-only session events remain authoritative when history projection is delayed or conflicted. |
4. Privacy & Multi-Tenancy
Isolation is managed at the Bank level using three scoping tiers:
- Global Host Bank (
user_idIS NULL,agent_def_idIS NULL):- Knowledge shared across all users and all agents within a specific
host_id. - Ideal for organization-wide SOPs, common facts, and shared documentation.
- Knowledge shared across all users and all agents within a specific
- Shared Agent Bank (
user_idIS NULL,agent_def_idIS NOT NULL):- Knowledge shared by all users interacting with a specific agent type.
- Used for maintaining a consistent agent “Persona” or specialized domain expertise.
- Private User Bank (
user_idIS NOT NULL):- Knowledge unique to a specific user.
- Can be scoped further by
agent_def_idto provide user-specific memory within a particular agent persona. - Used for personal preferences, private history, and individualized learning.
Memory-bank identity is not proof of access. Every recall, retain, reflection, export, resume, and deletion operation must derive host, principal, agent, and retention authority from authenticated server-side state. Personal-assistant profiles should default to a private user-and-agent bank unless an explicit sharing policy selects a broader scope.
Recalled memory is untrusted model context, not a system instruction or authorization grant. It cannot add tools, skills, credentials, network access, filesystem access, channel destinations, or execution capabilities.
5. Implementation Guide
To implement a “Learning Agent,” follow this sequence in your application logic:
- Ingestion: After an interaction or action result is durably accepted
into the agent session event stream, call
retainwith its source class and correlation. Never let a memory-write failure retry an effectful action. - Context Loading: Before calling the LLM, call
recallto fetch the most relevant 3-5 memories for the current prompt. - Scheduled Reflection: Run the
reflectoperation during idle time to compress raw experiences into high-level mental models.
Fine-Grained Authorization
The following is the raw material collected when I chat with AI about migrating the light-4j fine-grained authorization to light-rs to secure light-agent and light-gateway.
The information needs to be organized and documented into a design document for light-rs fine-grained authorization.
Question: We have a Java implementation of fine-grained authorization that we want to migrate to light-rs. In our java implementation, we have a middleware handler to verify the custom claims like role, group, attribute etc in the jwt token against endpoint specific rules defined in database and cached at the service. How should we handle the same requirement in rust?
Response:
In Rust, bridging the gap between dynamic, database-driven rules and high-performance web requests is a well-solved problem. If you are coming from Java, the architecture translates very cleanly, but takes advantage of Rust’s async runtime and thread-safe concurrency primitives.
To handle fine-grained authorization (Attribute-Based Access Control / ABAC or Role-Based Access Control / RBAC) where rules are cached locally but backed by a database, you have two primary paths in Rust:
- The Custom Middleware Approach (Closest to your current Java implementation)
- The Policy Engine Approach (Using an industry-standard Rust authorization crate)
Here is how you handle this requirement idiomatically in Rust, assuming you are using Axum (the current standard web framework in Rust).
Approach 1: Custom Middleware with Async Caching
In Java, you likely used something like Guava or Caffeine for your local cache. In Rust, the exact equivalent for high-performance, concurrent, async-aware caching is the moka crate.
Here is the architectural blueprint:
- Define your custom claims using
serde. - Setup a shared Application State holding your Database connection pool (e.g.,
sqlx) and your local Cache (moka). - Create an Axum Middleware to intercept the request, parse the JWT, consult the cache based on the route, and authorize the request.
1. Dependencies (Cargo.toml)
axum = "0.7"
jsonwebtoken = "9.0"
moka = { version = "0.12", features = ["future"] } # Async cache (Rust's Caffeine)
sqlx = { version = "0.7", features = ["postgres", "runtime-tokio"] }
serde = { version = "1.0", features = ["derive"] }
2. Defining Claims and State
#![allow(unused)]
fn main() {
use axum::{extract::State, http::Request, middleware::Next, response::Response};
use moka::future::Cache;
use serde::{Deserialize, Serialize};
use sqlx::PgPool;
use std::sync::Arc;
// 1. Define your custom JWT claims
#[derive(Debug, Serialize, Deserialize, Clone)]
struct CustomClaims {
sub: String,
roles: Vec<String>,
groups: Vec<String>,
attributes: std::collections::HashMap<String, String>,
exp: usize,
}
// 2. Define the Endpoint Rule (what you fetch from DB)
#[derive(Clone)]
struct EndpointRule {
required_roles: Vec<String>,
// add required_groups, attributes, etc.
}
// 3. Define the Application State to share across your app
struct AppState {
db: PgPool,
// Cache keyed by String (e.g., "GET:/api/data") returning an EndpointRule
rule_cache: Cache<String, EndpointRule>,
}
}
3. The Authorization Middleware
In Axum, middleware is written as an asynchronous function that takes the request, does some work, and either rejects it or passes it to the Next handler.
#![allow(unused)]
fn main() {
use axum::http::StatusCode;
use axum::response::IntoResponse;
async fn fine_grained_authz_middleware(
State(state): State<Arc<AppState>>,
req: Request<axum::body::Body>,
next: Next,
) -> Result<Response, StatusCode> {
// 1. Extract and decode the JWT from Authorization header
let auth_header = req.headers().get("Authorization").and_then(|h| h.to_str().ok());
let claims = decode_jwt(auth_header).ok_or(StatusCode::UNAUTHORIZED)?;
// 2. Identify the endpoint being accessed (e.g., "GET:/api/users/123")
// Note: In practice, use req.extensions() to get the matched generic route template
// rather than the raw URI so "/api/users/123" matches "/api/users/{id}"
let endpoint_key = format!("{}:{}", req.method(), req.uri().path());
// 3. Check Cache or Fallback to DB
let rule = match state.rule_cache.get(&endpoint_key).await {
Some(cached_rule) => cached_rule,
None => {
// Cache miss: Fetch from database
let db_rule = fetch_rule_from_db(&state.db, &endpoint_key)
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
// Populate cache
state.rule_cache.insert(endpoint_key, db_rule.clone()).await;
db_rule
}
};
// 4. Execute Fine-Grained Authorization Logic
let is_authorized = rule.required_roles.iter().any(|role| claims.roles.contains(role));
// ... check groups, evaluate custom attribute expressions, etc.
if !is_authorized {
return Err(StatusCode::FORBIDDEN);
}
// 5. Pass to the actual route handler!
Ok(next.run(req).await)
}
// Dummy helper functions to make the example compile mentally
fn decode_jwt(_header: Option<&str>) -> Option<CustomClaims> { /* ... */ unimplemented!() }
async fn fetch_rule_from_db(_db: &PgPool, _key: &str) -> Result<EndpointRule, ()> { /* ... */ unimplemented!() }
}
4. Wiring it up
use axum::{routing::get, Router};
#[tokio::main]
async fn main() {
let state = Arc::new(AppState {
db: PgPool::connect("postgres://...").await.unwrap(),
rule_cache: Cache::builder()
.max_capacity(10_000)
.time_to_live(std::time::Duration::from_secs(300)) // 5 minute TTL
.build(),
});
let app = Router::new()
.route("/api/secure-data", get(|| async { "Secret Data" }))
// Apply the middleware to all routes above this line
.route_layer(axum::middleware::from_fn_with_state(state.clone(), fine_grained_authz_middleware))
.with_state(state);
let listener = tokio::net::TcpListener::bind("0.0.0.0:3000").await.unwrap();
axum::serve(listener, app).await.unwrap();
}
Approach 2: Using a Dedicated Policy Engine (Recommended for complex rules)
If your endpoint-specific rules are highly complex (e.g., “User must be in Group X AND have Attribute Y = ‘finance’, UNLESS Role is ‘Admin’”), writing that evaluation logic manually in Rust gets messy.
Instead of writing custom evaluation logic, the Rust ecosystem heavily relies on external Policy Engines for ABAC/RBAC. The two leaders are:
1. AWS Cedar (cedar-policy) - Highly Recommended
Developed by Amazon Web Services in Rust, Cedar is a blazing-fast language for defining permissions.
- You store your Cedar policies in your database as text.
- Your Rust middleware fetches the Cedar policy string from the DB, compiles it (and caches the compiled policy using
moka). - You pass the JWT
CustomClaimsinto the Cedar engine as a context entity. - Cedar evaluates the claims against the policy in microseconds.
2. Casbin (casbin)
Casbin is a very popular authorization library that exists in Java, Go, and Rust.
- It supports ACL, RBAC, and ABAC.
- It has native adapters for
sqlx(Postgres/MySQL). You don’t even need to write the caching logic yourself;casbin-rshandles reading the rules from the database and keeping them cached in memory automatically.
Summary
- For simple Role/Group matching: Stick to the Custom Middleware approach using
mokafor async caching andaxum::middlewareto intercept requests. It is functionally identical to your Java approach but extremely fast. - For complex Attribute evaluation: Use your middleware to parse the JWT, but delegate the actual rule evaluation to
cedar-policyorcasbin.
Question: Could you please provide some examples for cedar-policy and casbin? For example, what is the format in database? In our system, we are using a yaml rule engine to evaluate the fine-grained authorization. We create some rules users can choose and they can also create their own rules if necessary.
Response:
Transitioning from a custom YAML rule engine to an industry-standard policy engine is a fantastic move. Both AWS Cedar and Casbin are designed precisely for this kind of “user-defined fine-grained rules” requirement, but they take very different approaches to how rules are written and stored.
Here is a breakdown of how both engines work, what their database format looks like, and how you map your JWT claims to them in Rust.
1. AWS Cedar (cedar-policy) - Best for Complex/Custom Rules
Cedar uses a custom, highly readable Domain Specific Language (DSL). It is specifically built to evaluate ABAC (Attribute-Based Access Control) using contextual data like JWT claims.
The Format (In the Database)
In your database, you do not need complex tables. You simply store the Cedar policies as plain text strings in a standard TEXT or VARCHAR column.
If a user wants to create a custom rule saying, “Allow access if the user’s JWT has the ‘finance’ group and a clearance level greater than 3,” the Cedar policy stored in the DB looks like this:
// Database column: policy_text
permit (
principal,
action == Action::"read_data",
resource == Resource::"secure_endpoint"
)
when {
context.jwt.groups contains "finance" &&
context.jwt.attributes.clearance_level > 3
};
For “predefined rules that users can choose,” Cedar has Policy Templates. You store the template once, and users just link their parameters to it.
// Template: "Require specific group"
permit(principal, action, resource)
when { context.jwt.groups contains ?required_group };
How you execute it in Rust
When a request comes in, you fetch the relevant Cedar text strings from your database, pass in the JWT claims as the Context, and let Cedar evaluate it.
#![allow(unused)]
fn main() {
use cedar_policy::{Authorizer, Context, Decision, Entities, PolicySet, Request};
use serde_json::json;
use std::str::FromStr;
fn evaluate_cedar_rule(db_policy_text: &str, jwt_claims: &serde_json::Value) {
// 1. Parse the text policy from the database
let policies = PolicySet::from_str(db_policy_text).expect("Invalid policy format");
// 2. Map your JWT claims directly into the Cedar "Context" using JSON
let context_json = json!({
"jwt": jwt_claims
});
let context = Context::from_json_value(context_json, None).unwrap();
// 3. Create the evaluation request
let request = Request::new(
Some("User::\"Alice\"".parse().unwrap()),
Some("Action::\"read_data\"".parse().unwrap()),
Some("Resource::\"secure_endpoint\"".parse().unwrap()),
context,
None,
).unwrap();
// 4. Evaluate
let authorizer = Authorizer::new();
let entities = Entities::empty(); // Used if fetching hierarchical data
let answer = authorizer.is_authorized(&request, &policies, &entities);
if answer.decision() == Decision::Allow {
println!("Access Granted!");
} else {
println!("Access Denied. Failed policies: {:?}", answer.diagnostics().errors());
}
}
}
2. Casbin (casbin-rs) - Best for Tabular/Database-Native Rules
Casbin works fundamentally differently. It splits authorization into two parts: the Model (a config file defining the structure of your rules) and the Policies/Rules (the actual data stored in the database).
The Format (In the Database)
Casbin doesn’t store a DSL string. Instead, it natively connects to your database using an adapter (like sqlx-adapter) and creates a table usually called casbin_rule.
The table structure looks exactly like a CSV file, with columns: ptype, v0, v1, v2, v3, v4, v5.
If a user selects a predefined rule to allow the finance_admin role to read /api/data, Casbin inserts a row into your database like this:
| ptype | v0 (Subject) | v1 (Object) | v2 (Action) | v3 (Condition - optional) |
|---|---|---|---|---|
p | finance_admin | /api/data | read |
If a user writes a custom ABAC rule using a UI (e.g., clearance > 3), Casbin can store the evaluation expression as a string in one of the columns (evaluated internally via the rhai scripting language):
| ptype | v0 (Subject/Condition) | v1 (Object) | v2 (Action) |
|---|---|---|---|
p | r.sub.groups.contains("finance") && r.sub.clearance > 3 | /api/data | read |
How you execute it in Rust
To make Casbin evaluate JWT attributes dynamically, you define a model.conf file that tells Casbin to evaluate (eval) the string stored in the database.
The model.conf File:
[request_definition]
r = sub, obj, act
[policy_definition]
p = sub_rule, obj, act
[policy_effect]
e = some(where (p.eft == allow))
[matchers]
# This matcher says: check if the object and action match,
# AND evaluate the ABAC rule stored in the database (p.sub_rule) against the request subject (r.sub)
m = r.obj == p.obj && r.act == p.act && eval(p.sub_rule)
The Rust Code:
#![allow(unused)]
fn main() {
use casbin::{CoreApi, Enforcer, DefaultModel};
use sqlx_adapter::SqlxAdapter;
use std::collections::HashMap;
// A struct representing your JWT Claims
#[derive(Clone)]
struct JwtClaims {
groups: Vec<String>,
clearance: i32,
}
async fn evaluate_casbin() {
// 1. Connect to DB. The adapter automatically fetches the rules from the `casbin_rule` table.
let adapter = SqlxAdapter::new("postgres://user:pass@localhost/db", 8).await.unwrap();
let mut enforcer = Enforcer::new("model.conf", adapter).await.unwrap();
// 2. Parse the JWT for the incoming request
let claims = JwtClaims {
groups: vec!["finance".to_string()],
clearance: 5,
};
// 3. Casbin in Rust uses `rhai` under the hood for dynamic object evaluation.
// You map your JWT fields into a format the engine understands.
let mut sub_map = HashMap::new();
sub_map.insert("groups", claims.groups);
sub_map.insert("clearance", claims.clearance);
// 4. Enforce the rules against the DB
// Request is: Subject (JWT Claims Map), Object (/api/data), Action (read)
if enforcer.enforce((sub_map, "/api/data", "read")).unwrap() {
println!("Access Granted!");
} else {
println!("Access Denied!");
}
}
}
Which one maps better to your YAML Rule Engine?
AWS Cedar is highly recommended for your use case.
Because you currently use a YAML rule engine where users define complex logical rules, your users are already accustomed to defining logical expressions.
- Translation: It is very easy to write a script that takes your existing YAML definitions and converts them into Cedar text policies.
- Safety: Cedar is statically typed and formally verified. If a user writes an invalid custom rule, Cedar will fail to parse it instantly and give you a helpful error message to show in your UI. Casbin’s
eval()logic can sometimes be brittle at runtime if the user types bad syntax. - Database Schema: With Cedar, you control your database schema. You just add a
policy_textcolumn to your existing “Rules” table. Casbin requires taking over a table (casbin_rule) to map to its strict CSV-like format.
Question: In our current implementation, we define all the persmissions and filters from UI and save the permissions and picked up rule in database. During server startup, it load everything locally in yaml format like the following. One of the pros is that our technical team write the rules, and the user or api owner just define the roles, groups, positions, attributes for the endpoint. The rule will be executed locally on the target server without going to the database so the performance is very good. If any permission changed for the target server, we can trigger a config reload from controller to ask the server to reload the config/permission. I am not sure if we should replace this implementation with Cedar. One of the feature we have is the column and row filter and I don’t think Cedar support it. What do you think?
rule.endpointRules:
echo@call:
req-acc:
- allow-role-based-access-control.lightapi.net
permission:
roles: mcp-reader
/v1/pets@get:
req-acc:
- allow-role-based-access-control.lightapi.net
permission:
roles: host-admin
/v1/accounts@get:
req-acc:
- allow-role-based-access-control.lightapi.net
res-fil:
- response-column-filter.lightapi.net
- response-row-filter.lightapi.net
permission:
col:
role:
teller: '["accountNo","userId","accountType","firstName","lastName"]'
account-manager: '["accountNo","userId","accountType","firstName","lastName","status"]'
row:
role:
teller:
- colName: accountType
colValue: C
operator: =
roles: account-manager teller
Response:
You have excellent architectural instincts. You are absolutely right to hesitate, and your concern about column and row filtering is the exact reason you probably should not replace your current system with Cedar.
Your current design is actually highly sophisticated. Let’s break down why your current implementation is so strong, why Cedar would struggle with your specific requirements, and how you should idiomatically build your current design in Rust.
Why Cedar is NOT a drop-in replacement for your system
AWS Cedar (and most standard policy engines like Casbin or OPA) are designed as Binary Decision Engines. You give them a Request, and they return an Allow or Deny.
The Missing Feature: Data Filtering (Masking) Your system handles two fundamentally different concerns in one elegant YAML file:
- Access Control (
req-acc): Can the user access the/v1/accountsendpoint? (Cedar does this perfectly). - Data Filtering (
res-fil): Which rows and columns can the user see in the response? (Cedar does not do this natively).
If you forced Cedar to do column filtering, you would have to misuse its “Annotations” feature to attach custom metadata to policies, read the output, and then write a custom Rust engine to parse that metadata and filter the JSON anyway. You would essentially be rebuilding your existing system, just with Cedar awkwardly wedged in the middle.
Strengths of your current architecture
What you have built is essentially a custom API Gateway / Sidecar Data Masking Engine.
- Performance: Loading rules into memory at startup means zero network latency. (Matches Cedar).
- Hot Reloading: Fetching updates via a controller is standard cloud-native practice. (Matches Cedar).
- Separation of Concerns: Tech team writes the underlying engine logic (
allow-role-based-access-control), while API owners just write YAML configs. (Matches Cedar’s Policy Templates). - Row/Column filtering: Natively understood by your engine. (Far exceeds Cedar).
Verdict: Do not switch to Cedar. Your current YAML-based, locally cached evaluation engine is a perfect fit for your business requirements.
How to implement your architecture in Rust
Since you are migrating to or building in Rust, your current Java-based architecture translates beautifully. In fact, Rust will make this system significantly faster and more memory-efficient.
Here is how you replicate your exact Java architecture—including the hot-reloading and filtering—in Rust.
1. Hot-Reloading using arc-swap
In Java, hot-reloading configuration safely across threads can be tricky. In Rust, the industry standard for “read-heavy, occasionally-updated” memory structures is the arc-swap crate. It allows your web server to read the YAML rules lock-free at millions of requests per second, while your controller can swap out the rules in the background instantly.
#![allow(unused)]
fn main() {
use arc_swap::ArcSwap;
use serde::Deserialize;
use std::sync::Arc;
use std::collections::HashMap;
// 1. Map your YAML exactly to Rust Structs
#[derive(Deserialize, Debug, Clone)]
struct EndpointRule {
#[serde(rename = "req-acc")]
req_acc: Option<Vec<String>>,
#[serde(rename = "res-fil")]
res_fil: Option<Vec<String>>,
permission: PermissionConfig,
}
#[derive(Deserialize, Debug, Clone)]
struct PermissionConfig {
roles: String,
col: Option<HashMap<String, HashMap<String, String>>>, // Map for roles to column JSON strings
}
// 2. Wrap your entire config in an ArcSwap
lazy_static::lazy_static! {
static ref CONFIG: ArcSwap<HashMap<String, EndpointRule>> = ArcSwap::from_pointee(HashMap::new());
}
// 3. Background task to listen to the Controller and Hot-Reload
async fn reload_config_from_controller(new_yaml_str: &str) {
let new_config: HashMap<String, EndpointRule> = serde_yaml::from_str(new_yaml_str).unwrap();
// Swaps the pointer atomically.
// All new requests see the new rules instantly. Old requests finish using the old rules.
CONFIG.store(Arc::new(new_config));
}
}
2. The Request Access (req-acc) Middleware
You would write an Axum middleware that intercepts the request, looks up the current endpoint in the arc-swap configuration, checks the JWT roles against permission.roles, and rejects the request if it fails.
#![allow(unused)]
fn main() {
// Inside your Axum middleware
let current_config = CONFIG.load();
if let Some(rule) = current_config.get("/v1/accounts@get") {
let required_roles: Vec<&str> = rule.permission.roles.split_whitespace().collect();
if !user_roles.iter().any(|r| required_roles.contains(&r.as_str())) {
return Err(StatusCode::FORBIDDEN); // req-acc fails
}
}
}
3. The Response Filtering (res-fil) Middleware
This is where Rust shines. Because Axum and Actix allow you to intercept the response body after the handler completes but before it goes to the user, you can apply your column filtering dynamically.
Using serde_json, you can strip out columns that the user isn’t allowed to see based on the col mapping in your YAML.
#![allow(unused)]
fn main() {
use serde_json::Value;
fn apply_column_filter(mut response_json: Value, allowed_columns_json: &str) -> Value {
// 1. Parse the allowed columns from the YAML string (e.g., '["accountNo","firstName"]')
let allowed_cols: Vec<String> = serde_json::from_str(allowed_columns_json).unwrap();
// 2. Filter the JSON object
if let Value::Object(ref mut map) = response_json {
// Retain only the keys that are in the allowed_cols list
map.retain(|key, _| allowed_cols.contains(key));
}
response_json
}
}
Summary
Keep your current system!
- The YAML approach is great. It acts as a perfect declarative contract between your API owners and your technical rules engine.
- Cedar is the wrong tool for Data Masking. Cedar is for strict Allow/Deny decisions, not mutating response bodies or filtering data.
- Rust will supercharge this. By using
serde_yamlto parse your config,arc-swapfor lock-free hot-reloading, andserde_jsonfor column masking inside an Axum middleware, you will get a system that does exactly what your Java system does, but with significantly less memory overhead and higher throughput.
User, Application, And Workflow Authorization
Historical design baseline. The A1 credential broker described below was retired by Workflow Invoke. Its browser callback and backend enrollment routes are no longer part of Workflow. For the current Gateway-to-Workflow user-token contract, see Workflow Invoke. The A1 qualification report and its activation receipt preserve historical evidence.
Historical status: accepted design, September 13, 2026. The A0 contract baseline v1 records the then-selected issuer profiles, contracts and migration inventory. A1–A3 qualification was pending at that checkpoint. This design was a foundation for issue #374 and the subsequent orchestration work. The unattended grant and renewal contract below was superseded by Workflow Invoke; its historical qualification claims do not describe the current deployment.
Scope And Related Designs
Cover user and immediate-caller identity at API/MCP and Resource/Knowledge
boundaries, original-token forwarding, scheduled and long-running execution,
credential renewal, revocation, and replacement of workflow-issued lad1
delegation tokens. Apply the same contracts to personal and enterprise profiles.
LLM User and Agent Authorization covers interactive model calls and explicitly excludes unattended delegation. Its user/workload separation, model assignments, billing principal, and alias restrictions remain applicable. Access Control and Fine-Grained Authorization provide the existing policy context; this design does not introduce a parallel role or ACL system.
The personal and enterprise orchestration designs own feature stages, reviews, budgets, and publication. Their grant references and reauthorization transitions depend on this design. Enterprise execution envelopes and accounting receipts remain separate from the user and app credentials specified here.
Decisions
- Preserve the original user access token in
Authorizationduring an interactive call chain within explicitly approved forwarding destinations, where issuer audience and sender-binding rules permit. The local issuer’s shared audience does not by itself approve a destination or calling app. Evaluate trusteduid,role,grp,pos,att, and other policy claims. - Put the immediate caller’s independently verified app credential in
X-Scope-Token. Every forwarding component supplies its own app token. - A valid user token does not bypass an endpoint’s allowed-caller policy. Gateway-only API/MCP endpoints reject direct agent, workflow, and test-tool calls unless a separate route profile explicitly authorizes them.
- Obtain fresh short-lived user credentials under a renewable, revocable grant for unattended work. The original expired JWT cannot remain the credential. Preserve the authenticated user identity and approved authority, not the original token bytes indefinitely.
- The OAuth issuer issues credentials. Workflow must not extend expiry,
disable expiry checks, sign copied user claims itself, or substitute a
service account when user authorization expires. Keep the issuer’s
auth_refresh_token_trecords and rotation, but load current user status and authorization claims on every refresh; saved claims are not continuing permission authority. - Separate credential replay protection, request authorization, and durable
effect idempotency. Removing
lad1must preserve or explicitly replace its request binding, replay, and workflow-permit protections.
Historical implementation boundary (2026-09-13)
Source inspected September 13, 2026. These are source observations, not a claim that current deployment settings or external identity providers are qualified.
| Area | Existing foundation and remaining gap |
|---|---|
| Workflow ingress | apps/light-workflow/src/rule_api.rs::authenticate verifies user and X-Scope-Token credentials separately; validate_invocation_caller checks configured service IDs, host, and environment |
| Workflow HTTP | executor.rs forwards stored user_authorization and the workflow service token; formatting these headers does not acquire a renewed user token |
| Interactive renewal | rule_api.rs::load_status can replace a stored bearer with a newer authenticated caller token after subject/claims checks; this depends on a live caller |
| Nested MCP | executor.rs still mints DelegationClaims from saved subject claims, with expiry bounded by the workflow deadline and now + 300; this does not refresh the user’s authority at the issuer |
| Gateway delegation | apps/light-gateway/src/main.rs verifies the workflow token and consumes a shared replay record before constructing the user principal |
| Agent execution origin | apps/light-agent/src/gateway_credentials.rs supplies the configured workload credential to each turn; it does not distinguish Chat from workflow execution. Separate Agent registrations, runtime credentials, and admission profiles are A2 work |
| Nested gateway protections | frameworks/light-pingora/src/mcp.rs derives child depth, execution class, and the synchronous permit pool from delegation; privateVersionTarget access requires its workflow invocation and tool_ref. All need verified replacements before removing lad1 |
| Workflow action decisions | Workflow already owns invocation/dependency/budget records; Gateway has delegation replay storage. The authenticated Workflow action-decision API and Gateway client specified below do not exist yet |
| Durable dispatch intent | agent_delegation_replay_t records consumed delegation IDs, not the proposed send-intent state machine. The Workflow-owned PostgreSQL dispatch ledger and Gateway write/recovery APIs below are A2 prerequisites |
| OAuth client runtime | frameworks/light-pingora/src/token.rs and spa_auth.rs implement acquisition/exchange/refresh paths. crates/light-client has configuration and other OAuth operations; a shared unattended credential provider remains integration work |
| OAuth issuer | In the portal-service repository, apps/light-oauth/src/main.rs has refresh and token-exchange handlers. Exchange is limited to configured MSAL/CCAC client profiles; this is not a general qualified workflow-grant endpoint |
| Issuer client authentication | light-oauth::authenticate_client checks client_secret; its current TLS setup does not supply verified client-certificate identity to token handlers. Issuer-side tls_client_auth, client certificate registration, and broker endpoint wiring are new A1 requirements |
| Issuer token purpose | generate_user_jwt and generate_client_jwt share the normal provider key and emit no purpose marker; only generate_long_lived_jwt selects the separate long-lived key. A1 must add issuer-owned token_use emission/reservation and shared verifier enforcement |
| Specialized issuer grants | client_authenticated_user accepts a trusted client’s user/role assertions; long_lived issues app tokens with registered/extra claims, including possible user-like fields. These supported features require explicit grant provenance and token-profile rules at workflow enrollment |
| Local user audience | generate_jwt_with_key uses the configured shared audience, default urn:com.networknt. Request-specific resource indicators are supported only for client_credentials; user-token forwarding needs a separate destination trust policy |
| Permission freshness | The issuer’s issue_access_token_from_refresh_token uses stored claims in both normal refresh and duplicate-retry handling. handle_refresh_token copies those claims into the replacement persisted by transactional replace_refresh_token; a tenant-bound current authorization-context query is required |
| Refresh recovery | The issuer authenticates the client before its configurable grace path returns the live replacement refresh token. The strict unattended profile needs issuer-enforced per-client/grant selection, consumed-token replay handling, and broker recovery changes |
| Workflow lifetime | workflow_invocation_t stores the bearer and expiry today. A protected renewal-secret store, grant references, and unattended reauthorization/revocation need implementation |
Inventory each affected API/MCP/Knowledge route before rollout. Existing support at Workflow ingress or the LLM gateway does not prove every resource verifies both identities or enforces a gateway-only caller policy.
Supported Issuer Grants And Token Profiles
Trusted-client custom claims remain supported for specialized integrations. Registered application metadata is owned by client administration; mutable user permissions are owned by their authoritative identity/policy source. Claims used for authorization need an explicit owner and update rule even when custom claims are uncommon. Do not require all application metadata to come from user tables, or let it silently override current user permissions.
client_authenticated_user remains a specialized issuer-approved assertion
grant. It is not used by the proposed Workflow broker and is not a fallback for
expired, revoked, or missing user authorization. long_lived remains supported
for app credentials in X-Scope-Token. Its custom uid/role/grp/pos/att
fields do not convert an app token into a user session or a WorkflowGrant.
For the unattended profile, the issuer restricts the broker client to its
approved enrollment and refresh grants. Workflow enrollment verifies grant
provenance through issuer-owned grant/session metadata, including grant type,
eligible client, subject/tenant, and consent. An ordinary signed JWT with a
uid is insufficient enrollment evidence. Build and qualify this metadata
lookup with the issuer in A1; the current handlers do not provide this contract.
Specialized clients retain their separately approved use cases; disabling those
features across the platform is not a prerequisite.
User and app verification profiles must be distinguishable using issuer-owned
token-purpose/key-purpose metadata or a qualified issuer lookup. Header position,
token lifetime, or the presence of uid alone is insufficient. Reject app
credentials presented as user credentials, including long-lived tokens carrying
user-like custom claims. These rules prevent substitution without removing
trusted-client customization. JWT validation profiles
The local A0 profile selects issuer-signed token_use=user|app, reserved against
custom-claim overrides, with A1 issuance and shared verifier implementation.
The A0 purpose contract
defines grant mapping and the app-only legacy long-lived key-purpose exception.
User purpose does not replace issuer provenance checks for Workflow enrollment.
User And Immediate Caller Identity
| Hop | Authorization | X-Scope-Token |
|---|---|---|
| Interactive Agent to light-gateway | Current user access token | Interactive Agent app token |
| Workflow Agent to light-gateway | Current user access token | Workflow Agent app token; action reference required |
| Workflow to light-gateway | Current user access token | Workflow app token |
| light-gateway to API/MCP | Current user access token | Gateway app token |
Each receiver validates signature, issuer, audience, expiry, and applicable
host/environment constraints independently for both tokens. Bind identities to
trusted issuer and tenant context; a bare uid from different issuers is not a
globally unique identity. Normalize claims through configured issuer mappings.
Never merge app roles into the user’s claims or accept unsigned identity headers.
Keep three concepts distinct: the application that obtained the login token,
the workload authorized to act for the user, and the immediate calling service.
The login JWT’s client_id does not establish the immediate caller. A trusted
service registration maps app-token identity, including the existing sid
where applicable, to the route’s allowed callers.
If a delegated JWT uses the RFC 8693 act claim, validate the current actor
under that issuer’s profile; prior actor history is audit information. It does
not replace authentication of the immediate hop. A profile that binds the user
token to a particular sender may require another issuer exchange before a
different service forwards it.
For gateway-only MCP access, both conditions are required: the current user may call the tool, and the authenticated app is an allowed gateway. The gateway evaluates the selected tool and request against existing fine-grained policy. The resource still enforces its user/resource ACL and caller restriction. Possession of a valid gateway token does not mean every user request is permitted.
Service-to-service forwarding replaces client-supplied scope headers with the forwarder’s own credential. Reject duplicate/malformed credentials and do not fall back from failed app authentication to user-only authentication. Direct interactive ingress uses its own explicit profile; a browser receives no gateway app secret to satisfy a downstream gateway-only rule.
Credential-Based Agent Execution Origin
For the initial profile, deploy separate interactive and workflow Agent service instances using the same binary but distinct app registrations, scope tokens, mTLS identities, and secret mounts. The workflow instance cannot access the interactive instance’s credentials. Trusted admission selects the instance; prompts, request headers, and tool arguments cannot select the credential profile. The interactive instance accepts the approved Chat ingress and rejects Workflow dispatch callers; it cannot provide a second route for running workflow jobs.
Gateway classifies the verified issuer/client identity through an administrator- owned registration mapping, checking its mTLS peer binding. Calls from Workflow or the workflow Agent registration always require an active action reference, including model, API/MCP, and Knowledge calls. Missing or invalid references deny the request even on an interactive URL. Root interactive admission accepts only its allowed registrations and ingress profile. Rewriting a mode header or path cannot turn a workflow credential into an interactive one.
The same verified app identity cannot be registered in both profiles in this pilot. A shared Agent runtime that holds both credentials and chooses between them from caller input is not qualified. A future shared-runtime design needs an independently verified job binding; it is not required for the pilot.
Gateway-Only Transport And App Credential Lifetime
A bearer token proves possession, not the physical origin of a request. Enforce the gateway-only boundary with backend ingress restrictions and authenticated service transport. The target profile uses mTLS, binding the allowed gateway workload to its peer identity and, where supported, certificate-bound app access tokens. A copied bearer alone must not allow an unapproved peer to impersonate the gateway. Any weaker rollout profile must state its residual replay exposure. Check that the authenticated peer maps to the same registered app as the scope token; accepting any certificate from the platform CA is insufficient. At a TLS terminator, propagate peer identity only through an authenticated internal path that strips externally supplied certificate/identity headers.
Use distinct app registrations and credentials for Gateway, Workflow, Agent,
and authorized test clients. The personal/local-issuer profile supports existing
long-lived app tokens with mTLS and app/peer binding. Short-lived
client_credentials tokens are an optional alternative; migrating to them is
not a prerequisite. Long-lived tokens still require expiry validation and an
operational revocation/replacement path. Keep official-environment tokens and
client keys/secrets in protected runtime configuration.
Checked-in app tokens and certificates may remain deliberate local-development fixtures so developers can start the stack without repeated credential setup. Treat them as public credentials: official environments must use independent keys/credentials and reject development fixtures. Native workers and test tools use their own registered identities; sharing a fixture gateway identity is not proof that gateway-only authentication has been qualified.
Shared Audience And Forwarding Destinations
The local issuer intentionally uses a shared platform audience because a user token traverses multiple services. Keep that audience and validate it at every receiver. Qualify a bounded set of trusted services for original-user-token forwarding through an administrator-controlled route/credential policy; an audience match or MCP registration alone does not enroll a recipient.
That policy binds each destination’s service identity and allowed origin/path to its credential profile. A demo or third-party MCP server does not receive the platform user bearer merely because it is behind Gateway. For destinations outside the approved set, use their own authorization flow or an issuer-qualified resource-specific exchange. The current local user-token grants do not provide that exchange; block user-token forwarding until an appropriate profile exists. Do not replace the missing profile with a service account and copied user claims.
Forward user tokens only to trusted configured destinations accepted by their issuer/audience and sender-binding rules. Do not disable audience validation or follow credential-bearing redirects to make forwarding work. Every recipient in the shared audience still enforces user policy and allowed-caller rules. A compromised approved recipient can expose a broadly usable user bearer; mTLS/app checks reduce where it can be replayed but do not make that bearer resource-specific. This is an explicit trust-domain tradeoff, not proof of isolation between every platform service. OAuth audience restrictions
Grant Ownership And Credential Storage
Use a trusted Workflow credential broker to mediate unattended user renewal.
For the first implementation it is a module hosted by light-workflow, using
shared OAuth client/provider code; it does not require a new deployed service.
It is separate from model prompts, native worker execution, ordinary workflow
inputs, and the artifact store. The OAuth issuer remains the token authority.
| Component | Responsibility |
|---|---|
| OAuth issuer or approved federation service | Authenticate the subject/client, approve renewable authorization, issue tokens, and enforce issuer-side expiry, revocation, and claim rules |
| Workflow credential broker | Redeem approved grants, protect renewal credentials, serialize renewal, verify issued tokens, and supply credentials to trusted outbound service calls |
| Workflow grant store | Persist user/issuer/client bindings, consent and resource ceilings, permitted workflow/schedule, grant version/status, deadlines, and credential reference |
| Gateway and resource services | Authenticate both identities, enforce current policy/ACLs, and validate applicable workflow/grant/action bindings |
| Controller/runner and native workers | Execute admitted work; do not receive refresh tokens, issuer client secrets, or signing authority |
A WorkflowGrant records its ID/generation, verified subject and issuer mapping,
host, authorized backend client, workflow definition/version or approved schedule,
allowed targets/resources/actions, authorization ceiling, consent evidence,
validity period, revocation state, and opaque renewal-credential reference.
These are proposed schema fields, not existing configuration keys.
Store refresh credentials in a dedicated encrypted credential store controlled by the broker; keep encryption keys outside workflow definitions and ordinary database rows. Persist only references and authorization metadata in workflow state. Neither a grant ID nor a workflow ID is bearer authority: broker access requires an authenticated allowed workload bound to that grant and execution.
Access tokens may be cached briefly in trusted service memory. Keep all reusable credentials out of prompts, conversations, review artifacts, issue comments, logs, and ordinary task results. Audit subject/app/grant/action identifiers, policy decisions, generations, and failure reasons without token bytes.
Interactive, Scheduled, And Long-Running Execution
Interactive execution may forward a currently valid original token. If it expires without an unattended grant, require authenticated renewal or pause. Do not turn a normal interactive login into unlimited background authority.
For unattended user-authorized work:
- While the user is present, obtain authorization for a bounded workflow or schedule. Use an approved confidential backend OAuth client to establish a renewable grant. Preserve the user’s identity and issuer-approved claims.
- Prefer an issuer-qualified token exchange profile that can issue a renewable grant for offline use. If the issuer instead requires a backend authorization code flow, complete that flow for the backend client. Do not copy a refresh token issued to the SPA into an unrelated client’s credential store.
- At a scheduled start, create the run bound to the grant and check its current status, generation, scope, and deadline. Acquire a fresh short-lived user token before dispatching protected calls; the old browser token is unnecessary.
- During execution, obtain or reuse a sufficiently fresh token for each target. Attach the immediate service’s own app token and apply current authorization and the run’s approved ceiling before every protected action.
- On one-off completion/cancellation, disable further use of that run’s grant binding. An approved recurring schedule may retain its parent grant until expiry/revocation; each scheduled run has separate execution/action identities.
sequenceDiagram
participant U as User
participant W as Workflow and credential broker
participant O as OAuth issuer
participant G as light-gateway
participant R as API or MCP resource
U->>W: Authorize bounded workflow or schedule
W->>O: Establish approved renewable user grant
O-->>W: Backend-bound renewal credential
Note over W,O: Later, the original access token has expired
W->>O: Redeem active grant with issuer-verified mTLS client authentication
O-->>W: Fresh short-lived user access token
W->>G: User token plus Workflow app token and action reference
G->>G: Validate identities, destination profile, and tool policy
G->>W: Fixed action-decision API with Gateway identity and user token
W->>W: Check live grant, run, dependency, and budget; claim dispatch
W-->>G: Bound decision, pinned target, depth, and dispatch lease
G->>G: Validate decision binding
G->>W: Begin dispatch for the same attempt and decision
W->>W: Recheck authority and atomically persist send intent
W-->>G: Acknowledge newly recorded send intent
G->>G: Acquire connection, complete TLS, and obtain required capacity
G->>G: Atomically check deadline and transition send guard at write boundary
alt Deadline valid and send guard acquired
G->>R: Accepted user token plus Gateway app token
R->>R: Validate caller and user/resource authorization
R-->>G: Result
G-->>W: Result and action evidence
else Deadline elapsed and request never initiated
G->>G: Atomically close send guard
G->>W: complete(NOT_INITIATED) for the same owner and generation
W->>W: Record outcome and release unused reservation
W-->>G: Safe to seek fresh authorization under retry policy
end
The final forwarding edge assumes an approved forwarding destination and a valid token for that resource. A required external-resource exchange must be qualified before enabling that route. The action-decision callback is the internal service operation defined below, not another workflow or MCP dispatch.
An organization-owned schedule may instead run under an explicitly authorized
service identity and its own permissions. Record the human creator for audit,
but do not fabricate their uid or impersonate an administrator. The execution
mode is fixed at authorization time; expiry cannot switch modes automatically.
RFC 8693 defines token exchange and permits refresh tokens for offline cases, but does not require issuers to provide them. A one-time exchange into another short-lived token is not an unattended renewal strategy. Issuer support for this profile must be qualified before enabling scheduled user execution. RFC 8693, section 2.2.1
Refresh-Token Claims And Authorization Freshness
Keep auth_refresh_token_t and refresh-token rotation. Change the source of
authorization claims when issuing a refreshed access token: query current
authoritative user status and permissions instead of treating the stored
snapshot as current authority. The record remains necessary to validate the
renewal credential, bind it to its user/client/provider/tenant and session,
preserve the approved grant scope, and support rotation and revocation.
This issuer-side record is separate from the Workflow broker’s protected store
of client-held renewal credentials.
Current Issuer Behavior
In portal-service, the normal username/password authorization-code login
already queries the database. portal-core::login_user_by_email assembles
roles, groups, positions, and attributes from user and relationship tables.
light-oauth::post_code copies these values into the authorization-code record;
handle_authorization_code then copies them into auth_refresh_token_t.
Creating that refresh record uses the code’s snapshot without another live
permission query.
On refresh, get_refresh_token_detail reads the refresh record, and
issue_access_token_from_refresh_token builds the access token from its saved
roles and other claims. handle_refresh_token copies those fields into the
replacement record. replace_refresh_token inserts the replacement, deletes
the old record using its aggregate version, and writes session/audit updates
in one transaction. Rotation therefore changes the credential, not the
freshness of the permission data.
The issuer also has a configurable duplicate-retry grace path:
find_recent_refresh_token_rotation can resolve a recently consumed token to
its replacement, and the same issuance helper mints another access token from
that replacement’s claims and returns the live replacement refresh token.
Client authentication happens first, so possession of the old token alone is
not sufficient for a confidential client. However, a presenter able to
authenticate as that client can follow the rotation within the grace window.
Rechecking claims does not restore theft detection. Retained interactive grace
profiles need current authorization checks; unattended grants use the strict
profile below.
A removed role can consequently survive repeated refreshes unless another mechanism revokes the session or grant. The exposure can last for the renewable session’s lifetime, not just one access token’s lifetime.
Alternatives And Tradeoffs
| Approach | Benefit | Cost or limitation |
|---|---|---|
| Reuse the stored claims snapshot | Simple; avoids user/permission joins during refresh | Permission and account-status changes can remain invisible across rotations |
| Load current claims on every refresh — initial choice | Direct freshness rule; future tokens reflect current account and permission state | Adds authorization-context queries and depends on that source being available |
| Cache claims with an authorization version | Reuses the snapshot while the current version matches; reduces repeated joins | Every relevant user, membership, role, group, position, attribute, and policy change must reliably invalidate the affected context; an unchecked version inside the token proves nothing |
| Revoke affected sessions/grants on permission changes | Forces a new authorization flow; useful for explicit revocation and security events | Disrupts users and background workflows; requires complete, race-safe invalidation and does not by itself invalidate issued access tokens |
Start with live claims on each refresh. Refresh already reads and writes the database; the incremental cost is resolving the current authorization context, not introducing the first database call. This occurs at renewal, not on every API request. Measure query cost before adding a versioned cache. Explicit session/grant revocation remains available alongside the live-query approach.
Required Refresh Behavior
- Authenticate the OAuth client and validate the refresh credential’s client/provider, subject/tenant, session/grant status, lifetime, and applicable sender binding. A successful credential lookup alone is insufficient.
- Load current account status and authorization context using the verified
issuer/provider mapping and the record’s stable
user_idandhost_id. Reject disabled, locked, removed, or otherwise ineligible users and invalid tenant memberships. Resolve currently effective roles, groups, positions, attributes, and other policy-bearing claims from their authoritative sources. - Issue a short-lived access token with that verified identity and current claims, constrained by current client policy and the existing OAuth grant. A refresh request cannot expand the originally granted scopes. Keep the workflow’s approved tools/resources/actions as a separate ceiling: newly acquired user permissions do not expand an already approved workflow. RFC 6749, section 6
- On normal refresh, atomically persist the replacement credential, consume the old one, and record the rotation/session audit before returning success. Preserve grant bindings and revocation semantics across generations. If claims columns remain, populate the replacement with the current snapshot; do not copy old permission values forward as authority.
- If a separately qualified interactive profile permits duplicate retries, recheck current status and claims before minting against the replacement. Never enter that grace path for an unattended grant. Neither profile may restore removed permissions or revive a revoked session.
- If the authoritative status/claims source is unavailable, fail renewal with an appropriate retryable outcome; do not fall back to saved permissions. Coordinate this with the broker’s bounded retry and expiry behavior below.
The implementation needs a dedicated tenant-bound authorization-context query.
The existing get_user_by_id returns NULL for roles, groups, positions, and
attributes, and selects the user’s currently chosen host. Reusing it would not
load the required claims. Nor should refresh simply call the email login query:
switching the user’s current Portal host must not move an existing grant to a
different tenant. Apply the relevant account, membership, and relationship
eligibility rules explicitly.
Classify stored claims during migration. Immutable grant bindings and approved
scope remain authoritative grant data; mutable permission claims must be
reloaded. Policy-bearing values in custom_claim need the same treatment as
role, grp, pos, and att. Existing snapshot columns may remain for
compatibility or controlled historical use, but renewal must not authorize from
them. Dropping the table or its claim columns is not required for this change.
For third-party identities, the live source may be the external issuer, an approved federation service, or a qualified synchronized authorization store. Portal’s database is authoritative only for claims it actually owns. Qualify provider mappings, account-status checks, and any synchronization delay before enabling unattended renewal.
Strict Rotation For Unattended Grants
The initial Workflow broker profile uses confidential-client authentication
bound to its registered mTLS workload/key, one-time refresh rotation, and no
replacement-token recovery through a consumed token. The issuer selects this
policy from protected client/grant metadata, including retained consumed-token
history; request parameters cannot opt into an interactive grace profile.
Per-client/grant selection is new A1 work, not an existing override of the global
refreshTokenRotationGraceSeconds setting.
Implement issuer-side OAuth tls_client_auth for the broker, not just mTLS
between Workflow and Gateway. The initial local-issuer profile terminates TLS
at a dedicated light-oauth broker listener, validates the client certificate
chain and validity, and passes verified peer identity to the handler through
server request context. Match the certificate’s registered subject/SAN to the
client/provider/tenant and enforce the configured certificate revocation policy.
Do not accept certificate identity from caller-supplied HTTP headers.
OAuth mTLS client authentication
Persist the broker client’s authentication method and certificate identity in issuer-controlled registration metadata. Every broker token, grant-lookup, and revocation operation requires that method. A correct client secret without the registered certificate fails, including at legacy token URLs; there is no secret-only fallback for this client. Other clients retain their configured authentication methods. A1 includes listener/trust configuration, handler peer identity plumbing, registration/schema publication, broker certificate/key loading, and certificate renewal/revocation qualification. This client authentication requirement does not automatically bind the forwarded user access token to the broker’s certificate; access-token sender binding is a separate issuer profile.
This protects against theft of the refresh token and client secret without the mTLS key. Full host compromise can still use or steal a software-held key; profiles requiring protection against key export need an isolated or non-exportable key facility. Do not claim that mTLS alone prevents that attack.
Keep enough family/session and consumed-token history to detect reuse across
rotations. An authenticated replay by the bound client returns invalid_grant,
revokes the affected refresh family, and records the event without returning
its live replacement. An unauthenticated or wrong-client request cannot obtain
replacement credentials or revoke another client’s family. Continue checking
current user authority on every successful refresh. Rotation and sender binding
serve complementary purposes. Refresh-token protection
The broker serializes renewal across replicas and durably records the attempt
and credential generation. If a response is lost after possible issuer commit,
or the broker crashes before saving the replacement, mark the credential state
uncertain, fence its use, and require reauthorization for the initial profile.
Revoke the uncertain family through an authenticated issuer operation. Do not
retry the old token or create a new grant automatically. Retry is safe only
when the issuer contract or transport establishes that no rotation occurred.
For the local provider, a connection-establishment failure (is_connect) or a
JWKS failure before the token POST is NOT_SENT. Only the same live renewal
owner/boot, attempt and generation may atomically restore an unrevoked grant to
ACTIVE, preserving its encrypted token and recording that outcome. Lost owner
proof, post-send failures and crashes still use the uncertain path. An HTTP error
response alone does not prove no commit.
Fetch/cache verification keys before rotating (five-minute cache lifetime), with one refetch for an unknown signing key ID. Publish new issuer keys before use. A failed refetch or invalid token after rotation remains uncertain. Recovery store errors are logged and retried on the next periodic tick; they must not close admission for unrelated workflows. Configuration errors remain startup failures. Save a valid rotation for its shared grant even if its requesting run was canceled, then deny that run the access token. Grant revocation and expired renewal ownership still prevent saving the rotation.
The retired enrollment API accepted an optional credentialBroker.legacyLongLivedAppKeys
list of approved issuer/key-ID pairs, default empty. This permitted only the A0
local long-lived app fixtures in X-Scope-Token; signatures, expiry, caller
service identity and explicit purpose markers remained enforced. It never relaxed
Authorization. This setting and enrollment API are no longer deployed.
A future recovery protocol must bind a durable refresh operation and its result to the original authenticated sender/key, recheck current grant state, and prevent a consumed token from retrieving its successor by itself. It requires separate qualification; an idempotency key or a time-based grace window alone does not supply that proof.
Permission Freshness, Revocation, And Recovery
Authorization requires all applicable checks: current user policy/ACL, allowed calling app, the approved grant/run ceiling, and the current action permit and budget. Newly acquired roles cannot silently widen the approved workflow. Removed roles, disabled accounts, and revoked grants must stop or narrow future actions according to current policy.
The live refresh query changes future tokens; it does not rewrite or invalidate an already issued access token. A resource that only checks a self-contained JWT may continue accepting its old claims until expiry. Earlier enforcement requires a current policy/status, authorization-version, revocation, or introspection check at the authorization boundary. A policy check using only the old JWT’s attributes cannot discover changed user membership. State the maximum enforcement delay for each qualified profile, including changes racing with token issuance; refresh-time checks alone are not immediate revocation. RFC 7009, section 3
Workflow and Gateway check current grant/run/action state before dispatching workflow-originated protected calls. Do not trust a caller-supplied invocation ID without checking the actor, tenant, target, request binding, and generation. For the initial unattended profile, failure to establish current grant authority blocks new dispatch. Any later caching profile must specify its maximum revocation delay; a self-contained JWT alone cannot prove immediate revocation.
Changes to stable user claims require authorization re-evaluation against the accepted disclosure ceiling. The current status API’s exact claims-digest check is not a general renewal protocol. Reauthorization records a verified new binding and decision; it cannot silently rewrite accepted inputs or reset workflow budgets, action identities, or review state.
Renewal is single-flight per credential generation, including across Workflow
replicas. Persist replacement credentials atomically with their generation.
A lost response during refresh-token rotation is uncertain credential state:
the initial unattended profile requires reauthorization as specified above.
Transient failures known not to have rotated the credential may retry within
bounded limits while credentials remain valid. After expiry, or when the grant
is fenced/revoked, stop protected dispatch and enter REAUTHORIZATION_REQUIRED
or an explicit retryable wait appropriate to the failure.
Reauthorization requires the same verified user/tenant and the necessary current grant. Grant revocation does not become permission to auto-create a new one. Unknown in-flight effects remain fenced and reconciled; renewed credentials must not replay those effects under a new attempt identity. Local saved work can remain recoverable while new model/tool/API actions wait for authorization.
Replay, Action Authorization, And Accounting
Short token lifetime limits exposure but does not stop a stolen bearer from being reused before expiry. Use protected transport, sender-constrained tokens where supported, and client-bound renewal credentials with appropriate rotation. OAuth Security BCP and certificate-bound access tokens describe those credential protections.
Separately, persist each workflow action’s identity and request digest, bound to the host, subject, app, invocation, target, and fencing/budget generations. Gateway validates an active server-owned action permit and records permit/deny with its policy/permit version and reason. Authentication success is not an authorization decision. A request reference selects state; it does not grant authority by itself.
For effects, an identical retry reconciles the recorded result or uncertain execution; the same action identity with a different request digest fails. Changing a client-supplied idempotency key cannot allocate an additional authorized workflow attempt or budget. Revalidate access before disclosing a cached result. Do not mark the user access token itself as consumed after one request: it can legitimately authorize several calls.
Keep invocation depth/class, policy and response ceilings, action reservations, budget ledger/generation, and live cancellation state in authoritative workflow records or separately qualified execution envelopes. Do not clone them into user tokens as mutable authorization truth. Gateway must resolve or validate them at decision time when replacing the current delegation verifier.
Gateway To Workflow Action Decisions
Implement fixed internal HTTP operations in light-workflow, proposed as
POST /internal/workflow/actions/authorize, .../begin-dispatch, .../complete,
and .../status, with a trusted Gateway client. These are service operations,
not model-visible tools, workflow definitions, or calls routed back through
Gateway. Workflow owns the state and transactions; Gateway receives bounded
results and needs no direct database credentials.
Before outbound dispatch, Workflow reserves an action against the admitted run and its pinned dependency/budget records. Calls carry an opaque action reference alongside the user and immediate-app tokens. Workflow and workflow-bound Agent registrations require this reference, using the credential-based origin rules above. Omitting it must not reclassify the request as a root interactive call.
Gateway validates the incoming user and app credentials, app/peer binding,
destination profile, and tool policy, then calls the fixed operation directly
over authenticated service transport. It sends the current user token in
Authorization and its own app token in X-Scope-Token; Workflow verifies both
and requires an allowed Gateway peer matching that app. This internal caller
profile permits only these fixed control operations, not arbitrary bypass of
Gateway-protected business APIs.
The request includes the action reference, invocation/tenant, canonical request digest, selected logical tool/operation and target, verified original calling app, and a durable Gateway dispatch-attempt/replica identity. Original-caller fields are assertions from this authenticated Gateway, checked against the action’s stored actor binding; public request headers are not authority. Workflow compares the user identity and current claims with the admitted grant and applies the claims-change rules above.
In one transaction against authoritative current state, Workflow locks the
relevant grant, run, action, and budget records, validates their generations,
deadline, cancellation/revocation status, actor, request digest, and pinned
dependency, and atomically claims the dispatch against its budget reservation.
It must not authorize from a lagging read replica or a Gateway cache. The
decision returns the bound subject/app/tenant, action and attempt IDs, grant/
run/budget/fencing generations, exact tool reference and target/contract digest,
execution class, parent/child depth where applicable, policy/disclosure ceilings,
and a short dispatch lease. Persist allow/deny evidence with the decision ID;
only an allow creates an AUTHORIZED dispatch-ledger row. The result is
authenticated service data, not a replacement user JWT or a
transferable bearer credential.
Durable Dispatch Ledger And Recovery
Use a new workflow_action_dispatch_t table in Workflow’s operational
PostgreSQL database, owned and migrated by light-workflow. It is distinct
from Gateway’s agent_delegation_replay_t and survives removal of that table.
Gateway durably records intent by calling the fixed API; no local file, memory
cache, or additional Gateway database is the authority. Database unavailability
blocks sending. A2 must build the table, APIs, Gateway client, and recovery logic.
The ledger binds tenant, action, canonical effect/dispatch-attempt ID, request digest, target, decision/lease generation, Gateway owner/boot identity and fencing generation, reservation, state, and outcome/evidence references. Uniqueness and compare-and-set transitions allow only one active dispatch attempt per action. Keep transition history and retain unresolved rows irrespective of lease expiry; retain terminal evidence for the invocation’s recovery/audit period. No tokens or private keys are stored in this ledger.
Before sending, Gateway calls begin-dispatch for the authorized attempt.
Workflow checks ownership, generations, lease, live grant/run/action state and
budget again, then atomically changes AUTHORIZED to SEND_INTENT. Only the
acknowledgement of that new transition permits the owner to send. A duplicate
call reports existing state; it does not issue another send permission. Gateway
records the outcome through complete, transitioning to a known terminal state
or UNCERTAIN, or recording retryable NOT_INITIATED as specified below.
Status/result reads recheck caller and disclosure authority.
Completion reporting uses a dedicated service profile: the authenticated
Gateway may record evidence for its stored send-intent binding after user-token
expiry or grant revocation. It cannot authorize another send or disclose a
business result. Such reporting does not require renewed user authority;
authorize/begin and business-result disclosure retain their current user checks.
complete(NOT_INITIATED) handles deadline expiry before the target request
starts, including during connection preparation or after a late send-intent
acknowledgement. Before reporting it, the original live Gateway owner must
atomically close a local send guard:
NOT_STARTED -> ABORTED_NOT_INITIATED competes exclusively with
NOT_STARTED -> STARTED. Keep the guard NOT_STARTED while acquiring the
pooled connection, completing TLS/protocol setup, and obtaining all capacity
permits. Preparation may establish a connection but must not send the target
operation’s headers or body, including through early data or automatic retries.
At the final transport write boundary, use one guarded operation to check the
original monotonic deadline and transition NOT_STARTED -> STARTED only if the
deadline is still valid and the send-intent acknowledgement is valid. Immediately
initiate the first request write on the prepared connection, with no intervening
queue, connection/permit acquisition, handshake, or asynchronous yield. This
guard belongs inside the transport at that boundary; wrapping a high-level
HTTP client’s queued send call is insufficient.
Disable automatic request retries for these routes throughout the outbound
transport, including retrying a failed reused connection on a fresh connection.
A send-intent acknowledgement permits only one guarded request initiation;
client or proxy retry logic cannot obtain another by resetting the local guard.
Expiry before this transition closes the guard as ABORTED_NOT_INITIATED and
allows the owner completion. Closing it prevents late callbacks or queued
continuations from sending. Once STARTED wins, failure or expiry remains
UNCERTAIN unless the outcome is known; the guard never authorizes a late write
or an automatic retry. Abort further initiation if the deadline has elapsed.
A transport timeout, absent result, or lack of a receipt is not proof of
non-initiation. If the guard was crossed or its state is unknown, use
UNCERTAIN instead.
Workflow accepts this completion only from the authenticated Gateway whose owner,
boot, fencing, decision/lease, action, and attempt identities match the current
SEND_INTENT record. Lease expiry alone does not prohibit this cleanup, but a
superseded owner/generation cannot assert it. In one transaction record
NOT_INITIATED, invalidate that send permission, and release the unused
reservation exactly once. A repeated accepted completion returns its receipt;
it cannot release a newer generation’s reservation or change another outcome.
The trusted owner’s assertion is required; Workflow cannot infer it from time.
Recovery is state-based across all Gateway replicas:
AUTHORIZEDwith no send intent: after fencing the old owner, fresh online authorization may replace the expired lease under the same canonical attempt and budget reservation. The old decision generation can no longer begin a send.NOT_INITIATED: the recorded owner completion proves that this send was abandoned before transport initiation. Under normal backoff and retry limits, fresh online authorization may reuse the same canonical action/effect attempt and request digest with a new decision/lease and reservation generation. It must reacquire available budget and recheck current grant/run authority; incurred costs and retry counters are not reset. No operator is needed.SEND_INTENTorUNCERTAIN: a send may have happened, even if an acknowledgement was lost. A status read or Gateway restart never authorizes another send. Reconcile through a target receipt or qualified target idempotency protocol using the same effect identity; otherwise require operator resolution.- Known terminal outcome: return or reconcile recorded evidence after current authorization checks, without executing the effect again.
Lost authorize/begin/complete responses and a crash before or after a network
send follow these rules. If the owner crashes before NOT_INITIATED is durably
accepted, a new boot or another replica cannot reconstruct the local guard and
assert non-initiation. If the completion committed but its response was lost,
the recorded NOT_INITIATED receipt is sufficient for the safe retry path.
Do not reset an ambiguous attempt to AUTHORIZED,
allocate another reservation, or infer that lease expiry proves no effect.
Admission failure known to precede send intent settles the unused reservation.
Qualified receiver reconciliation uses the existing complete operation, not a
new dispatch operation. It authenticates the receiver app and its exact mTLS
peer, checks that the receiver is registered for the action’s pinned tool, and
accepts only a terminal evidence digest for the same action and generation while
that generation is UNCERTAIN. It can settle the retained reservation after the
user token or run deadline expires, but it cannot authorize, begin, retry, read a
business result, or change an existing terminal receipt. An identical receipt is
idempotent; conflicting evidence fails closed.
Dispatch Lease Without Cross-Host Clock Comparison
There is no positive authorization cache. Gateway records local monotonic time
t0 immediately before sending authorize, and its send deadline is
t0 + 5 seconds. Count authorization latency, local admission, the
begin-dispatch round trip, connection/TLS/protocol preparation, capacity
acquisition, and all intervening waits against that same budget. Check the
deadline and transition the send guard together at the final write boundary
described above, after all preparation. A response received after the deadline
cannot permit a send.
Do not reset t0 on retries or on receipt of either response. Since the Workflow
decision occurs after t0, this conservatively bounds decision-to-send elapsed
time to five seconds without subtracting timestamps from different hosts.
Workflow separately expires its authorization lease at its own decision time
plus five seconds, or the earlier grant/run/action deadline, using its database
clock sampled after acquiring the relevant locks. begin-dispatch rejects a
late claim on that clock and rechecks current authority in the same transaction
as SEND_INTENT. It also returns the remaining authorized duration after
applying those deadlines. Gateway sets its deadline
to the earlier of t0 + 5 seconds and the monotonic begin-request start plus
that duration; it never extends the original limit. UTC times are useful for audit
and issuer expiry checks, not for reconstructing the Gateway elapsed timer.
Gateway restart, suspend/resume, migration, or loss of timer continuity
invalidates the outstanding local lease. Recovery consults the durable ledger;
it cannot rebuild a send deadline from wall-clock timestamps. A fresh lease
requires the safe AUTHORIZED or recorded NOT_INITIATED recovery path above.
Unknown references, invalid responses, unavailable storage, or elapsed deadlines
fail closed before sending.
An elapsed deadline after send-intent acknowledgement uses NOT_INITIATED only
when the original owner can close the unstarted send guard; otherwise it retains
the uncertain-execution recovery path.
begin-dispatch is the final authorization decision point. Revocation committed
before it denies the action; afterwards the action is admitted/in flight and
subject to cancellation/fencing. The lease bounds initiation of the request,
not network delivery, completion, or rollback of an already sent effect.
Long-running work rechecks permits at subsequent protected actions.
Nested Depth, Capacity, And Private Version Targets
Replace each use of context.delegation in Gateway’s workflow admission with
the verified action context, not fields copied from the user JWT or request.
Only an independently admitted root starts at depth zero. Workflow derives a
nested workflow’s depth with checked parent.permit_depth + 1, bounds it by
the parent grant/dependency ceiling and the target’s
maximum_delegation_depth, and preserves the permitted execution class and
remaining deadline. Missing context, depth overflow, or an unknown class denies
the nested call; it must not silently select root defaults.
Gateway continues selecting the synchronous permit pool by the verified child
depth. Preserve the existing nonblocking try_acquire_owned behavior: an absent
pool or exhausted capacity returns the defined capacity error. A parent holding
a synchronous permit must not make its child wait for that same pool. Async
children also obey the depth/lineage limits. When Gateway starts a child through
StartInvocationRequest, carry the parent action/decision reference; Workflow
revalidates it and atomically binds the child to that action and lineage.
Retries cannot create another child or reset depth, class, or budget.
For privateVersionTarget, Workflow resolves the exact target from the admitted
run’s pinned dependency registry. Gateway compares the verified stableToolRef,
private target name/version, endpoint, and contract digest with its published
mapping and authorizes using the logical authorizationToolName. An invocation
ID or tool reference supplied by a caller is insufficient. Keep private targets
out of public discovery; direct guesses, cross-run/tenant references, wrong
versions, and binding drift fail without revealing the private target. Valid
pinned calls remain reachable even when a public alias advances to a new version.
Implementation Order And Exit Gates
A0: Freeze Contracts And Issuer Profiles
The A0 baseline captures the initial local-issuer selection, current client inventory, contract records, claim ownership, transport inventory and deployment-manifest requirements. It selects backend authorization-code enrollment with PKCE S256 for the broker; the issuer must persist and validate the challenge and consent/provenance in A1. This baseline does not qualify a running stack or an external issuer.
Define the identity context, allowed-caller rules, grant/action schemas,
credential-store boundary, issuer mappings, renewal/revocation behavior, and
failure states. Inventory existing endpoint bypasses and all lad1 consumers.
Pin the supported light-oauth and third-party issuer profiles. Provider settings
or a token-exchange handler alone do not satisfy offline execution qualification.
Classify grant bindings versus mutable claims, identify each claim’s authoritative
source, and define tenant binding and permission-change enforcement delays.
Inventory trusted clients and their client_authenticated_user, long_lived,
exchange, and refresh permissions. Freeze the broker’s eligible grant provenance,
strict refresh policy, user/app verification profiles, shared-audience forwarding
destinations, separate Agent registrations/admission, issuer mTLS registration,
and the internal action-decision/dispatch-ledger contracts. Define the local
monotonic deadline and restart/fencing rules without a cross-host skew allowance.
A1: Implement Issuer Grants And Broker Renewal
See the A1 implementation and qualification report for source changes, test evidence and the remaining selected-stack activation gate.
Implement the A0 token_use marker in every local access-token issuance path,
including refresh and specialized grants. Derive it from validated grant
semantics, reserve it against registered/request custom claims, and implement
shared verifier checks for user versus app positions. Preserve only the explicit
legacy long-lived app key-purpose exception; markerless short-lived tokens do
not gain user authority from their claims or shared signing key. A2 must wire
these checks at every receiver in the selected profile.
Emit purpose as a dedicated JwtClaims field serialized as token_use, with
the validated purpose passed explicitly to generate_jwt_with_key. Keep the
name reserved in custom claims, but do not insert the issuer value into extra:
the helper applies remove_reserved_claims(extra) before signing and would
remove it. A1 tests must verify exactly one correct purpose marker in the
actual signed token after filtering and serialization, including override cases.
Implement the shared OAuth acquisition/provider layer and Workflow-hosted
broker, encrypted renewal storage, grant lifecycle, and serialized refresh
recovery. Extend/qualify the local issuer for backend-bound unattended grants,
revocation, and the live refresh behavior above. Implement the tenant-bound
authorization-context query for normal refresh and any retained interactive
grace profile. Implement broker-client grant restrictions, issuer grant-provenance
lookup, and strict unattended rotation with consumed-token history and family
revocation. The broker must fence uncertain rotations and require reauthorization
instead of recovering through the legacy grace path. Retain transactional
rotation and grant ceilings.
Implement tls_client_auth at light-oauth’s dedicated broker listener, verified
TLS-peer request context, registered certificate identities/authentication
methods, revocation checks, and broker key/certificate loading. Publish and wire
the listener, trust, registration, and certificate-rotation settings in the
selected stack. All broker grant/refresh/lookup/revocation paths enforce the
registered method, including requests sent to legacy endpoints.
Wire trusted client settings and runtime-secret mounts without placing client
secrets in workflow definitions.
Exit gate: after the original user JWT expires and the browser disconnects, a scheduled run and an hours-long run obtain valid credentials under the same approved identity/ceiling. Test restart, key rotation, concurrent renewal, lost refresh responses, revoked grants, disabled users, removed roles, and issuer outage. No stale-claim or service-account fallback is accepted.
For a multi-hour normal-load soak, record elapsed run-hours, renewal lead time, refresh attempts, successful rotations, uncertain rotations, and reauthorization counts/reasons. Report uncertainty per refresh attempt and the fraction of runs requiring reauthorization. Report injected lost-response cases separately from ordinary-load observations. With 600-second tokens, renewal occurs before each expiry, so an hours-long run crosses many rotation boundaries. Measure this operational cost without weakening strict rotation or treating a short, failure-free sample as a production reliability guarantee.
Issuer-specific checks must also prove that:
- A
client_credentialstoken carrying registereduid/roleclaims still hastoken_use=appand fails userAuthorizationverification. Registered and request-supplied purpose overrides cannot alter issuer purpose. Exercise all issuance/refresh paths, missing/unknown/conflicting markers, user tokens inX-Scope-Token, and markerless long-lived fixtures under the explicit app-only key-purpose profile. The exception never admits a token as a user. - Removing a role/group/position or changing an authorization attribute affects the next refreshed token, including through any retained interactive grace path; unattended grants can never opt into that path.
- Locked/deleted users, removed tenant membership, and revoked sessions cannot renew; changing the current Portal host cannot change the grant’s tenant.
- Added user permissions cannot expand the OAuth grant or workflow ceiling; permission-bearing custom claims cannot bypass the live query.
- A failed authorization query returns no token based on saved claims, and failed/concurrent rotation cannot return an uncommitted replacement.
- Authenticated reuse of a consumed unattended refresh token fails and revokes its family even inside the interactive grace window. Wrong-client and unauthenticated requests neither retrieve replacements nor revoke that family.
- The broker succeeds with its registered mTLS identity. Missing certificates, a different client’s certificate, invalid/revoked certificates, forged peer headers, and secret-only requests fail before issuing or rotating a token, including at legacy endpoints. Exercise approved certificate renewal and retirement; stolen refresh-token/client-secret material without the mTLS key cannot renew. Existing non-broker client authentication remains compatible.
- Lost refresh responses and a crash before broker persistence fence the grant and require reauthorization; no automatic old-token retry retrieves a successor.
- Workflow enrollment rejects app tokens containing user-like custom claims and ineligible assertion grants. The broker client cannot use specialized grants to manufacture a substitute user session; separately approved clients retain their supported specialized use cases.
- An existing access token and a permission change racing with refresh obey the profile’s documented enforcement delay; renewal is not claimed to revoke all previously issued access tokens.
A2: Enforce Both Identities And Replace Nested Delegation
Apply the per-route user/app contract to Workflow, Agent, Gateway, API/MCP, and
Resource/Knowledge paths that the feature uses. Replace nested MCP minting with
broker-provided user credentials and the workflow’s own app credential. Integrate
the fixed Workflow authorize/begin/complete/status APIs and Gateway client,
workflow_action_dispatch_t migrations and durable transitions, online dispatch
claims, audit decisions, and effect reconciliation. Implement the monotonic
send deadline, transport-level atomic send guard after connection/capacity
preparation, owner-only NOT_INITIATED completion,
and recovery/fencing rules. For the selected Pingora/reqwest outbound paths,
identify and qualify the transport write hook or adapter against pinned dependency
versions. A custom connector alone is sufficient only if it controls that final
write boundary, including reused connections. Explicitly wire and verify the
route-specific retry settings at every client/proxy layer; a handler-level
guard or reliance on default retry behavior does not qualify the path.
Provision separate interactive and
workflow Agent registrations/instances, isolated credentials, peer mappings,
and admission policies; Gateway classifies origin from verified credentials.
Replace delegated depth/class and private-target checks with the verified action
context, including Workflow-side child-lineage validation.
Exit gate: direct calls using a valid user token and an unauthorized app fail; permitted gateway calls still require user tool/resource permission. A copied gateway bearer from an unapproved transport peer fails the qualified gateway-only profile. Test wrong issuer/audience/host, duplicate headers, missing claims, renewal during a tool loop, replay across replicas, changed request payloads, revoked action permits, exhausted budgets, and cancellation races.
Additional exit cases:
- Approved shared-audience service chains retain original-token forwarding; an unapproved demo/third-party target receives no platform bearer. App tokens cannot authenticate as users, and valid certificates from the wrong workload cannot authenticate as Gateway.
- A workflow Agent credential with no action reference, a forged interactive header, or an interactive URL is denied. Workflow callers cannot enter through the Chat Agent or select its credentials. Normal Chat still succeeds with its own registration; workflow Agent calls succeed only with active bound actions.
- Nested sync and async calls retain increasing depth and inherited limits; missing context, overflow, and excess depth deny admission. Saturating a root sync pool still permits a qualified child in its separate depth pool; a full child pool fails promptly without waiting on the parent. Restart/retry cannot reset depth or create a second child.
- Valid pinned private-version calls succeed and stay absent from discovery. Guessed IDs/names, wrong tenants/runs/versions, and dependency/contract drift fail; updating the public alias does not redirect an accepted pinned call.
- Across Gateway replicas, duplicate action claims authorize at most one dispatch. Workflow outage, stale generations, altered request digests, expired leases, and a lost decision response cannot bypass authorization or replay an uncertain effect. Crash before send intent, after its commit but before its acknowledgement, after sending, and before completion persistence. Recover on a different Gateway replica from the Workflow ledger; never resend a possibly executed effect without qualified reconciliation. Repeat with the old delegation replay table absent and with the new store unavailable.
- Use independently controlled wall clocks and local elapsed timers: large host
clock offsets and wall-clock jumps cannot extend Gateway’s five-second bound.
Delayed authorize/begin replies, lock waits, local pauses, and earlier grant
deadlines cause timely rejection. Restart or timer discontinuity invalidates
local leases. Verify late begin claims fail on Workflow’s own clock and retries
cannot reset the original timer. Test revocation before and after the
SEND_INTENTtransaction, with no cached approvals. - Delay the send-intent acknowledgement beyond Gateway’s deadline: assert zero
target requests, accepted owner-only
NOT_INITIATED, one reservation release, and successful fresh authorization of the same canonical attempt without an operator. Also delay pooled-connection acquisition, TLS/protocol readiness, and capacity preparation after send intent: the guard staysNOT_STARTED, expiry producesNOT_INITIATED, and no target operation headers/body are sent. Race readiness and the deadline at the final write boundary: exactly one local guard transition wins, with no further queueing before the write. A started request cannot reportNOT_INITIATED; expiry after that transition cannot trigger a late write or automatic resend. Duplicate completions cannot double-release; stale owners/boots/generations are rejected. Test a lost completion response, a crash before completion persistence, and grant revocation before retry: only recorded non-initiation enables this safe retry of aSEND_INTENTattempt, and fresh authorization still enforces revocation and retry budgets. - Fail a reused upstream connection at the first write and after a partial
write. Verify the actual configured transport never replays the operation on
a fresh connection, even if the request body is replayable. Count outbound
request initiations as well as target observations: only one initiation may
cross the guard, and an unknown outcome remains
UNCERTAIN. Exercise each qualified outbound client/proxy path with its effective retry settings.
A3: Migrate, Remove Old Authority, And Qualify Orchestration
Migrate active invocations to valid grant references through authorized reauthorization or drain them; an expired persisted bearer is not enrollment proof. Stop persisting reusable access tokens in ordinary invocation state. Coordinate sender/receiver rollout so routes cannot fall back to weaker authentication when a new credential fails.
Inventory the app tokens, issuer/signing trust, mTLS identities, and forwarding
profiles used by the selected stack. Keep deliberate local-development fixtures
in Git; official installations use independent runtime credentials and reject
those fixtures. If an official installation has reused development credentials
or exposed its own tokens, replace/revoke the affected credentials and remove
their trust before qualification. A long expiry alone is not a reason to remove
the supported app-token profile, and blanket rotation of development fixtures
is not a prerequisite. Test the actual user/app/certificate combinations at the
official boundary rather than relying on a dev label.
Remove workflow delegation signer/verifier wiring, its deployment secrets,
and crates/agent-delegation/agent_delegation_replay_t only after verifying no
remaining consumer and qualifying their replacement protections. Preserve
active evidence and recovery behavior during schema/config migration. Install
and qualify workflow_action_dispatch_t and all fixed dispatch APIs before
removing the old replay table. Drain/reconcile old delegation attempts first;
their replay rows cannot be treated as evidence of send outcome. The new ledger
and its unresolved attempts are retained across A3 and subsequent restarts.
Exit gate: run the complete authenticated API/MCP/Knowledge paths, scheduled execution, renewal, revocation, and effect-recovery cases on the selected stack. Verify that specialized grant support, local shared-audience forwarding, and development startup remain compatible with the explicitly selected profiles. Then admit development orchestration Phase 1. Enterprise accounting and sandbox qualification remain additional profile gates, not replacements for this work.
References
- Issue #374
- LLM User and Agent Authorization
- Workflow-Backed MCP Tools
- Stateless Auth Handler
- MSAL Exchange Handler
- Token Handler
Authorization A0: Contracts And Issuer Profiles
Frozen historical baseline. The browser authentication, registered callback, and credential-broker provisioning below describe the original A0 proposal. Workflow Invoke later retired both browser callback and backend broker enrollment. See Workflow Invoke for the current user-token and LONG registration contracts. The frozen text below remains for provenance and is not current deployment guidance.
Status: A0 contract baseline v1 frozen after review, September 13, 2026. The
qualified-receiver complete variant recorded below is part of the frozen A0
contract; it does not change the frozen operation set or authorize sends.
It defines the initial implementation contracts for the accepted
authorization design.
It records source and local registration evidence, not completed A1–A3 runtime
qualification. Personal orchestration Phase 1 remains gated on those phases.
The baseline selects portal-config-loc/all-in-lt with local light-oauth for
the first unattended user profile. light-portal-install must publish the same
contracts when that distribution is qualified. External issuers and official
environments need their own completed profile manifest and qualification record;
neither inherits approval from a working local login.
Baseline And Evidence
| Repository | Inspected HEAD |
|---|---|
light-fabric | 3a13e20fd7a522d73a6d64a02e6e735ad681af73 |
portal-service | 82895c28bcf41c9a0c65c54596d8cc5d8b20041c |
portal-config-loc | 161ae97e5d1b1ea0beaf6d3786f31f1220cdbda5 |
light-portal-install | 50c6121e2dbc34f97a776df3edb69378def04427 |
These identify inspected source, not a claim that running images were built
from those commits. The running Compose project was confirmed as
portal-config-loc/all-in-lt, including its runner credentials override.
No credentials are included in this record or the accompanying inventory.
The client inventory
contains the 32 active provider/client bindings read from local
configserver.auth_client_t joined to configserver.auth_provider_client_t for
host 01964b05-552a-7c4b-9184-6857e7f3dc5f. Both rows must be active. The export
includes IDs, names, types, profiles, exchange types, scopes and aggregate
versions; it excludes secrets, custom-claim values and user information.
Thirty-one bindings use provider AZZRJE52eXu3t1hseacnGQ; one uses
AZ7MrzXYcz2X8kdd44FPLw. Only the former is selected below.
For the selected provider, 30 clients are trusted and one is confidential.
Light Portal Client has exchange type msal; pylon has ccac; the other
bindings have no exchange type. These are observed registrations, not proof
that their traffic uses every grant available to them. In current light-oauth:
password,client_authenticated_userandlong_livedcheckclient_type == trusted;client_profile == servicedoes not disable them.- Exchange is selected by
token_ex_type. It is not a general offline grant. - Refresh requires an authenticated client and a matching refresh record. The registry has no broker-specific strict-rotation or mTLS-authentication fields yet. A1 must add and enforce those restrictions.
Selected Issuer Profiles
The following names identify this document’s profiles, not existing config keys.
| Profile | Enrollment and credentials | Permitted use |
|---|---|---|
local-workflow-user-v1 | Dedicated confidential broker client; user-present backend authorization_code, then refresh_token; issuer-verified tls_client_auth | Initial scheduled and long-running user workflow profile, enabled only after A1–A3 |
local-app-v1 | Existing issuer-approved long_lived app tokens plus registered mTLS workload identity | Immediate app identity in X-Scope-Token; never user enrollment authority |
| Existing interactive/specialized profiles | Current supported interactive grants, client_authenticated_user, client-credentials and MSAL/CCAC exchange | Preserve their separately approved uses; no automatic conversion to local-workflow-user-v1 |
| External unattended profiles | None selected in v1 | Disabled until an issuer-specific renewable grant, identity mapping, freshness policy and recovery protocol are qualified |
The selected local provider uses issuer urn:com:networknt:oauth2:v1, shared
user audience urn:com.networknt, and provider ID AZZRJE52eXu3t1hseacnGQ.
The checked-in public issuer base is https://oauth.localhost; the provider’s
discovery URL is not a substitute for the configured JWT issuer. Retain the
issuer’s RS256 verification profile and separate long-lived app signing-key
purpose. Issuer-owned purpose/provenance must distinguish user tokens from app
tokens, including short-lived client-credentials tokens; uid, sub, token
lifetime or the header carrying a token is not sufficient evidence.
Freeze the local purpose marker as an issuer-signed token_use claim with
exact values user and app. A1 adds it to every local access-token issuance
path: user grants and their refreshes emit user; client_credentials and
long_lived emit app. Specialized user assertion/exchange grants retain their
supported user-token purpose, but that marker alone never proves eligibility
for Workflow enrollment; issuer grant provenance remains mandatory.
The issuer derives purpose from the validated grant/issuance path. Reserve
token_use against both registered custom claims and request extra claims, and
emit it as a dedicated JwtClaims field, separate from filtered extra.
generate_jwt_with_key calls remove_reserved_claims(extra) internally, so
inserting the reserved marker into extra would strip it before signing.
Pass the validated purpose explicitly into issuance and serialize it once as
token_use. Refresh must not derive it from a saved custom-claim snapshot.
A1 implements shared verifier
checks; A2 wires them at every selected receiver: user Authorization requires
user, and X-Scope-Token requires app. Unknown, malformed or conflicting
purpose is rejected. Markerless user or short-lived app tokens require new
issuance; they cannot fall back to uid, role or the shared provider key.
Existing markerless long-lived app fixtures may retain the explicitly registered
issuer/key-purpose app verification profile, only in the app position. A
conflicting marker is rejected even there; the normal shared signing key never
qualifies for this exception.
For local-workflow-user-v1, freeze these choices:
- Create a dedicated broker registration. Allow only its backend code grant
and refresh grant; reject
password,client_authenticated_user,client_credentials,long_livedand exchange for this client. Do not repurpose either existing Workflow app registration from the inventory. - The issuer authenticates the user and records explicit workflow/schedule
consent. Bind the code to the broker, provider, user, tenant, exact registered
callback and enrollment intent. Use one-time state and PKCE S256, with the
verifier retained by the broker. The browser supplies neither a refresh token
nor a broker credential. A1 must implement the complete challenge path:
current
post_codestorescode_challengeandchallenge_methodasNone, even though redemption has averify_pkcehelper. Test code substitution, missing/wrong verifier, replay and consent/tenant mismatch. See PKCE and authorization-code protections. - Use the dedicated issuer TLS listener for broker token, provenance lookup and revocation operations. Match the actual verified certificate to issuer registration; reject secret-only broker requests at every legacy URL too. This is mTLS client authentication, not automatic certificate binding of the forwarded user’s access token.
- Issue a bounded renewable family for this consent and backend client. Keep
auth_refresh_token_t; apply live tenant-bound user/claim lookup on renewal, strict one-time rotation and no consumed-token successor recovery. A1 must expose issuer-owned provenance and authenticated family revocation. An existing refresh token with no eligible provenance cannot be enrolled. - Keep the initial user access-token lifetime at 600 seconds, matching current local issuance. Grant expiry and any shorter issuer limit cap renewal. An hours-long workflow does not obtain a longer-lived user access token.
- Preserve the shared audience only inside explicitly approved forwarding destinations. Start with an empty destination allowlist until A2 publishes exact routes and receiver identities. Demo or third-party targets do not receive the platform user token merely because Gateway can reach them.
The broker’s callback address, new client UUID, certificate identities, trust roots and secret references are deployment bindings, not values to guess from client names. A1 provisions them in the versioned manifest below. Enrollment must fail while a required binding is absent.
Claim Ownership And Freshness
| Claim or binding | Authority | Refresh/action rule |
|---|---|---|
| Issuer/provider, user ID, tenant, broker client, consent, family and scope ceiling | Issuer grant/session plus accepted Workflow grant | Immutable binding; tenant cannot follow the user’s currently selected Portal host |
| Account active/locked/verified and tenant membership | Current user_t and user_host_t, with applicable eligibility rules | Recheck on issuance/renewal; disabled or removed users cannot renew |
role | Current role_user_t / role_t for the grant tenant | Rebuild eligible memberships; snapshot roles are not authority |
grp | Current group_user_t / group_t for that tenant | Rebuild eligible memberships |
pos | Current employee_t, user_position_t / position_t | Rebuild eligible positions in the grant tenant |
att | Current attribute_user_t / attribute_t | Reload values used by policy |
| Custom policy-bearing claims | Explicit registered issuer/policy source per claim | Reload or fail; no fallback to a saved value |
| Custom application metadata | Client administration | Preserve supported metadata without allowing it to override user identity or permission authority |
| Run/action budget, depth, target and disclosure ceiling | Authoritative Workflow records | Resolve at authorize/begin; never take mutable values from copied JWT claims |
portal-core::login_user_by_email shows the relationship sources but selects
the current host and has login-specific filtering. get_user_by_id returns
NULL permission columns and also selects the current host. A1 needs a dedicated
authorization-context query; neither helper is the refresh contract unchanged.
Absent memberships mean an empty permission set, not permission to reuse the
old snapshot. Failed lookups issue no replacement token.
Freeze the initial freshness mode as live refresh, existing access tokens bounded by expiry. Re-read applicable live policy/ACL data at each resource decision. This does not discover a changed claim if that policy only consumes the old JWT. The deployment manifest must record validator expiry leeway and qualified issuer/receiver clock error; the worst-case old-claim interval is the 600-second token lifetime plus those allowances. No immediate user-role revocation claim is made. A stronger online user-status/version profile needs separate qualification, not an undocumented cache assumption.
Workflow grant/run/action revocation has no positive authorization cache: authorize and begin-dispatch read current primary-store state. A revocation committed before begin denies a new send intent; a later revocation follows the accepted in-flight cancellation rules. The local five-second initiation deadline has no cross-host clock-skew allowance.
Identity And Caller Contract
All new service records use contractVersion: 1, camelCase property names,
explicit tagged variants and rejection of unknown versions/fields. UUIDs are
opaque identities; counters are nonnegative integers with checked increment,
and overflow fails closed. Timestamps are UTC audit/expiry values. They never
reconstruct a Gateway monotonic lease after restart. An ID alone grants nothing.
VerifiedIdentityContext is constructed by trusted credential validation,
never deserialized as authority from public request headers:
| Field group | Required content |
|---|---|
user | issuerProfileId, issuer, providerId, subjectId, hostId, tokenClientId, tokenExpiresAt, claimsDigest |
callerApp | issuer, providerId, clientId, serviceId, hostId, environment, registrationVersion, peerIdentity |
origin | INTERACTIVE, WORKFLOW, or SYSTEM, derived from the verified registration; system execution never fabricates a user |
actionBinding | Required for every workflow-origin protected call: stored action reference, invocation and admitted actor/job binding |
For the SYSTEM variant, a verified service subject replaces user; it has
no impersonated user claims. That separate execution profile is not enabled by
selecting local-workflow-user-v1.
The local user subject is the verified uid mapped under the issuer profile;
client_id/cid identifies the login client, not the immediate workload.
The app sid is checked against registration, client and actual mTLS peer.
Conflicting identity claims fail instead of being merged. Existing Knowledge
subject/group/organization normalization and disclosure rules must be preserved
by the issuer mapping and verified policy context, not caller assertions.
| Receiver/operation | Allowed immediate caller for v1 | Additional checks |
|---|---|---|
| Gateway business/model/API/MCP/Knowledge ingress | Explicit interactive clients, Workflow, or isolated workflow Agents | Verify user and caller independently; Workflow/workflow-Agent origin always needs an active action binding |
| Gateway-only business API/MCP/Knowledge receiver | Gateway registration and matching peer | Current user policy/resource ACL plus exact forwarding destination; direct Agent/Workflow/test-tool calls denied |
| Workflow action authorize/begin/status | Allowed Gateway and matching peer, directly | Current user and stored action/actor/tenant/attempt binding; not a general business-API bypass |
| Workflow action completion | Recorded Gateway owner, or an exact registered receiver/peer for qualified reconciliation | Can report evidence after user expiry; receiver evidence settles only the same uncertain action generation and cannot authorize a send or disclose results |
| Interactive Agent | Registered interactive ingress | Reject workflow jobs; do not expose interactive credentials to workflow execution |
| Workflow Agent | Trusted workflow admission bound to the selected job | Isolated registration, mTLS identity and secret mounts; cannot select interactive credentials |
| Issuer broker operations | Registered broker certificate/client | Backend consent/family/tenant binding; no secret-only fallback |
For the initial gateway-only profile, migrate Agent Knowledge calls through
Gateway with the same dual-token and disclosure checks. Do not leave the old
direct Agent-to-Knowledge lad1 route as a workflow bypass. Any future direct
interactive Knowledge route needs an explicitly qualified separate profile.
Native personal workers remain consumers of admitted local work; they receive
no refresh credentials or authority to choose a caller profile.
Grant, Action And Credential Records
These are new logical schemas. A1/A2 migrations and shared Rust types implement them; this document does not claim the fields already exist in configuration.
| Record | Frozen fields and invariants | Owner |
|---|---|---|
WorkflowGrant | grantId, generation, issuerProfileId, issuerGrantRef, subjectId, hostId, brokerClientId, consentRef, scopeCeiling, resourceCeilingRef, workflowOrScheduleBinding, notBefore, expiresAt, status, credentialRef; references are tenant-bound and versioned | Workflow grant store; issuer provenance verified by broker |
WorkflowRunAuthorization | invocationId, grantId, grantGeneration, runGeneration, accepted workflow version/digest, admitted actors, effective ceiling, budgetLedgerId, budgetGeneration, deadline and cancellation state | Workflow; recurring schedules create distinct run bindings |
WorkflowAction | actionId, invocationId, grantId, actor/job binding, parentActionId, pinned logical operation/target/contract, requestDigest, policy/disclosure references, permitDepth, executionClass, reservation, deadline, generation and state | Workflow; untrusted tool arguments cannot allocate their own authority |
BrokerCredential | credentialRef, grant/client/provider/tenant/family binding, encrypted refresh material, encryption-key reference, generation, renewal owner/fence, renewal-attempt identity and state | Dedicated broker store; no access from workflow DSL, runners or artifacts |
DispatchDecision | decisionId, action/attempt/request/identity bindings, all grant/run/budget/owner generations, target/contract, depth/class/ceilings, reservation and lease | Workflow decision transaction; not a bearer credential |
DispatchReceipt | Decision/attempt/generation, previous/new state, completing owner/boot/fence or qualified receiver identity, outcome/evidence reference, released reservation and audit time | workflow_action_dispatch_t and its retained transition evidence |
Grant admission requires ACTIVE; REAUTHORIZATION_REQUIRED, REVOKED and
EXPIRED forbid new protected actions. Reauthorization binds the same verified
user/tenant under a new generation without resetting budgets or uncertain
effects. One-off run completion retires that run binding; a recurring schedule’s
parent grant survives only within its original consent and expiry.
The broker stores a durable renewal attempt before redeeming a refresh token. Only the lease/fence owner may commit its replacement generation. A crash or lost response after possible issuer commit fences the credential as uncertain, requests family revocation and requires reauthorization. A database restore must not resurrect an older refresh generation. Revocation retry cannot make the fenced credential usable while the issuer is unavailable.
Encryption keys are outside ordinary database rows; ciphertext is separate from
ordinary workflow state. Access tokens may exist briefly in trusted memory.
Prompts, worker sessions, artifacts, logs, results and GitHub comments contain
no reusable credentials. The issuer keeps its own refresh/family history;
the Workflow broker store does not replace auth_refresh_token_t.
Action Decision And Dispatch Contract
Freeze the four fixed internal POST operations from the accepted design:
Path suffix under /internal/workflow/actions/ | Request binding | Successful result |
|---|---|---|
authorize | Action/invocation/tenant, verified original caller, request/target binding, canonical attemptId, gatewayOwnerId, gatewayBootId, expected fencing generation | decisionId, exact bound decision and AUTHORIZED state; denial records evidence without an authorized row |
begin-dispatch | Same attempt/decision/owner/boot and expected grant/run/budget/lease/fencing generations | First successful AUTHORIZED -> SEND_INTENT acknowledgement; remaining lease duration can only shorten Gateway’s original deadline |
complete | Decision/attempt/generation and either the recorded owner binding or an exact registered receiver/peer for an already UNCERTAIN action, plus tagged outcome and evidence reference | Durable receipt; identical duplicate completion returns the same receipt without another release; receiver evidence cannot begin or retry dispatch |
status | Tenant/action/attempt and authenticated disclosure context | Current recorded state/evidence; never a new permission to send |
Workflow creates action records against admitted run/job and pinned dependency records before protected dispatch. The executor and trusted Agent adapter pass the resulting opaque action reference; it is not model-selected authority. Each concrete protected operation has its own binding. Reserving a turn does not authorize arbitrary later tool names, parameters or destinations. The original caller in the internal request is an authenticated Gateway assertion checked against that record, not a forwarded public identity header.
Required request/version/binding fields cannot be omitted or defaulted. A
duplicate begin reports existing state, not a fresh send acknowledgement.
A conflicting body or owner/generation fails. Lost responses recover through
the ledger; neither generic HTTP retries nor a status result authorizes replay.
complete distinguishes known terminal outcome, UNCERTAIN and owner-proven
NOT_INITIATED; business-result disclosure always needs current authorization.
For JSON request digests, reuse workflow-invocation-contract’s strict
rfc8785-safe-json-v1 canonicalizer and sha256: lowercase-hex format. Keep
requestDigest separate from user-token bytes and from the logical input digest:
bind the pinned operation/target, actual method/path/query, effect-bearing
headers and payload. Credential headers are verified separately. Preserve
query/array order where meaningful. Binary payloads use a byte digest within
the canonical descriptor. A2 must verify the final request against this binding
after tool-to-HTTP mapping; changed targets, aliases or bodies cannot reuse it.
Unsupported/non-materializable request forms fail qualification rather than
omitting part of the operation from the digest.
Uniqueness is scoped to tenant/action/canonical attempt, with only one live dispatch generation per action. Workflow locks current grant/run/action/budget state in its operational PostgreSQL transaction. No lagging replica or positive Gateway cache supplies approval. Lease expiry never deletes unresolved evidence.
| Recorded state | Permitted recovery |
|---|---|
AUTHORIZED | Fence the old owner and seek fresh authorization for the same canonical attempt if no send intent exists |
SEND_INTENT | Only the original first-transition acknowledgement permits its single guarded send; a lost acknowledgement or restart must not resend |
NOT_INITIATED | Live original owner proves its guard aborted before start; release once, then reauthorize the same attempt under a new generation and current limits |
UNCERTAIN | Reconcile through qualified target evidence/idempotency or an operator; no automatic effect replay |
| Known terminal | Reuse the authorized recorded result/evidence, not the operation |
The five-second timer starts before authorize and includes both control RPCs,
connection/TLS/protocol preparation and capacity waits. Only at the final
transport write boundary may the atomic deadline/acknowledgement check change
NOT_STARTED to STARTED. Expiry before it can close as
ABORTED_NOT_INITIATED; failure after it cannot claim non-initiation. Follow
the accepted design’s owner/boot/fence checks, duration shortening and uncertain
recovery rules. No automatic request retry or redirect can introduce a second
send. Safe NOT_INITIATED retry retains incurred costs and retry counters.
Preserve existing ExecutionClass values interactive, standard, batch,
the checked u16 depth ceiling and separate nonblocking synchronous pools per
depth. StartInvocationRequest needs verified parent action/decision binding
before replacing delegation-derived depth. Root classification is independent
admission; omission cannot reset a child to depth zero. Pin private-version
targets and their logical authorization name across alias changes.
Migration And Transport Inventory
Paths below are source evidence relative to light-fabric unless prefixed
with another repository. Removal is gated on replacement behavior, not on a
text search for the word delegation becoming empty.
| Current path | A1–A3 disposition |
|---|---|
apps/light-workflow/src/executor.rs: DelegationSigner::mint for nested calls | Replace copied user-claim lad1 issuance with broker user credentials, Workflow app identity and online action decisions |
apps/light-gateway/src/main.rs: authenticate_agent_delegation, workflow verifier and replay-store adapter | Replace with dual-token/peer verification and Workflow dispatch ledger; drain old effects before deleting replay authority |
frameworks/light-pingora/src/mcp.rs: delegation context, private targets, depth/class and permits | Carry verified action context; preserve pinned target and nested-call protections |
apps/light-agent/src/main.rs: knowledge_authorization and its upload/retrieve callers | Replace Agent-issued Knowledge tokens and direct route with the approved Gateway path; retain normalized subject and disclosure constraints |
apps/light-knowledge/src/lib.rs: authenticated_context | Replace DelegationVerifier at retrieve, upload, MCP, document-version and passage routes; do not leave its legacy-acceptance window as a fallback |
crates/agent-delegation and its Cargo dependents | Remove only after both Workflow/Gateway and Agent/Knowledge consumers migrate |
crates/workflow-invocation-contract/src/lib.rs: WorkflowDelegationClaims | Migrate required depth, allowed-tool, tenant, deadline and budget checks; retain shared invocation/canonicalization contracts |
crates/agent-store/src/lib.rs: replay-table inventory; operational SQL/schema bundles | Reconcile ownership/removal of agent_delegation_replay_t; keep new Workflow-owned dispatch evidence |
crates/llm-gateway/src/authorization.rs, crates/agent-runtime-protocol/src/gateway_delegation.rs | Preserve dual-token gatewayDelegation policy and its explicit lad1 rejection; the name does not mean it mints legacy tokens |
apps/light-workflow/src/invocation.rs and rule_api.rs | Replace ordinary persisted bearer renewal with grant references; load_status token replacement is not unattended renewal |
workflow.invocation.ignoreUserJwtExpiry | Must be false in the selected qualification profile, including dev; its current dev-only compatibility exception cannot qualify this design |
| Both distributions’ Compose/startup/publication settings | Publish isolated identities, mTLS/trust, broker and dispatch settings; remove old signer secrets only after remaining consumers are drained |
The selected dependency baseline in Cargo.lock is Pingora/core 0.8.1,
local patches/pingora-proxy 0.8.1, reqwest 0.12.28, and hyper-util
0.1.20. Hyper 0.14.32 and 1.9.0 both occur; trace the actual outbound
path instead of assuming either one owns every write. Dependency changes require
repeating transport qualification with the replacement lockfile.
| Outbound path | Concrete A2 qualification target |
|---|---|
| Pingora HTTP proxy | patches/pingora-proxy/src/proxy_trait.rs::error_while_proxy enables reused-connection retry; the loop in src/lib.rs consumes retry decisions. Disable these retries for protected dispatch |
| Pingora HTTP/1 and HTTP/2 | proxy_h1.rs and proxy_h2.rs call write_request_header; locate the actual first transport write below these calls. A handler hook or merely entering an async write is not proof that all queues are past |
| MCP HTTP/backend paths | frameworks/light-pingora/src/mcp.rs uses public/private reqwest clients and ToolRetryPolicy. Disable tool-level status/timeout/connect retry as well as lower-client retry for protected requests |
| Workflow-backed MCP control and child start | Bind the child start to the verified parent decision and retain existing idempotency. Do not confuse repeatable control/status operations with permission to resend a business effect |
| Model and Knowledge calls used by a workflow | Include their real client/proxy paths in the same inventory; no workflow-origin bypass through the interactive profile |
A2 must exercise new and reused connections, first/partial write failures, delayed TLS/connection/capacity, replayable bodies, no hidden retries, and monotonic deadline races on each qualified path. Count initiations as well as target receipts. Until the required write boundary is controlled, that path is not enabled for this profile.
Deployment Manifest And Phase Handoff
A1/A2 must materialize one versioned, non-secret manifest for the selected stack. Freeze its required content now:
- Issuer/profile/provider/tenant, allowed JWT algorithms,
token_useenforcement and any explicit legacy long-lived app key-purpose mapping, broker registration/allowed grants, exact callback and issuer endpoints, certificate registration/trust/revocation settings and opaque secret references. - Token lifetime, validator leeway, qualified clock-error bound and resulting maximum old-claim interval; current claim sources and custom-claim ownership.
- Exact Gateway, Workflow, interactive Agent and workflow-Agent client/service IDs, certificate peers, environment and allowed-caller/route mappings. Client names and Compose service names are inventory hints, not authentication.
- Exact approved forwarding destinations, pinned tool/endpoint/contract versions, resource/disclosure ceilings, supported transport paths and retry settings.
- Source revisions, built image digests, published configuration versions, migration versions, gate results and retained evidence references.
The current local startup defaults include
com.networknt.portal.gateway-1.0.0, com.networknt.workflow-1.0.0,
com.networknt.agent.codex-personal-1.0.0 and
com.networknt.agent.claude-personal-1.0.0. They do not yet establish the new
workflow-only registrations. Gateway’s startup environment defaults to loc
while Workflow/Agents default to dev; reconcile and validate the effective
published peer mappings during provisioning, not by weakening equality checks.
Neither existing Workflow client name proves it is the new broker.
Each isolated workflow Agent also needs its own service ID in
codingProfile.workspaceBindings[].agents and runner-local
RunnerWorkspaceConfig.bindings[].agents, with the complete bindings equal.
Keep intentional dev credential fixtures; official profiles must reject their
trust and use independent credentials.
| Handoff | Required deliverable |
|---|---|
| A0 to A1 | This v1 baseline, selected issuer/grant restrictions, claim ownership, current client/source inventory and the required deployment manifest fields |
| A1 | Issuer/broker implementation, issuer-owned token_use emission/reservation and shared verifier checks, enrollment including PKCE and consent/provenance, mTLS enforcement, live claims, strict rotation and failure evidence |
| A2 | Manifest-bound caller/destination policy, isolated Agents, action/dispatch stores and APIs, nested protections, and qualified transport behavior |
| A3 | Selected-stack end-to-end gates, old-authority migration/drain, compatible dev startup and independent official trust where applicable |
| Personal orchestration Phase 1 | Only after the selected A1–A3 qualification record passes; A0 by itself does not admit development workflows |
A1’s purpose matrix must include a client_credentials token with registered
uid/role claims: it remains token_use=app and is rejected in user
Authorization. Test custom/request attempts to override purpose, every
issuance/refresh path, missing/unknown markers, user tokens in the app position,
and the narrowly scoped legacy long-lived app exception. A2 repeats receiver
integration tests. Inspect the actual signed JWT payload to prove exactly one
issuer-selected token_use survives filtering and serialization, including
custom-claim override attempts. A1 also records refresh attempts, uncertain
rotations and
reauthorizations during a multi-hour soak, separating ordinary-load observations
from injected response loss; see the parent design’s A1 exit gate.
Changes to contract meaning, issuer grant eligibility or recovery transitions require a new baseline revision and affected gate review. Reassigning environment bindings also requires requalification of those bindings; it cannot silently change origin classification, token purpose or lifecycle transitions.
A1 Implementation And Qualification Status
Historical record. This report records the former A1 credential-broker implementation and its 2026 qualification evidence. Workflow Invoke later retired the broker, including both browser callback and backend enrollment routes. Current invocation sends the acting user’s bearer through Gateway to Workflow. Asynchronous
workflow_startruns that can outlive that bearer may use a LONG registration exchanged through light-oauth; synchronousworkflow_invokedoes not register LONG. See Workflow Invoke. The historical commands, deployment steps, and acceptance plans below are not instructions for the current stack. The linked activation receipt retains its original observations, including the retired callback, as audit evidence.
Historical status: A1 source implementation complete; scope-consent reuse correction qualified in the short integration suite and deployed locally before broker retirement.
Historical acceptance amendment (2026-09-14)
The user requires only the existing Portal portal.r / portal.w scope consent.
There must be no additional workflow-specific consent screen or second login.
Workflow identity, policy binding, credential ceilings, expiry and revocation
remain backend enforcement requirements; existing scope consent is not permission
to skip issuer provenance validation.
The user waived scheduled and hours-long renewal qualification after the original user token expires. Record these tests as waived / not run, not passed; they are no longer A1 acceptance blockers. Historical requirements below describe the previous baseline and are superseded by this amendment.
The extra Portal enrollment prompt and generated Gateway workflow_authorize
tool remained removed. At that time, root HTTPS Workflow invocation without a
supplied grant acquired one internally through /workflow/credentials/enroll. Workflow used
the issuer’s mTLS /oauth2/{provider}/workflow/enrollments/acquire endpoint,
then redeemed the returned one-time PKCE code server-to-server. Only a grant UUID
was returned to Gateway; no redirect, password or second consent was involved.
The issuer records SHA-256 access-token fingerprints in ordinary authorization-code and refresh issuance audits. Acquisition requires exact issuance evidence tied to an active authorization-code session, matching user/Host/client/provider and scope ceilings. Tokens issued before this change need a normal Portal refresh to obtain this evidence; unsigned claims, app tokens and unrecorded user tokens cannot bootstrap renewal. Acquired grants retain the source-session relationship; revocation, scope removal or missing provenance blocks renewal. Legacy browser grants require their original login evidence and remain only for compatibility.
Short real-HTTPS/mTLS integration passed: initial acquisition, acquisition after normal Portal refresh, direct broker renewal, scope-expansion rejection, wrong Host rejection, unrecorded-token rejection, missing-provenance rejection and source-session revocation. OAuth unit tests: 22 passed, 2 ignored. Workflow library: 86 passed. MCP module: 175 passed, 3 ignored. Scheduled/hours-long tests were waived, not passed. No live personal/native workflow was launched for this change.
At that checkpoint, the updated OAuth, Workflow and Gateway binaries were running
locally with zero restarts; the unauthenticated enrollment probe returned 403.
This was container-layer qualification, not a rebuilt release image, and would
have been lost on container recreation. Rollback binaries were recorded under
/tmp/phase1-backend-acquisition.4FBLYy/; their current presence is not assumed. Snapshot
01a0a222-a6dc-7288-9845-f709ead64d8e was current at the time. The editable
instance property then contained an enrollment ACL assignment; its present
state requires a separate check. Application databases were not wiped. A separate
oauth_a1_qualification database held schema-only fixtures and test evidence.
The implementation follows the frozen A0 baseline. It does not admit personal orchestration Phase 1: production receiver enforcement and dispatch are A2/A3, and live deployment qualification remains separate.
Historical implementation
- Issuer: dedicated, reserved
JwtClaims.token_use; live tenant-bound refresh authority; explicit custom-claim sources; PKCE and browser consent; dedicated certificate-authenticated broker listener; strict rotation, consumed-token history, grant lookup and revocation. Broker clients cannot use secret-only authentication through another provider binding. Specialized non-broker grants retain their supported behavior. - Recovery: revocation tombstones serialize with enrollment and redemption, including when revocation arrives before the grant exists. Uncertain refreshes require reauthorization; the broker never retrieves a replacement through interactive retry grace.
- Workflow: encrypted credential storage, a store-to-issuer/client binding, durable enrollment and renewal ownership, run binding, key rotation, restart recovery, cancellation/revocation fencing, and retryable issuer revocation. At the time, enrollment APIs returned references and an authorization URL, never refresh credentials. The browser callback used a separate optional TLS listener with stored OAuth state and backend PKCE. Both enrollment routes and the callback listener have since been retired.
- Shared client/verifiers: fixed HTTPS endpoints, explicit CA trust, mTLS, bounded responses, signed-token validation and disabled redirects/retries; user/app purpose validation including duplicate-marker rejection and the explicit app-only legacy-key exception. Receiver-wide wiring remains A2.
- Deployment: the canonical Portal patch and fresh schema, regenerated bootstrap SQL for both distributions, a separate credential database/runtime principal, private mount preparation and ownership, registration/replay checks, certificate/key rotation procedures, opt-in Compose overlays, and a null-default Portal configuration catalog delta. The OAuth image build now includes the complete external Cargo path-dependency manifests needed by A1.
The former issuer schema came from portal-db/postgres/patch_20260913_01_workflow_broker.sql
and was once reflected in canonical DDL and distribution bootstrap SQL. The
later portal-db/postgres/migrations/patch_20260928_04_retire_workflow_broker.sql dropped
the broker tables. The old patch and bootstrap state are historical evidence,
not installation instructions; applying the old patch would recreate retired tables.
The former Workflow credential schema was
light-fabric/apps/light-workflow/migrations/credential_broker.sql.
Provisioning assets formerly lived in light-fabric/deployment/workflow-broker
and the matching local and installer overlays. These source assets were retired;
the paths here identify historical qualification inputs.
Historical qualification
Machine-readable results and local image IDs and working-tree source hashes identify the earlier qualified artifacts. The image IDs are local builds, not published registry receipts, and predate the availability fixes below. That historical issuer image was not a release qualification and is no longer a deployment target.
| Check | Result |
|---|---|
| OAuth unit suite | 21 passed; its two database tests were also explicitly run |
| Workflow library suite | 80 passed |
| Shared security and purpose suites | 17 and 2 passed |
| Existing workload grant compatibility | Passed against isolated PostgreSQL |
| Live authority | Passed: tenant pinning, removed membership and locked users |
| Complete consent flow | Passed over real HTTPS: consent form, PKCE redemption and production callback listener |
| Purpose and ceilings | App token with user claims rejected; scope/duration expansion rejected; purpose override filtering verified |
| Interactive refresh grace | Passed: removed role/group/position and changed attribute appear in the next token; account/session/membership revocation and failed live queries reject renewal |
| Strict issuer concurrency | Exactly one committed rotation; authenticated duplicate revokes the family; wrong-client reuse does not revoke it |
| Broker concurrency/recovery | Passed: concurrent renewals and callbacks, process replacement, key-file rotation and store/client mismatch rejection |
| Certificate lifecycle | Actual mTLS renewal/retirement, missing/wrong peer and secret-only rejection passed |
| Failure injection | Lost committed HTTP response, issuer outage, revocation retry and post-rotation persistence failure passed without refresh replay |
| In-flight fencing | Response delayed after issuer commit is rejected after run cancellation or recovery fencing |
| Enrollment revocation | Revocation before creation and before delayed code redemption prevents resurrection |
| Canonical schema | PostgreSQL 17.10 fresh/upgrade/replay, cascade policy and deterministic regeneration gates passed |
| Distribution bootstrap | Exact regenerated installer SQL loaded successfully; both distributions use identical canonical bootstrap SQL |
| Private provisioning | Credential database installation/replay and runtime role separation passed; registration replay cannot revive a retired certificate |
| Mounts | Both Compose overlays validated; cached non-root images can read only their own prepared test mounts |
| Two-hour renewal soak | Stopped at user request; manual qualification pending. Final-source run logged six successful rotations through 3,301 seconds; this is not a completed two-hour gate |
| Qualification images | Both separate A1 tags built; OAuth startup/mTLS smoke passed on the disposable schema; no latest replacement or live restart |
The HTTP tests use the real Workflow broker and local issuer together, with PostgreSQL and actual TLS handshakes. Faults are injected around transport delivery or database persistence, rather than replacing renewal with a mock success. These are isolated integration tests, not a claim that the live Portal frontend and deployed configuration have been qualified.
The historical issuer/Workflow integration gates used
portal-service/apps/light-oauth/scripts/run-a1-gates.sh. That script was
removed with the broker; the deferred two-hour soak was never recorded as passed.
The historical gate used the disposable oauth_a1_qualification database with a
schema-only configserver fixture. Credentials were supplied privately. The
script recorded base revisions and working-tree source hashes and rejected source
drift during qualification. The A1 source was later committed and retired; the
recorded base revision alone was not qualification evidence. Remote GitHub CI
had not been run at the time.
Availability Review Follow-up
Both portal-service availability findings are fixed:
- The broker listener starts TLS handshakes and certificate extraction in separate tasks, capped at 128 pending handshakes with a five-second timeout per task. Silent peers do not serialize acceptance. A real mTLS regression keeps four earlier TCP connections idle while an authenticated request completes.
- Broker redemption and renewal use their existing transaction connection for session authority, live claims, custom-claim sources and signing-key reads. They never acquire another pool connection while holding rotation locks. Twelve concurrent redemptions, then twelve concurrent renewals, complete with a one-connection pool alongside ordinary issuance. Strict reuse detection is preserved.
The complete short gate suite passed again on a fresh isolated PostgreSQL database: 21 OAuth unit tests, explicit live-authority and workload-compatibility tests, and real HTTPS/mTLS broker integration including the new contention regression. Source hashes for this follow-up identify this tested revision. The earlier image receipts and partial soak describe older source; no images or soak were rebuilt/rerun for this follow-up. The deployed scheduled and hours-long runs remain pending, and the A1 exit gate has not passed. The subsequent Workflow/client review below refines pre-request error classification.
Workflow And Client Review Follow-up
All five light-fabric findings are addressed:
- Periodic recovery logs store failures and retries; it no longer terminates the managed task and closes unrelated Workflow admission. Startup configuration validation remains strict.
- Connection refusal, TLS establishment failure and pre-request JWKS failures
return
NotSent. The same live owner atomically recordsNOT_SENTand restoresACTIVEwithout changing the token or generation. Later caller attempts may retry. Post-send failures and lost ownership still fence the grant. - Verification keys are cached for five minutes before rotation, with one refetch for an unknown key ID. A post-rotation verification failure remains uncertain; issuers must publish new keys before using them.
- Canceling one run preserves a valid committed shared-grant rotation and denies only that run’s token. A sibling run can renew; grant revocation and owner fencing still reject late responses.
- At that time,
credentialBroker.legacyLongLivedAppKeysconfigured approved local issuer/key pairs for markerless app fixtures. It defaulted empty, applied only toX-Scope-Token, and cannot override explicit invalid purpose markers or authenticate a user. Former distribution preparation scripts carried the setting.
Short gates passed on a disposable PostgreSQL database with actual HTTPS/mTLS:
21 OAuth unit tests, explicit live-authority/workload tests, full broker integration,
80 Workflow library tests, 20 light-client tests, 17 security tests, two purpose
contract tests, and two preparation tests. The integration exercises refused and
failed-TLS connections without token POSTs, JWKS outage/rollover, post-rotation
verification failure, recovery store failure, sibling-run cancellation, and the
actual API legacy-key allowlist. Source hashes
identify this follow-up. No soak or deployed-stack acceptance was run. Both
previous image receipts predate these fixes and require rebuilding/requalification.
At that stage, the credential-store SQL (the NOT_SENT result constraint) and
issuer database patch were prerequisites for the then-current images. Both broker
paths were retired later. No live database or service was changed in this follow-up.
Local Database Migration Applied
At this historical migration checkpoint, the local all-in-lt PostgreSQL instance had the issuer patch in
configserver.configserver and the credential migration in the newly provisioned
workflow_credentials.workflow_secret. All seven issuer tables, five credential
tables, the NOT_SENT constraint, runtime password authentication and restricted
role privileges were verified. Cascade validation passes after repairing three
stale A2A schema references and adding four missing canonical gateway policies.
Migration receipt records hashes
and checks without secrets. The runtime URL is in the ignored mode-0600 file
portal-config-loc/all-in-lt/postgres-db/secrets/workflow-broker-database-url.
That file was an input to the retired broker preparation, not a current setup input.
This supersedes earlier statements that no live database was changed. No images were rebuilt and no application services were restarted. Broker registration, certificates, configuration activation and selected-stack qualification remain pending; the broker profile is not enabled by this migration alone.
Historical local broker activation
The user rebuilt the issuer and Workflow images. At the time, the local broker was enabled with a dedicated private client CA and a registered 180-day client certificate. The existing local issuer HTTPS certificate is used for its internal server and localhost callback. Private mounts remain ignored by Git, directories 0700 and files 0600, owned by the measured service UID/GID 999:999. No user grant was created.
The catalog and instance configuration were imported through event-importer using
the registered ConfigInstanceCreatedEvent with commandkind: MUTATION for the
instance property. Workflow snapshot 04c798da-c95b-499f-9167-a86bd248a145 is active.
The broker JWKS URL uses light-oauth:6881 inside Docker; its authorization URL
uses localhost:6881 for the browser. The local legacy app exception is restricted
to the verified LC signing key. OAuth and Workflow were recreated and are healthy
with zero restarts. Registered mTLS reaches grant lookup; no certificate fails TLS,
wrong client and secret-only broker authentication return 401. JWKS returns 200.
The callback returns 400 without state/code over verified HTTPS.
Activation receipt records image
IDs, snapshot IDs and certificate fingerprint from that historical activation.
At the time, deploy-local.sh lt included a broker overlay when the ignored
workflow-broker/.runtime/enabled marker existed; the overlay was later removed.
At the time, the host system trust store did not trust the local issuer CA;
explicit CA-file verification passed. The browser consent path was later retired.
This activation supersedes the historical no-import/no-restart statements above. The live preflight returned nine unclassified names on four clients: Support Triage Local Demo, Tech Support LLM Workload Dev, mcp379-local-qualification and pylon. No claim-source classifications were changed. The broker has no custom claims, so these do not block its registration or activation.
Superseded selected-stack acceptance plan
The former selected-stack plan called for reviewing custom-claim sources, enrolling a user through Gateway and a browser authorization flow, and running scheduled and hours-long renewal qualification. Those steps were not completed as an A1 exit gate. Workflow Invoke subsequently retired that enrollment flow, so this plan is no longer a deployment or test procedure.
The A1 implementation described here was later committed and retired.
Authorization A2 Implementation Progress
Historical progress snapshot. The broker grant, enrollment, and renewal paths described below predate Workflow Invoke and were later retired. See Workflow Invoke for current Gateway-to-Workflow invocation and LONG registration. The old gates below are not instructions for the current stack.
Historical status: the A2 source foundation and selected personal coding paths were implemented, and all selected services, including both workflow-only Agents, passed local runtime qualification; selected-stack A2 acceptance had not passed. At that checkpoint, the remaining authorize/begin/send/complete, renewal, revocation, receiver and recovery matrix was still required. This record did not admit personal orchestration Phase 1.
Historical implemented foundation
workflow-actiondefines server-owned action bindings, canonical attempts, owner/boot/fencing identities, exact-byte request digests and child-lineage validation. Opaque references resolve stored permits; callers cannot supply replacement depth, class, grant or budget authority.- The PostgreSQL ledger implements authorize, begin-dispatch, completion and
status. Concurrent duplicates replay the same authorization; only a new
SEND_INTENTacknowledgement permits initiation. Owner-boundNOT_INITIATEDcompletion releases once and permits a new authorization generation. Historical completion receipts cannot release a newer reservation.UNCERTAINcannot become retryable by lease expiry. - Migration
0007_workflow_action_dispatchadds authority, permit, dispatch and audit tables. It is present in canonical operational-store schema bundle2.2.0, with matching manifest/order/checksums. That schema-bundle version is independent of the container image release tag; the selected deployment uses image tag2.3.5-dev.20260909.2338. The local operational database has the schema and its canonical migration-ledger receipt. light-security::dual_identitychecks separate user/app purposes, explicit issuer/audience/tenant, route-approved app service IDs and verified TLS peer fingerprints. Duplicate/coalesced credential headers are rejected. Workflow and receiver origins require an action reference. Missing verified peer context is rejected before any issuer/JWKS work. This helper does not replace route policy authorization.light-axum::mtlsprovides a certificate-verifying listener with bounded, concurrent handshakes and peer fingerprints from TLS, not forwarded headers.- Workflow has a dedicated, optional mTLS action API and a fixed HTTPS control
client in
light-client. Authorize/begin/status verify user claims and hold broker grant/run read locks while checking the ledger. Completion authenticates the original service/peer without requiring an unexpired user token. workflow.actionAuthorizationdefaults to null. Enabling the listener requires the action migration, broker and strict JWT expiry verification. This switch alone does not establish a working A2 deployment.- Pingora core 0.8.1 is vendored with an optional socket-level guard below TLS and buffering. Connection/TLS preparation precedes the guard; deadline check and first socket poll happen synchronously. Cancellation/expiry fences cleanup writes. Guarded HTTP/2 and unknown/custom transports fail closed.
- The proxy patch passes an optional guard to HTTP/1 and disables its default
reused-connection retry for guarded requests. The new
guarded_httpadapter makes one connection attempt and one guarded send, with bounded response size and timeout, no redirects and no retry loop.
Verification performed
These are source/foundation checks, not deployed-stack acceptance:
| Check | Result |
|---|---|
workflow-action contract tests | 5 passed |
| Disposable PostgreSQL action-ledger test | 1 passed, executed against PostgreSQL |
light-security unit tests | 19 passed |
| Real mTLS listener test | 1 passed |
light-workflow --lib | 81 passed |
light-agent --lib | 21 passed, 4 database tests ignored by their existing environment gates |
light-agent --bin light-agent | 44 passed |
light-knowledge --lib | 8 passed |
| Pingora core socket-guard tests | 5 passed |
light-pingora --lib | 448 passed, 5 existing tests ignored |
| Guarded HTTP adapter tests, included above | 3 passed |
| Gateway action response classification | 2 passed |
light-gateway --lib | 4 passed |
light-gateway --bin light-gateway | 79 passed, 3 existing tests ignored |
| Combined Workflow/Agent/Knowledge/Gateway compile | Passed |
Selected-package cargo clippy --all-targets | Passed; existing repository warnings remain |
| Operational bundle checksum validation | Passed |
git diff --check | Passed |
cargo fmt --all -- --check | Passed |
mdbook build docs | Passed; existing large search-index warning remains |
The PostgreSQL test uses its own disposable database, not the live operational or credential stores. It covers simultaneous duplicate authorization/begin, conflicting owners and changed bindings, opaque-reference mismatches, non-initiation replay and reservation release, stale decisions, cancellation in the foundation authority table, uncertain outcomes, child-run lineage, and idempotent qualified-receiver reconciliation with conflicting evidence rejected. The second pass extends this test to actual invocation cancellation and the existing budget ledger, as described below.
Transport tests cover real plain connection reuse, TLS buffering after expiry, pending writes, cancellation cleanup, and partial-send failure of a replayable body without reconnection in the adapter. Full Gateway proxy-path first-write and partial-write failure tests, with its actual retry configuration, remain pending.
Reproducible bounded checks are in
scripts/run-authorization-a2-foundation-gates.sh. The PostgreSQL test requires an
explicit empty disposable database URL; absence is reported as not run. No A1
soak is started by this script.
Integration added in the second implementation pass
This pass adds real call-site wiring, but does not complete A2.
- Invocation admission accepts
renewableGrantIdonly in the explicitly enabled A2 profile. The broker checks the exact consent binding, and the operational transaction installs run authority alongside invocation acceptance. Repeated broker run binding requires the same grant, user, tenant, work binding and expiry. Consent binding is{profile: "workflow-action-v1", workflowDefinitionId, definitionDigest, policyDigest, responsePolicyDigest}. - The action ledger locks the actual
workflow_invocation_tandworkflow_invocation_budget_trecords. Actual cancellation, deadline, policy, subject and budget-generation checks now apply. Action reservations charge the existing nested-call, byte and cost counters. Non-initiation refunds once; uncertainty retains the reservation. Initiated effects conservatively consume their configured byte/cost bounds until qualified actual-cost receipts exist. bound_mcp::Runtimeis installed on the Workflow executor when A2 is enabled. It resolves the published dependency, renews through the broker, creates the action permit, and sends the original user credential plus the Workflow app credential over mTLS. Model-supplied authorization headers are not used on this path. The current producer supports root-run MCP tool actions only.- Workflow’s dedicated mTLS action listener also serves protected invocation routes. The ordinary listener rejects A2 invocation requests without verified peer context. Retrieval/cancellation/result routes receive the same strict user/app check. Existing route-level invocation authorization still applies.
- Gateway loads the restart-required
gateway.workflowActionsprofile fromworkflow-actions.yml. Its TLS listener requests certificates from the configured client CA; ordinary browser connections may omit a certificate, but strict MCP routes require an approved verified peer. Origin is classified from verified credentials.X-Workflow-Grantselects a grant for root invocation; it is only a reference, and Workflow verifies the actual consent binding. - The MCP request context now carries the verified A2 caller. Private targets inspect the stored action and require matching stable tool and contract metadata. Existing MCP policy and argument-mapping checks are retained. HTTP-tool and stateless MCP backend branches call the fixed action client and guarded transport. The destination registry checks the exact outbound URL. Platform credentials are stripped for destinations not configured for original user forwarding.
- Gateway boot registration is durable. The client generates a fresh boot ID; Workflow assigns a fencing generation. Repeated registrations replay, while a superseded boot cannot regain ownership. Authorize, begin and completion lock the current owner within their dispatch transaction. A new replica/boot can query latest state without acquiring send permission for an uncertain effect.
- Gateway does not treat HTTP
202 Accepted, informational responses or redirects as terminal execution evidence. These responses retain an uncertain reservation and do not complete the Workflow step. A focused regression test covers these classifications. Target-receipt reconciliation remains pending. - The socket guard now checks the deadline throughout request writes, including after a successful TLS control-record write. This conservative implementation also bounds request-body writing to the lease; response reads may take longer.
The PostgreSQL test now installs the real Workflow schema and constraints, seeds an actual invocation/budget, and tests real cancellation, reservations and boot replacement. It passed. Workflow’s 81 tests, invocation-contract’s 4 tests, and proxy-framework’s 448 tests passed (5 existing proxy tests ignored). The core socket guard now has 5 passing tests. Combined Workflow/Gateway compilation passed. These checks do not constitute an end-to-end A2 deployment test.
Integration added in the third implementation pass
- Nested Workflow starts now carry
parentActionId. Gateway derives depth, execution class and deadline from the verified parent binding, sends the child start through the guarded action transport, and includes the parent in its idempotency digest. Workflow proves the current Gateway owner andSEND_INTENT, verifies the child invocation fields, records parent action/run lineage, and inherits the exact renewable grant without copying refresh material. - Workflow service Agent jobs are limited to bound coding/workspace inputs in the A2 profile. Separate interactive and workflow Agent instances use immutable service/definition/origin configuration, distinct credentials and verified mTLS listeners. Interactive instances do not consume the workflow job queue; workflow instances reject Chat, A2A, UI and upload ingress. Workflow checks the job’s live invocation, grant, cancellation, deadline, depth and exact published Agent definition both before turn creation and before coding dispatch.
- Knowledge has an optional restart-required receiver profile and dedicated mTLS
listener. Gateway forwards the current user token, its own scope token and the
opaque action reference only to approved internal targets. Knowledge verifies
Gateway user/app/peer identity, then uses its separate receiver identity to ask
Workflow for the live action binding. Exact tool/contract/policy/disclosure
mappings are checked before existing Knowledge user/resource ACLs run. When the
profile is enabled, failure cannot fall back to the old
lad1delegation. - The frozen
completecontract also accepts a qualified receiver receipt for an alreadyUNCERTAINgeneration. Workflow authenticates the receiver app and exact mTLS peer, checks its registered tool set, and accepts only terminal evidence for the same generation. Identical receipts replay; conflicting evidence fails; the retained reservation is settled once. Receivers cannot authorize or begin a dispatch through this path. - Gateway terminal-response classification is target-specific. Informational,
redirect and generic
202responses are not completion evidence. A Workflow child202qualifies only when its bounded response contains the exact stable tool, a non-nil instance and a positive durable state version.
Selected-stack runtime check on 2026-09-14
After rebuilding the selected local images and restoring the disposable
operational schemas, ./scripts/deploy-local.sh lt start exited successfully and
reported that the required Compose services passed runtime qualification.
PostgreSQL, Config Server, Workflow, Gateway, Knowledge Admin, and the ordinary
interactive Agents were running; Gateway loaded its authorization policy and
registered with Controller.
On a subsequent 2026-09-14 run, both workflow-only snapshots were published,
their app credentials gained the required execution.invoke scope, and both
services registered with Controller without execution-result authorization
errors. The deployment qualification contract now includes both services when
the A2 profile is active. This is still not the A2 exit gate: no complete
authorize/begin/send/complete, renewal, revocation, receiver receipt, or recovery
matrix was completed in that run.
Required work before A2 acceptance
- Exercise the two workflow-only Agents through real Workflow job admission, coding dispatch, cancellation and result reconciliation. Startup and Controller registration alone do not qualify those paths.
- Connect any asynchronous effect receiver that needs automatic recovery to the
qualified
completereceipt contract. Synchronous Knowledge responses and durable Workflow acceptance have immediate receipt policies; other unknown effects correctly remain non-retryable until their receiver supplies evidence. - Add and run the complete A2 integration/security/race matrix against the real receiving and sending paths. Existing regression suites predominantly exercise the default profile; they are not evidence that the new enabled profile works end to end. GitNexus reports critical aggregate change risk because Gateway and Workflow entry points, MCP dispatch, and Agent/Knowledge admission are all in scope. Do not remove legacy delegation before its replacement qualifies.
Deployment state
The qualified-receiver completion contract is frozen. The local database contains
the A2 action migration and receipt. Portal contains separate codex-wf and
claude-wf definitions with service IDs
com.networknt.agent.codex-personal-workflow-1.0.0 and
com.networknt.agent.claude-personal-workflow-1.0.0. Portal and runner-local
workspace bindings use authorization revision 3 and grant both interactive and
workflow-only identities. The local distribution has an enabled, ignored A2
runtime package containing marked app credentials, mTLS identities and exact peer
mappings; the installer contains the matching non-secret provisioning assets.
The rebuilt selected services pass local runtime qualification, including the
workflow-only identities. Their full A2 action and job paths remain unqualified.
No commit or push was performed.
The 2026-09-13 local inventory confirms that the published Codex and Claude
coding profiles and the owner workspace grant the existing
com.networknt.agent.codex-personal-1.0.0 and
com.networknt.agent.claude-personal-1.0.0 identities. Those running definitions
already have pre-A2 history and cannot be relabeled as workflow-only instances.
Those identities remain interactive. The separately published workflow-only
definitions preserve their history boundary; configuration aliases or a
caller-supplied origin header cannot cross it.
Authorization A3 Migration and Removal Progress
Historical progress snapshot. This record predates Workflow Invoke and the retirement of the A1 credential broker. Its grant-backed enrollment and renewal gates are superseded; see Workflow Invoke for the current invocation and LONG credential flow.
Historical status: A3 had started; removal was not yet safe. The selected local A2 stack used separate workflow-only Agent identities, and grant-backed action invocations no longer persisted the caller’s reusable access token in ordinary Workflow invocation state. The complete A2 integration matrix and A3 consumer drain remained open at that checkpoint.
First migration slice
- Workflow action admission still authenticates the current user and immediate caller, binds the run to the approved renewable grant, and records the action authority and dispatch ledger.
- For the dedicated grant-backed action listener,
user_authorizationanduser_authorization_expare inserted as null. The broker remains the source of fresh credentials for bound MCP actions and workflow Agent jobs. - A later authenticated status read cannot repopulate a deliberately null token. Ordinary interactive invocations retain the existing token behavior until a qualified non-durable credential handoff replaces it.
- Terminal cleanup remains idempotent. Existing active rows have not been rewritten: rows with valid grant authority must be migrated, while legacy rows without it must drain or be reauthorized.
Selected-stack credential inventory
| Workload | App credential | Peer identity | Current purpose |
|---|---|---|---|
| Gateway | issuer-signed app token | Gateway client certificate | Workflow action control and backend dispatch |
| Workflow | issuer-signed app token | Workflow server/client certificates | Action API and Gateway calls |
| Codex workflow Agent | issuer-signed app token with execution.invoke | Codex client/server certificates | Workflow jobs and Controller execution API |
| Claude workflow Agent | issuer-signed app token with execution.invoke | Claude client/server certificates | Workflow jobs and Controller execution API |
The local preparation tool generates separate app tokens and mTLS material for
these identities. Gateway and Workflow keep portal.r portal.w; only the two
Agent identities receive execution.invoke. Private PKI remains runtime-owned
and is not checked into Git.
Legacy authority still present
The agent-delegation crate still has five direct package consumers:
light-agent, light-workflow, light-gateway, light-knowledge, and
light-pingora. Workflow still has signer configuration; Gateway and Knowledge
still verify legacy delegation; Pingora still derives nested workflow context
from it. The Agent now requires its legacy signer only when a direct Knowledge
endpoint is configured, which removes the dependency from workflow-only Agents
without weakening configured Knowledge access.
agent_delegation_replay_t also remains in the Agent operational schema and its
validation/reset tooling. It cannot be dropped until old attempts are drained
and the workflow_action_dispatch_t evidence/recovery path passes the selected
stack’s first-write, partial-write, uncertainty, restart and receiver-receipt
tests.
Superseded gate plan before removal
The following list records the former A3 plan and is not a current rollout procedure.
- Run the complete A2 authenticated action matrix, including workflow Agent and Knowledge paths, cancellation, unauthorized callers, nested depth, private targets and effect recovery.
- Qualify enrollment, scheduled renewal, rotation, revocation, permission changes and crash recovery against the selected issuer profile.
- Migrate active grant-backed invocation rows to null token storage. Drain or explicitly reauthorize active legacy rows that have no valid grant authority.
- Replace each remaining Gateway/Pingora/Knowledge delegation decision with the live action and receiver contracts, then prove no legacy consumer remains.
- Remove signer/verifier configuration and deployment secrets in one
coordinated rollout. Only then remove
agent-delegationand apply a migration that dropsagent_delegation_replay_twhile retaining the action ledger and unresolved dispatch evidence.
Agent Engine Pattern
The Agent Engine Pattern is the architectural standard for building industrial-grade, metadata-driven AI platforms within the Light-Fabric ecosystem.
In this model, the Rust Runtime acts as a high-performance Orchestrator, while the Application Logic resides in externalized metadata (JSON/YAML) and the Hindsight Memory database.
1. Why the Metadata-Driven Approach?
- Separation of Concerns: Complex platform logic (security, retries, database connectivity, LLM integration) is implemented once in Rust. Business logic—defining agent personas, goals, and steps—is “programmed” via JSON or Database records.
- Hot-Reloading: Using the
arc-swapcrate and YAML-based rule engines, agent personas, model parameters, and tool access can be updated in real-time without a server restart. - Elastic Scalability: Deploy one shared agent engine and specialize it from
registry metadata. The public
light-agentservice owns sessions and reasoning,light-agent-workerhosts sandboxed coding/runtime adapters, and the optionallight-agent-channelowns messaging connections. These are thin trust-boundary executables over shared domain crates, not separate persona engines. - High Performance: Rust’s asynchronous
tokioruntime allows a single engine instance to manage thousands of concurrent agentic sessions with minimal memory overhead.
2. The Core Architecture: Engine vs. Content
To function as a generic interpreter, the Light-Fabric Engine relies on four primary components:
A. The Tool & Skill Registry (The “Hands”)
The engine maps string identifiers in the workflow JSON (e.g., "call": "get_customer_data") to governed API/MCP capabilities, fixed actions, or immutable sandbox packages.
- Implementation: Uses a
ToolRegistrywith trait objects (Box<dyn Tool>) or dynamic dispatch to MCP (Model Context Protocol) servers. - Logic: When the LLM requests a tool call, the engine verifies permissions via Fine-Grained Authorization, executes the tool, and feeds the result back into the context.
The registry is not an authorization or execution boundary. Mutable script
source is never trusted because it appears in metadata; executable packages
must be content-addressed, reviewed, and run through an approved
ExecutionBackend.
B. Hindsight State Manager (The “Memory”)
Unlike simple session storage, the state manager persists every step of the agentic interaction into biomimetic memory banks.
- Implementation: Every “turn” in the conversation is saved as a
unit_tin the Hindsight database. - Benefit: Provides fault tolerance (resuming from a crashed step) and “Recall” capabilities, allowing agents to remember past interactions across different sessions.
C. Prompt Templating (The “Mind”)
System prompts and instructions are stored as templates rather than hardcoded strings.
- Implementation: Uses the
teraorrinjaengines for high-performance string interpolation. - Example:
"You are a {{agent_role}}. Your current objective is to {{agent_goal}}." - Rust Logic: The engine merges runtime context (user input, memory recall, tool results) into the template before calling the LLM.
D. Policy Engine (The “Shield”)
Before any tool execution or data retrieval, the engine consults the Light-Rule middleware.
- Logic: Ensures the agent has the authority to access specific data or execute specific functions, preventing “prompt injection” from leading to unauthorized actions.
3. Conceptual Implementation in Rust
The AgentEngine in Light-Fabric follows a non-blocking, async loop:
#![allow(unused)]
fn main() {
pub struct AgentEngine {
registry: Arc<ToolRegistry>,
memory: Arc<HindsightClient>,
rules: Arc<RuleEngine>,
}
impl AgentEngine {
pub async fn execute_step(&self, session_id: Uuid, task: Task) -> anyhow::Result<()> {
// 1. Fetch current context from Hindsight Memory
let mut context = self.memory.get_context(session_id).await?;
// 2. Resolve Task Type (Agentic vs. Tool Call)
match task {
Task::LlmCall { agent_id, prompt_template } => {
// Render prompt with Tera
let prompt = self.render_prompt(prompt_template, &context)?;
// Call LLM Provider
let response = self.llm_provider.chat(prompt, &context).await?;
// Retain turn in Hindsight
self.memory.retain_turn(session_id, response).await?;
},
Task::ToolCall { tool_name, params } => {
// 3. Enforce Fine-Grained Authorization
if self.rules.authorize(session_id, &tool_name).await? {
let result = self.registry.call(&tool_name, params).await?;
context.add_result(tool_name, result);
}
}
}
// 4. Update Session State
self.memory.checkpoint(session_id, context).await
}
}
}
4. Operational Challenges & Solutions
- Tool Versioning: As the platform evolves, tools may change. Light-Fabric handles this by versioning tool definitions in the Registry, ensuring old workflows remain compatible with the tools they were designed for.
- Safe Execution: A logical agent does not automatically own a sandbox. Remote model and gateway-only work can remain in the long-lived service; shell, filesystem, browser, local MCP, CLI-model, or untrusted execution uses an approved
ExecutionBackendsuch as a microVM, rootless container, Kubernetes Job, dedicated VM, or fixed external action. The effective policy must match the backend’s proven boundary. - Observability: Because the engine is generic, tracing is built into
light-runtime. Traces record session, turn, model-call, tool-action, policy, lease, and result metadata without treating private hidden reasoning as an observable platform contract.
The Recommendation
Light-Fabric adopts this “Engine-first” philosophy to keep one durable agent model across enterprise, coding, and personal-assistant profiles. Agent definitions and skills are data; shared Rust crates implement sessions, policy, memory, runtime protocols, and audit; thin service, sandbox-worker, and channel binaries enforce their distinct lifecycles and trust boundaries.
See Light-Agent Execution for the concrete service, session, turn, tool, runner, and sandbox boundaries.
Light-Agent Execution
Status
Proposed.
This design defines how interactive Light agents are hosted, how agent turns and tool actions are authorized and recovered, and when execution must move from a long-lived agent service into a runner-managed sandbox.
It complements:
- Light-Workflow Runner
- Execution Backends And Sandbox Execution
- Agent Engine Pattern
- Hindsight Memory
Decision
A logical agent is not an isolation unit.
Do not create a container, VM, or sandbox for every agent definition by default. Use a hybrid model:
- run interactive API-based agents in a long-lived light-agent service;
- group service instances by tenant, trust, model-provider, network, and data boundary;
- execute local CLI providers, code, shell, browser, filesystem, private local MCP, and other effectful work through the shared controller/runner execution substrate;
- select the sandbox backend from server-owned policy and runner capabilities;
- use a dedicated VM or external fixed service only when the workload, credential, regulatory, or host-exposure requirement justifies it.
The same agent session may use more than one execution boundary. A remote model call can remain in the service, an HTTP or MCP API tool can run through light-gateway, a code-repair action can run in a Cube or Docker sandbox, and a publish action can run in a separate fixed service.
Support three product profiles through the same agent control plane:
- enterprise business agents use remote model providers and typed API/MCP tools through light-gateway;
- coding agents run a workspace-aware model/tool loop in a runner-managed sandbox;
- personal assistants use the same session, memory, policy, and skill model, but receive messages and proactive triggers through a separately deployed channel gateway and use an edge runner for local-device effects.
These are runtime profiles, not forks of the agent engine. Share the durable agent domain model, policy evaluator, skill resolver, runtime protocol, and audit vocabulary. Add a separate executable only where the trust boundary or process lifecycle is materially different.
Problem
Agent definitions, chat sessions, model calls, tool calls, and local execution have different lifecycles and trust boundaries. Treating all of them as one process creates two bad extremes:
- one shared process receives every tenant credential, workspace, tool, and side effect; or
- every logical agent permanently owns a container or VM even when it only makes a bounded remote model call.
The first is unsafe. The second is expensive, slow to scale, and ties metadata to infrastructure unnecessarily.
Agent execution also differs from a workflow task:
- a session can contain many turns;
- a turn can contain several model calls and tool actions;
- the client expects interactive streaming and cancellation;
- multiple turns may share conversation memory;
- a coding session may optionally reuse a workspace;
- an agent can ask for human approval without keeping compute alive.
The execution substrate can be shared with workflows, but session and turn orchestration remain owned by light-agent.
Current Runtime Boundary
The current apps/light-agent executable is a long-lived Axum service.
At startup it creates one process-wide AgentState containing:
- one model provider and model;
- one MCP gateway client;
- one portal query client and portal credential;
- one PostgreSQL pool and memory store;
- one host identity;
- one optional agent definition ID;
- one catalog cache.
The service exposes a WebSocket chat route. Each connection:
- accepts or creates a session ID;
- uses the session UUID as its memory-bank ID;
- loads conversation history;
- accepts user messages sequentially on that socket;
- recalls memory;
- selects portal catalog tools and, because every currently executable entry is gateway-placed, intersects them with gateway tools/list;
- runs up to ten model/tool iterations;
- calls selected tools through light-gateway;
- persists the final conversation history and experience.
Docker Compose and Kubernetes deploy light-agent as a persistent service. The current account, advisor, and technical-support scripts use distinct service identities, which makes one deployment per configured agent profile the practical short-term model.
This implementation is a useful service foundation, but it is not yet a durable or strongly isolated agent execution engine.
Current Gaps
The first implementation work must close these gaps before broad multi-user or effectful use:
- A caller-supplied sessionId is not bound inside light-agent to an authenticated user and agent definition before memory is loaded.
- The existing memory schema can store user_id and agent_def_id, but current session-bank creation does not populate those ownership fields.
- Catalog selection limits the tool specifications shown to the model, but a returned tool name is not revalidated against the accepted set immediately before tools/call.
- Tool arguments fall back to an empty object on malformed JSON instead of failing closed and being checked against the selected input schema.
- There is no durable agent-turn or tool-attempt record. A crash after an effectful tool call and before history persistence can leave an unknown outcome that a reconnect may repeat.
- Concurrent connections using the same session can race history updates.
- Turn-level deadlines, token/cost budgets, tool-call budgets, cancellation, output limits, and concurrency quotas are incomplete.
- Tool results are inserted into model context without a strict byte/token limit or an explicit untrusted-content boundary.
- CLI providers spawn local child processes. A child inherits the service environment unless explicitly scrubbed and can therefore see process-wide credentials.
- Claude Code agent mode currently requests its permission-bypass mode. It must never run inside a shared credential-bearing agent service.
- Local helper scripts contain default bearer-token literals. Those values must be removed and rotated regardless of whether they were intended only for development.
- Portal-command is the production memory-write default. Direct PostgreSQL writes are retained only as an explicitly enabled local/development compatibility mode.
- The current MCP client path forwards the caller Authorization header to light-gateway. It does not yet exchange it for a token narrowed to the agent, turn/action, tool, data boundary, and policy digest.
Goals
- Preserve low-latency interactive chat and streaming.
- Keep logical agent definitions independent from deployment units.
- Bind every session and turn to authenticated tenant, host, user, agent, and policy identities.
- Serialize concurrent same-session prompts through a bounded durable queue.
- Provide durable, idempotent turn and tool-action state.
- Reuse runner registration, scheduling, leases, fencing, watchdog, artifact, credential, and backend contracts.
- Keep workflow tasks and agent turns under their respective orchestrators.
- Route remote API tools through light-gateway.
- Downscope caller authority to the exact agent turn/action and data boundary before gateway dispatch.
- Route local or effectful execution through an approved ExecutionBackend.
- Support task-scoped and bounded session-scoped sandboxes.
- Support enterprise, coding, and personal-assistant profiles without forking the agent domain model.
- Treat Codex, Pi, Claude Code, Gemini CLI, Kilo, and similar products as agent runtime adapters rather than ordinary model providers.
- Materialize one centrally governed skill into profile-specific prompt, schema, package, and sandbox inputs.
- Normalize messaging channels and proactive triggers into authenticated, idempotent agent turns.
- Allow typed agent-to-workflow and workflow-to-agent handoffs without moving interactive turn ownership into light-workflow.
- Fail closed when a deployment cannot satisfy the required boundary.
- Release action leases, model channels, and action credentials while waiting for human approval. Clean task sandboxes; retain a non-secret session workspace only through a distinct bounded hold/checkpoint policy.
Non-Goals
- Do not convert every chat message into a workflow instance.
- Do not give every agent definition a permanent container or VM.
- Do not make controller-rs the owner of conversation or workflow state.
- Do not let the model choose its isolation boundary or credentials.
- Do not let local runner configuration weaken server-owned policy.
- Do not treat a Docker container, Kubernetes pod, or Toolbx as a universal security boundary.
- Do not expose publish, signing, deployment, or unrestricted shell credentials to a general agent loop.
- Do not require a sandbox for a bounded remote model call with no local effects.
- Do not rely on the UI disabling the composer to serialize session mutations.
- Do not use a workflow instance as the inner loop for every chat message, coding command, or personal-assistant action.
- Do not let light-workflow or a channel gateway directly spawn an external agent CLI.
- Do not treat a skill package, repository instruction, plugin, or generated skill as authorization to gain tools, credentials, network, or filesystem access.
- Do not place messaging-channel credentials, model-provider credentials, tenant API credentials, and unrestricted local-device access in one shared process.
Concepts
Agent Definition
Versioned metadata describing instructions, skills, model policy, tool policy, memory policy, data boundary, and default execution profile.
An agent definition is content. It does not own a process.
Agent Product Profile
A server-owned profile selecting the turn lifecycle, ingress surfaces, runtime placement, default tools, memory policy, sandbox requirements, and deployment boundary for an agent definition.
The initial values are enterprise, coding, and personal-assistant. A profile narrows the effective policy; it does not grant authority by itself.
Agent Runtime Adapter
A versioned adapter that runs one model/tool loop and emits normalized runtime events. Native light-agent reasoning, Pi RPC/SDK, Codex, Claude Code, Gemini CLI, and other external harnesses implement this boundary.
An agent runtime is not a model provider. A model provider performs inference; an agent runtime may own a session, tools, local state, approvals, and repeated model calls.
Agent Runtime Host
The small light-agent-worker executable launched inside a runner-managed
sandbox. It verifies the leased runtime specification, materializes approved
skills and context, starts exactly one runtime adapter, streams normalized
events, and exits or checkpoints at the lease boundary.
It does not authenticate end users, own conversation history, choose policy, or accept arbitrary executable paths from a prompt.
Channel Gateway
The optional light-agent-channel executable that owns messaging-platform
connections, webhook verification, user/channel pairing, delivery receipts,
and channel credentials. It converts inbound messages, scheduled triggers, and
connector events into authenticated idempotent turn requests.
It is an ingress and delivery adapter, not an agent engine and not a general execution environment.
Agent Service Instance
A long-lived light-agent process or replica serving compatible agent definitions and sessions. An instance has one deployment trust boundary, network zone, service identity, and set of provider/credential capabilities.
The current implementation has one model/provider and one optional agent definition per instance. Supporting multiple definitions in one pool requires request-time immutable definition resolution and a cache keyed by host and agent definition.
Agent Session
An authenticated conversation scope bound to:
- tenant and host;
- user or service principal;
- agent definition and version;
- memory policy and bank;
- data boundary;
- optional sandbox session;
- creation, idle, and maximum lifetime;
- optimistic version or active-turn fence.
A session ID is an opaque server-issued reference, not proof of access.
Agent Turn
One accepted user or service request and its resulting model/tool loop. A turn has a durable ID, idempotency key, policy snapshot, budgets, state, timestamps, and terminal result.
Agent Action Attempt
One effectful tool or local execution attempt within a turn. It has an idempotency key, attempt number, lease, fencing token, approval state, result, and reconciliation state.
Read-only remote gateway calls may use a lighter audit record, but side-effecting or sandboxed actions require a durable attempt.
Execution Subject
The origin-neutral identity carried by the controller/runner protocol:
subject.kind = workflow-task | agent-turn | agent-action
subject.id
subject.attempt
origin.service
origin.instance
Workflow correlation and agent-session correlation are optional typed extensions. They are not mandatory fields in the runner transport.
Sandbox Session
An optional backend environment reused across related turns or actions under one immutable policy, principal, agent definition, workspace base, and expiry. It is separate from the chat session. Most chat sessions need no sandbox.
Ownership
| Component | Authority |
|---|---|
| light-agent | Agent session, turn, model loop, memory policy, action intent, approval wait, final response |
| light-workflow | Workflow instance, workflow task, workflow retry, workflow approval, workflow transition |
| controller-rs | Runner admission, capacity, reservation, lease transport, heartbeat, quarantine |
| light-workflow-runner | Lease validation, local journal, backend lifecycle, bounded execution, cleanup evidence |
| light-agent-worker | Leased sandbox-side runtime hosting, skill/context materialization, normalized event streaming, process-tree shutdown |
| light-agent-channel | Messaging connection, webhook verification, channel/principal binding, trigger normalization, response delivery |
| ExecutionBackend | Backend-specific preparation, inspection, execution, logs, artifacts, cancellation, cleanup |
| light-gateway | API/MCP authentication, authorization, routing, network policy, and response controls |
| model provider | Model inference only; its output is untrusted input to policy enforcement |
| agent runtime adapter | One bounded model/tool loop behind the normalized runtime protocol; no authority to widen its lease |
| fixed action/service | Structured publish, signing, deploy, push, or other high-value operation |
controller-rs and the runner are origin-neutral. They do not advance an agent turn or workflow task. The origin service reconciles the fenced result into its own state.
Execution Modes
Use explicit modes. Do not silently redirect one mode to another.
| Mode | Owner | Intended use | Local execution |
|---|---|---|---|
| native-workflow | light-workflow | Classification, summarization, branching, schema-bound JSON | None |
| agent-service | light-agent | Interactive chat, memory, remote model, gateway tool loop | None by default |
| runner-agent | light-agent or light-workflow | Files, shell, code, browser, local MCP, private tenant tools | ExecutionBackend |
| channel-agent | light-agent-channel plus light-agent | Messaging, scheduled triggers, personal-assistant ingress and delivery | None in channel gateway |
| fixed-action | dedicated typed service or runner template | Publish, sign, deploy, branch/PR, high-value credentials | Fixed structured operation |
Native Workflow Agent
Keep bounded workflow reasoning in light-workflow. It receives workflow-safe context, calls an approved remote model, validates structured output, and returns control to explicit workflow tasks.
It receives no filesystem, local shell, dynamic tools, tenant workspace, or release credentials.
Agent Service
Use a long-lived light-agent service for interactive sessions:
- WebSocket or streaming chat;
- Hindsight memory;
- remote model providers;
- portal catalog caching;
- dynamic gateway tools/list and tools/call;
- independently scaled specialist agents.
The service container is an application isolation boundary, not a safe place to execute arbitrary code. It should have no workspace mount, host container socket, build tools, browser automation, or unrestricted local MCP server.
Runner Agent
Use runner-agent mode when a turn needs:
- checked-out repositories or mutable files;
- shell or language runtimes;
- browser automation;
- CLI-based model agents;
- local MCP servers;
- private tenant network access;
- code generation, repair, or tests;
- untrusted tool packages or scripts.
For these cases, either:
- keep the model loop in light-agent and lease individual local actions; or
- place the entire model/tool loop in the sandbox when a CLI agent or workspace-aware model must observe and mutate local state.
The second model is required for Codex-, Pi-, and Claude Code-style execution.
Host it with light-agent-worker; do not start a second copy of the public
light-agent service inside the sandbox. The shared service must not spawn an
external agent CLI with its own environment.
Per-action leasing remains useful for a native service-side loop that needs one isolated command. A workspace-aware external runtime receives one bounded agent-turn lease, with optional policy-compatible session reuse, because its filesystem observations, command sequence, and model context form one local execution loop.
Fixed Action
Publishing, signing, deployment, final tags, branch push, and pull-request creation use fixed actions with structured inputs. They consume immutable artifacts or an accepted canonical patch and receive a fresh scoped credential.
They do not execute arbitrary commands from the agent or mutable workspace.
Recommended Architecture
Portal / CLI / API Messaging / schedule / connector event
| |
| light-agent-channel
| |
+------------- authenticated turn ----+
|
v
light-agent
session and policy
+-------------+-------------+
| | |
v v v
model provider light-gateway light-workflow
API / MCP durable process
|
v
controller-rs
|
v
light-workflow-runner
|
v
ExecutionBackend
|
v
task/session sandbox
|
v
light-agent-worker
+ runtime adapter
Fixed high-value effects remain separate typed services/actions.
The normal interactive path does not allocate a sandbox. A sandbox is allocated only when effective policy and the requested action require local execution.
Product Profiles
Enterprise Business Agent
Use the enterprise profile for API- and MCP-centered business processing. The model loop stays in the long-lived light-agent service. The effective catalog exposes only assigned and currently executable gateway tools. Durable or regulated multi-step processing is delegated to light-workflow.
This profile has no workspace mount, local shell, browser process, external agent CLI, or personal channel credential.
Coding Agent
Use the coding profile for repository inspection, code changes, builds, tests,
local MCP, and developer tooling. The whole workspace-aware loop runs through
light-agent-worker in a task-scoped sandbox by default. A bounded
session-scoped workspace is an optimization that requires the same principal,
agent definition, repository base, policy, runtime adapter, backend, and
expiry.
The runtime adapter may be native or may wrap an external harness such as Pi, Codex, or Claude Code. The adapter is selected by immutable server policy and image/package identity, never by a prompt-supplied command. Provider access is brokered by a runner-owned service outside the untrusted payload boundary; raw provider keys and reusable proxy bearer tokens are not copied into the sandbox. The worker receives only a peer-bound local channel for the current attempt, model allowlist, data-boundary and policy digests, token/cost budget, rate, and expiry. The generated-code process runs under a different identity and process/mount namespace and cannot inspect, reconnect to, or inherit that channel.
The untrusted workspace can produce a patch and diagnostic artifacts. Trusted runner code computes the canonical diff, enforces protected paths, and exports immutable artifacts. Branch, pull-request, push, publish, signing, and deploy remain fixed actions over the accepted patch or commit.
Personal Assistant
Use the personal-assistant profile for long-lived user memory, messaging channels, proactive schedules, personal connectors, browser tasks, and optional local-device access.
light-agent-channel owns platform-specific connections and principal pairing.
It does not hold model-provider or general tenant credentials. Typed remote
connectors execute through light-gateway. Browser, filesystem, desktop,
home-automation, or other user-local effects execute through a dedicated or
user-owned edge runner with an explicit capability policy.
A logical personal assistant does not require a permanent VM. A dedicated service or runner is required when personal OAuth grants, private-network access, legal boundaries, or local-device capabilities cannot share a service pool safely.
Scheduled and connector-triggered turns enter the same durable per-session queue as user prompts, carry an origin and idempotency key, and obey quiet hours, rate, cost, approval, and notification policy. A proactive trigger cannot interrupt an active turn or silently act as the user.
Agent Runtime Protocol
Define a versioned agent-runtime-protocol shared by light-agent,
light-agent-worker, runner adapters, and test fixtures. It is separate from the
model-provider trait and from the controller/runner lease protocol.
The runtime specification includes:
- runtime adapter ID, version, immutable image/package digest, and capability digest;
- agent, session, turn, and execution correlation;
- bounded context and selected skill-package digests;
- workspace base, writable roots, protected paths, and change policy;
- model, tool, network, credential, approval, resource, artifact, and deadline policy;
- optional checkpoint/session identity and compatibility digest;
- a one-time event-stream authentication handle.
The runtime emits ordered, bounded events such as:
runtime.startedandruntime.ready;model.started,model.delta,model.completed, and usage;tool.requested,tool.started,tool.result, andtool.failed;approval.requestedandapproval.resolved;workspace.changedandartifact.proposed;checkpoint.created;turn.completed,turn.failed,turn.cancelled, orturn.unknown.
Each event carries the execution ID, turn/action identity, monotonically increasing sequence, event ID, policy digest, timestamp, and bounded payload or artifact reference. Duplicate events are idempotent. Missing sequences can be resumed from the worker journal. An event is evidence; only light-agent can accept it into agent-domain state.
The protocol supports start, cancel, inspect, checkpoint, resume, and bounded input/approval responses. It does not expose a generic remote shell endpoint.
Runtime Capabilities And Adapters
A runtime capability document declares whether the adapter supports workspace mutation, native tools, streaming, interruption, approval suspension, checkpoint/resume, project-local instructions, local MCP, and session reuse. Server-owned compatibility policy maps a tested adapter version to the capabilities it may claim.
For local execution it also carries an immutable runtime-tool manifest: stable
internal tool reference, model-facing alias, input/output schema digests,
effect class, required capability, and dispatch adapter for each shell,
filesystem, browser, or local-MCP operation. At turn admission, server policy
intersects that manifest with the execution profile and lease allowedTools;
where the runtime supports live enumeration, the worker intersects it again
with the current local tool set. A runtime self-report can narrow availability
but cannot add authority absent from the server-owned compatibility record.
The first adapters should be:
- a deterministic mock adapter for protocol, recovery, and fencing tests;
- a native bounded adapter using shared Light-Agent core logic;
- one SDK/RPC-based coding adapter, with Pi as the preferred first candidate;
- subprocess adapters for Codex, Claude Code, Gemini CLI, or Kilo only after their non-interactive event and approval contracts are pinned and tested.
Do not scrape terminal presentation output when a structured SDK, RPC, or JSON event mode exists. Never enable a permission-bypass flag as a substitute for the platform sandbox and approval policy.
Agent And Workflow Interoperation
Agent and workflow orchestration are bidirectional but retain separate domain ownership.
An agent starts a workflow through a typed gateway/API tool when a skill needs durable branching, retries, assertions, human tasks, or long waits. The agent stores the workflow instance reference and may stream or poll its public status; it does not reproduce the workflow steps in its own model loop.
The existing workflow call.agent behavior remains the backward-compatible
native-workflow mode: light-workflow performs a bounded model call and
validates schema-bound JSON without an interactive session or local tools. A
future explicit agent-service mode submits a typed agent job to light-agent.
Light-agent may satisfy that job in its service or through runner-agent
placement according to the selected definition and policy.
light-workflow never directly launches an external agent binary and never mutates agent session history. light-agent never advances workflow tasks. Handoffs carry a correlation ID, caller and tenant binding, input/output schema, deadline, idempotency key, budget, cancellation policy, and bounded delegation depth. Cyclic delegation and unbounded agent/workflow recursion are rejected.
Shared Runner Contract
The workflow runner protocol should be origin-neutral before its first stable version. Do not require processId and taskId in every wire message.
Every scheduling request and lease carries:
- execution ID;
- origin service and authenticated origin instance;
- subject kind, ID, and attempt;
- tenant and host derived from trusted identity;
- policy snapshot and digest;
- execution profile and compatibility digest;
- runner/backend selection;
- lease ID and fencing token;
- deadlines and cleanup policy;
- idempotency key;
- optional typed workflow or agent correlation.
Example standalone agent action:
{
"executionId": "01970f5d-2222-7000-8000-000000000001",
"origin": {
"service": "light-agent",
"instance": "account-agent-east"
},
"subject": {
"kind": "agent-action",
"id": "01970f5d-2222-7000-8000-000000000020",
"attempt": 1
},
"agent": {
"sessionId": "01970f5d-2222-7000-8000-000000000010",
"turnId": "01970f5d-2222-7000-8000-000000000011",
"agentDefId": "01970f5d-2222-7000-8000-000000000012"
},
"leaseId": "01970f5d-2222-7000-8000-000000000030",
"fencingToken": 19,
"policyDigest": "sha256:...",
"executionProfile": "agent-microvm",
"commandTemplateId": "agent-tool-cargo-test",
"deadlineAt": "2026-07-10T20:30:00Z",
"expiresAt": "2026-07-10T20:10:30Z"
}
The controller authenticates which services may submit each subject kind. light-agent cannot submit workflow-task work, and light-workflow cannot mutate an agent turn merely because both use the same runner.
Persistence Split
Use common execution tables for controller/runner state:
- runner_scheduling_request_t;
- execution_attempt_t;
- runner session/backend capability records;
- execution session, artifact, and runtime audit records where sharing is appropriate.
Use origin-specific tables for domain state:
- task_info_t and workflow approval/transition records for workflow tasks;
- agent_session_t, agent_turn_t, agent_action_attempt_t, and the ordered session event stream for agent work.
The common attempt stores subject identity, lease, fencing, backend, normalized result, and cleanup. The origin transaction conditionally accepts that result and advances only its own domain object.
This split is a design decision:
- every runner-backed agent action references one shared execution_attempt_t row;
- execution_attempt_t contains only controller/runner concerns such as origin, subject, attempt, reservation, lease, fencing, runner/backend, normalized result, and cleanup;
- agent_action_attempt_t contains agent-domain concerns such as tool identity, model iteration, argument digest, effect class, approval, budgets, recovery policy, and acceptance into the turn;
- gateway-only actions can have an agent_action_attempt_t without an execution_attempt_t and instead record the gateway request/idempotency identity;
- controller-rs never writes agent turn, history, or approval state.
Result-Ready Wakeup
The common execution row remains the durable source of truth. In the same
PostgreSQL transaction that conditionally stores a newly terminal
execution_attempt_t, controller-rs emits a versioned
execution_result_ready_v1 notification containing only the attempt ID,
authenticated origin, subject kind, and correlation ID. It contains no result
bytes, tenant content, or authorization.
light-agent listens for the notification, loads the authoritative row, verifies
origin/subject/fencing bindings, and conditionally accepts the result in its
own domain transaction. It must also scan indexed unaccepted terminal attempts
at startup and periodically. The listener uses a dedicated connection and, on
startup or reconnect, establishes LISTEN before its catch-up scan so a commit
in that handoff window is either scanned or queued. LISTEN/NOTIFY is only a
low-latency wakeup:
delivery can be missed, duplicated, or reordered. A future typed controller
callback may provide another wakeup, but it cannot replace the authoritative
query or make controller-rs write agent tables.
Session And Turn Model
Session Admission
The front door authenticates the caller before accepting or resuming a session. The server derives:
- tenant and host;
- user or service principal;
- allowed agent definition;
- model/provider and data-boundary policy;
- memory scope;
- maximum session and idle lifetime.
On new session, light-agent creates a server-issued session ID and stores the ownership binding. On resume, all binding fields must match. A valid UUID alone never grants access.
The existing agent_memory_bank_t.user_id and agent_def_id fields should be populated. agent_session_history_t should either gain explicit ownership columns or reference a new agent_session_t that contains them.
Proposed Agent Tables
agent_session_t:
- tenant, host, session, user/service principal, and agent definition;
- definition, model, tool, memory, and execution policy digests;
- memory bank ID;
- optional execution session ID;
- state, optimistic version, active turn, created/last/idle/max expiry;
- cancellation, revocation, and retention state;
- durable execution-session cleanup request, state, and evidence correlation.
agent_turn_t:
- session and monotonically increasing turn sequence;
- origin kind such as user, channel, workflow, scheduler, or connector plus an immutable origin reference and bounded delegation depth;
- client message ID and idempotency key;
- immutable prompt/input reference and policy snapshot;
- model/provider reference and data boundary;
- QUEUED, RECEIVED, RUNNING_MODEL, WAITING_ACTION, RUNNING_ACTION, WAITING_RECONCILIATION, WAITING_APPROVAL, COMPLETED, FAILED, CANCELLED, or UNKNOWN state;
- enqueue sequence, queue deadline, activation time, and optional cancellation reason;
- token, cost, model-call, action-call, and wall-clock budgets;
- accepted result, error class, timestamps, and audit correlation.
agent_action_attempt_t:
- turn and action/tool identity;
- stable internal tool reference, model-facing alias, tool source, schema digest, effect classification, and selected execution placement;
- optional runtime adapter ID/version, runtime action ID, and capability digest;
- input schema and canonical argument digest;
- approval requirement and binding;
- numbered logical attempt and optional superseded/resumed-from attempt;
- nullable execution_attempt_id referencing the common execution_attempt_t row where runner-backed;
- gateway request/idempotency identity where remotely executed;
- known-success, known-failure, cancelled, or unknown outcome;
- recovery classification and remaining correction budget;
- bounded result/artifact references and reconciliation state.
agent_approval_t:
- approval ID, session, turn, canonical action intent and argument digest;
- tool/operation, destination, data-boundary and policy digests, artifact or patch bindings where applicable, actor authority, state, and expiry;
- source attempt when a running runtime discovered the approval boundary;
- optional execution-session approval-hold ID and bounded hold expiry;
- consuming post-approval agent-action and common execution-attempt IDs;
- a unique active approval per exact subject and single-use consumption.
agent_session_event_t, or an equivalent append-only portal event stream:
- session sequence, event ID, turn ID, optional action-attempt ID, and event type;
- USER_MESSAGE, MODEL_RESPONSE, ACTION_DISPATCHED, ACTION_RESULT, APPROVAL_REQUESTED, APPROVAL_DECIDED, TURN_TERMINAL, or SYSTEM event;
- immutable content reference/digest, source class, policy digest, timestamp, and actor;
- unique action-result event per accepted agent action attempt.
agent_session_history_t is a rebuildable conversation-context projection over the ordered event stream. It is not the authoritative ledger proving that an effectful action occurred.
Personal-assistant deployments also require channel-domain records, either in the GenAI schema or a dedicated channel service:
agent_channel_binding_tbinds tenant, principal, agent, platform, channel, remote identity, pairing/verification state, and revocation without storing raw channel secrets;agent_channel_delivery_tdeduplicates inbound platform events and outbound responses, records delivery state, and references the resulting turn;- scheduled trigger records bind agent, session/origin, schedule, quiet-hours, notification, idempotency, and maximum-delay policy.
Channel records do not replace agent turns. They prove ingress and delivery; light-agent remains authoritative for reasoning and action state.
Turn State Machine
QUEUED -> RECEIVED -> RUNNING_MODEL
RUNNING_MODEL -- no tool --> COMPLETED
RUNNING_MODEL -- tool --> WAITING_ACTION
WAITING_ACTION -- no approval --> RUNNING_ACTION
WAITING_ACTION -- approval required --> WAITING_APPROVAL
WAITING_APPROVAL -- approved; allocate new attempts --> RUNNING_ACTION
WAITING_APPROVAL -- rejected --> RUNNING_MODEL or FAILED by policy
RUNNING_ACTION -- known success/recoverable failure --> RUNNING_MODEL
RUNNING_ACTION -- uncertain outcome --> WAITING_RECONCILIATION
WAITING_RECONCILIATION -- known recoverable result --> RUNNING_MODEL
WAITING_RECONCILIATION -- terminal/unsafe --> FAILED
WAITING_RECONCILIATION -- cannot determine --> UNKNOWN
Any active state -- accepted cancellation --> CANCELLED
Policy/security violation or exhausted hard budget --> FAILED
A model call may be safely retried only when it has no external effect or the provider request is idempotent. An action with an unknown outcome is inspected or reconciled before another attempt.
An action failure does not automatically fail the turn. A known, policy-allowed recoverable failure such as a compiler error, failed test, linter result, or non-zero diagnostic command is persisted as untrusted tool output and returned to RUNNING_MODEL when correction budgets remain. The model may explain the failure or propose a new action.
The turn becomes FAILED or CANCELLED when:
- policy, authentication, schema, or security enforcement rejects the action;
- approval is rejected and policy treats rejection as terminal;
- a hard turn deadline or token/cost/action budget is exhausted;
- the client or control plane cancels the turn;
- the action is classified non-recoverable;
- an unknown side effect cannot be reconciled and policy requires termination.
A correction is a new action identity. Reusing an attempt is allowed only when the backend/external operation has a proven idempotency or inspection contract. maxCorrectionActions and per-tool retry limits prevent an agent from looping on the same failure. A turn may still finish COMPLETED with a user-facing explanation that one or more actions failed; completion means the response was durably delivered, not that every action succeeded.
WAITING_APPROVAL is durable agent state. It always ends the current action lease, model-broker capability, and action credential. A task-scoped sandbox is cleaned after its required evidence is exported.
A reusable agent-session workspace has a separate lifecycle. Under an
explicit bounded non-secret retention policy, its execution_session_t may
enter IDLE_APPROVAL_HOLD with no executable action, tool credential, or model
channel. The hold expires at the earliest of approval expiry, session idle/max
expiry, cost/retention policy, broker/credential boundary, and backend-native
TTL. Pause or a verified checkpoint is preferred over consuming active compute.
Absence of an action lease is not by itself a session-cleanup signal.
The origin may renew the session-retention record only from authenticated session activity and never beyond the fixed maximum; it must not fake action lease heartbeats while a person decides. If the backend cannot safely retain or checkpoint the non-secret workspace, the runner exports an approved immutable patch/checkpoint and cleans the sandbox. Preserving important uncommitted work must not depend only on a live sandbox.
The origin transition into WAITING_APPROVAL atomically persists exactly one
session disposition: cleanup, or a policy-valid bounded hold. If common session
state later lives in another database, use an idempotent transactional outbox.
There must be no interval where a session reaper can interpret the ended action
lease as abandonment before the hold is durable.
Approval never reactivates an execution attempt. If policy knows approval is
required before dispatch, light-agent records the bound action intent and
approval but creates no common execution attempt. If a running runtime
discovers an approval boundary, it returns a known approval_required
terminal result; controller-rs ends the action lease and the runner revokes
grants and cleans or checkpoints according to policy. After approval,
light-agent consumes
the approval into a new numbered agent action attempt and a new common
execution attempt with a fresh lease and monotonic fencing token. The previous
attempt and its backend handle, grants, and fencing token remain immutable and
cannot resume execution. A retained session workspace is reused only after
principal/base/policy/runtime/expiry compatibility and cleanup state are
revalidated; otherwise the new action starts in a fresh sandbox and restores
only a verified policy-permitted checkpoint or patch.
Concurrency
Only one mutating turn should own a session version at a time by default. Multiple WebSockets or replicas must not overwrite the same history.
The default user experience is a bounded durable server-side FIFO per session, not immediate rejection:
- Authenticate and authorize the prompt, deduplicate its client message ID, assign the next session enqueue sequence, and persist a QUEUED turn.
- Return the turn ID, state, queue position, and estimated/retry timing to the client. The UI may disable or label the composer, but correctness does not depend on client-side serialization.
- When no active mutating turn exists, conditionally acquire the session version, revalidate revocation and current policy, snapshot the effective turn policy, and activate the oldest non-expired queued turn.
- Allow a user to cancel a queued turn. Interrupting an active turn requires an explicit cancel-and-enqueue operation; a second prompt never implicitly cancels in-flight work.
A queued prompt is durable but is not added to the active turn’s model context or mutable history projection. It becomes eligible for conversation context only after it wins FIFO activation, so a later prompt cannot change the meaning of an in-flight action.
Queue depth and wait time are bounded per tenant, principal, agent, and session. A full queue returns a retryable admission response such as 429 with retryAfter. A 409 is reserved for a stale explicit session version or an operation that semantically requires exclusive ownership; it is not the normal second-prompt response.
Use a conditional active-turn or aggregate-version update across replicas. A read-only secondary view can stream state, but it cannot bypass the FIFO or append another active user turn without winning session activation.
Agent Definition And Policy Snapshot
Resolve and snapshot at turn admission:
- agent definition/version;
- system instructions and selected skills;
- model/provider and regional/data-boundary policy;
- memory scope and retention;
- permitted catalog and tool policy;
- action execution profile;
- network and credential profile;
- turn/model/tool/token/cost limits;
- approval rules;
- protected workspace policy;
- artifact and audit policy.
Do not execute a long turn from mutable current rows. A catalog refresh may narrow executable tools immediately for emergency revocation, but it cannot widen the accepted snapshot without a new authorization decision.
For future multi-agent pooling, cache by host, agent definition ID, version, and policy digest. Never use one global catalog entry across definitions.
Centralized Skills Across Profiles
The centralized registry is the source of assigned skill identity, version, instructions, taxonomy, tool/workflow links, runtime compatibility, and governance metadata. It is not the process that executes a skill.
At turn admission, light-agent resolves an immutable effective skill set and records every selected version and digest. A profile-specific materializer then produces only the inputs required by the selected runtime:
| Profile/runtime | Materialized form |
|---|---|
| Enterprise agent | Bounded prompt instructions plus selected gateway tool schemas |
| Native workflow agent | Bounded instructions, structured input, and required output schema |
| Coding runtime | Read-only SKILL.md, references, and signed script/assets package inside the sandbox |
| Personal assistant | Instructions, connector/tool mappings, schedule/notification constraints, and optional reviewed package |
| Workflow-backed skill | Instructions plus a typed workflow reference and start contract |
Skill content is layered in decreasing authority:
- server and execution policy;
- signed platform/tenant skill versions assigned to the agent;
- reviewed user-specific skill configuration;
- repository or workspace-local instructions;
- user prompts, retrieved content, and tool output.
Lower layers cannot override higher-layer policy. Repository instructions and downloaded or generated skills are untrusted content even when useful to the model. A self-generated skill is stored as an inactive proposal and requires validation, scanning, review, immutable packaging, and an explicit assignment before another turn can load it.
Do not execute source code copied directly from a mutable database row. Script
or binary content belongs in an immutable artifact with digest, provenance,
scanner results, entrypoint metadata, and a required sandbox profile. The
trusted runner—not light-agent-worker or generated code—downloads the
selected immutable packages before sandbox creation, verifies their
digest/signature/size and archive safety, and stages them as read-only mounts
with nodev, nosuid, and noexec unless a reviewed profile requires an
executable entrypoint. The worker revalidates the mounted manifest before use.
Neither the worker nor payload receives artifact-store credentials or package
download egress.
See Centralized Skills for the catalog and package model and Skill Workflow Orchestration for workflow-backed skills.
Tool Authorization And Execution
Treat model tool calls as untrusted requests.
For each model iteration:
- Resolve the effective catalog for the authenticated agent and turn.
- Apply lifecycle, sensitivity, effect, approval, tenant, cost, and network policy.
- Partition candidates by the server-owned execution placement recorded in the catalog/policy snapshot: gateway, runner, workflow, or fixed service.
- For gateway candidates, intersect with live gateway
tools/listandtoolsListAccessControlunder the downscoped turn identity. - For runner candidates, intersect with the execution profile, lease
allowedTools, server-approved runtime-tool manifest, and live worker or sandbox-local MCP enumeration where supported. Do not require these tools to exist in gatewaytools/list. - Expose workflow and fixed-service candidates only through their typed contracts; they are never converted to free-form local or gateway tools.
- Form a collision-free union. Bind each model-facing name to its internal tool reference, placement, schema digest, and policy snapshot. A duplicate alias across placements fails closed unless server policy assigned distinct deterministic aliases.
- Send only that accepted set to the model.
- On returned tool call, recheck that the exact bound tool remains in the accepted set.
- Parse arguments strictly. Malformed JSON fails; it does not become an empty object.
- Validate arguments against the accepted input schema and routing metadata.
- Re-evaluate effect, approval, quotas, cancellation, policy revocation, and destination immediately before dispatch.
- Compute the effective delegation as the intersection of caller authority, agent-definition policy, turn/action policy, tool policy, and current revocation state.
- For gateway execution, exchange the caller identity for a short-lived downscoped gateway token bound to the turn or exact action.
- Create a durable attempt and idempotency key when the action can have an effect.
- Dispatch only through the placement bound at disclosure; model output cannot change the route.
- Bound, redact, classify, and persist the result before giving it back to the model.
The gateway remains the final API authorization and routing boundary. Agent catalog policy is an additional restriction and must not be bypassed merely because the gateway would accept a broader caller token.
The runner lease and runtime-tool manifest are the corresponding final local availability boundaries. A local tool name is not authority by itself, and the model broker, credential broker, runner control socket, and backend lifecycle API are never included in the tool union.
Gateway Delegation
Production agent calls do not forward the caller’s full bearer token directly to light-gateway. light-agent uses a trusted token-exchange or credential-broker service to mint a signed, short-lived delegated token whose authority can only narrow the caller.
The effective authority is:
caller grants
intersect agent-definition policy
intersect turn/action policy
intersect tool and data-boundary policy
intersect current revocation and quota state
A tools/list token is scoped to the turn and accepted tool set. A tools/call token should be scoped to one action and include or cryptographically bind:
- gateway audience;
- tenant, host, caller subject, and light-agent actor identity;
- agent definition, session, turn, and action IDs;
- exact tool or narrowly bounded tool set;
- allowed scopes, destination/service, sensitivity ceiling, and data boundary;
- policy snapshot/digest and argument or request digest where practical;
- issued-at, short expiry, unique token ID, and replay/idempotency binding.
light-gateway validates the signature, audience, expiry, actor/delegation chain, policy binding, tool, destination, and current authorization. It intersects the delegated token with its own access-control and tool metadata; possession of a more powerful original user token cannot widen an agent turn.
If token exchange is unavailable or a requested binding cannot be enforced, the production call fails closed. Direct forwarding may exist only as an explicit local-development compatibility mode and must never be the default for effectful or sensitive tools.
Placement
| Tool/action | Default placement |
|---|---|
| Remote read-only HTTP/MCP | light-gateway |
| Remote effectful HTTP/MCP | light-gateway plus durable action attempt and approval/idempotency |
| Local command or language runtime | runner ExecutionBackend |
| Filesystem or repository mutation | runner task/session sandbox |
| Browser automation | runner sandbox with network policy |
| Local MCP server | runner sandbox or dedicated tenant service |
| Branch/PR creation | fixed action over accepted patch |
| Publish/sign/deploy | fixed external service or dedicated fixed runner action |
Tool Results
Tool output is untrusted content even when the tool is authorized.
Result handling distinguishes action outcome from turn outcome:
-
known success is persisted and normally returns to RUNNING_MODEL;
-
known recoverable failure is persisted with bounded diagnostics and returns to RUNNING_MODEL when correction policy and budgets allow;
-
known terminal failure ends or cancels the turn according to policy;
-
unknown outcome enters WAITING_RECONCILIATION and cannot be represented to the model as if the action definitely failed;
-
a new corrective tool call receives a new action ID and idempotency decision.
-
enforce byte, item, nesting, and token limits;
-
preserve truncation markers and full artifact references when policy permits;
-
separate tool data from system instructions;
-
do not follow instructions found in tool output unless the agent policy explicitly treats that source as instructions;
-
redact secrets before persistence and again before model context;
-
store the action ID, tool/version, argument digest, authorization decision, destination, result digest, and model iteration.
Model Provider Boundary
Remote API Providers
Remote API providers can run from the long-lived service when:
- the service data boundary permits the prompt;
- provider credentials are service-owned or tenant-approved;
- no local executable is spawned;
- the turn has token, cost, timeout, and concurrency limits;
- response and tool calls are treated as untrusted.
Do not send tenant-local repositories, private logs, or private-network data to a SaaS provider unless the effective policy authorizes that transfer.
CLI And External Agent Runtimes
Codex, Pi, Claude Code, Gemini CLI, Kilo CLI, and similar harnesses are agent
runtimes, not ordinary model API adapters. Existing CLI implementations under
model-provider are compatibility code and should migrate behind the
agent-runtime adapter boundary.
They run under light-agent-worker with runner-agent placement and require:
- fresh task or bounded session sandbox;
- minimal allowlisted environment;
- no inherited portal token, database URL, unrelated provider keys, or controller credential;
- explicit workspace, network, tool, and resource policy;
- local deadline and process-tree cancellation;
- bounded stdout/stderr;
- immutable binary/image identity and capability digest;
- structured SDK, RPC, or JSON event integration where available;
- normalized approval, cancellation, usage, patch, and terminal events;
- cleanup journal and backend-native expiry.
Model access for these runtimes terminates at a runner-owned broker. Prefer a
preconnected descriptor, peer-credential-checked Unix-domain socket, vsock, or
an equivalent backend-local transport. A socket pathname alone is not an
authorization boundary: the broker authenticates the attempt and peer, and
the runner prevents descriptor inheritance, cross-process /proc inspection,
and ptrace. The broker independently enforces the approved model, data
boundary, policy digest, token/cost budget, rate, cancellation, and expiry.
An adapter that can operate only with an extractable provider key is ineligible
for an untrusted coding profile.
Permission-bypass flags are prohibited. If an adapter needs an unattended mode, the platform sandbox and approval policy—not a CLI bypass option—provide the effective boundary. The shared service never invokes these binaries directly, and light-workflow never invokes them at all.
Sandbox Scope And Backend
Backend and session scope are separate decisions.
| Scope | Use | Default |
|---|---|---|
| none | Remote model plus gateway-only tools | Long-lived agent service |
| turn/task | CLI agent, untrusted tool, one repair/action | Preferred strong isolation |
| agent-session | Interactive coding workspace reused across turns | Explicit TTL, same principal/policy/base |
| dedicated | Privileged, regulated, or long-running tenant agent | Dedicated VM or service |
| Workload | Minimum boundary | Candidate |
|---|---|---|
| Bounded remote reasoning | service container | light-agent pod |
| Trusted internal command | shared-kernel-container | Rootless OCI or ordinary Kubernetes Job |
| Autonomous code or untrusted package | microvm | Cube Sandbox or Docker Sandboxes |
| Strong tenant isolation | dedicated-vm | Approved dedicated VM |
| Trusted local developer helper | host-integrated | Toolbx, never represented as a sandbox |
| Publish/sign/deploy | external-service | Fixed typed action or service |
The deployment advertises available backends. Server-owned policy chooses an eligible backend or defers/denies execution. It never silently downgrades.
Session Reuse
An agent sandbox session can be reused only when all of these match:
- tenant, host, principal, agent definition, and policy digest;
- workspace base revision and change policy;
- backend, template/image, and compatibility digest;
- network, credential, model-provider, and tool policy;
- maximum lifetime, idle timeout, and cleanup state.
Credentials remain task-scoped even when the workspace is reused. A session that received a high-value credential is destroyed after the action unless an explicit policy proves safe cleanup.
The effective physical-session expiry is the earliest of the agent session’s idle/max expiry, execution-session policy, broker/grant expiry, and backend-native TTL. An approval hold can preserve a compatible non-secret workspace only until that same effective expiry; it does not refresh or extend the fixed maximum and it carries no action credential or model channel.
When light-agent closes, revokes, or expires a logical
session, the same durable transaction creates an idempotent common
execution-session cleanup request. controller-rs immediately fences and
cancels active attempts and dispatches cleanup; the runner destroys the
backend session and records evidence. Cleanup is retried across restarts and
the backend TTL remains only a final fail-safe, not the expected reclamation
path. An EXPIRED agent session whose physical sandbox is merely waiting for
its independent TTL is a reconciliation defect.
Memory Boundary
Conversation history and distilled memory are domain state, not sandbox state. The sandbox may receive a bounded prompt/context projection, but it does not own the memory database.
Required controls:
- bind memory bank and session history to authenticated host, user/principal, and agent definition;
- authorize every resume, recall, retain, and history update;
- apply optimistic versioning to conversation projections;
- keep accepted action attempts and append-only session events authoritative over the mutable history projection;
- distinguish user-authored, tool-derived, model-derived, and operator instruction sources;
- prevent tool output or retrieved memory from becoming privileged system instructions;
- enforce retention, deletion, legal hold, export, and audit policy;
- redact or tokenize sensitive values before embedding or cross-boundary model transfer.
History Conflict After An Effect
An optimistic history conflict must never cause an effectful action to be forgotten or repeated.
When an action reaches a known terminal result, light-agent performs an idempotent origin-acceptance transaction that:
- conditionally accepts the current agent_action_attempt_t;
- records or references the common execution_attempt_t/gateway result;
- appends one ACTION_RESULT event with the action/result digest;
- advances the agent turn to RUNNING_MODEL, WAITING_RECONCILIATION, or a terminal state.
Updating agent_session_history_t is a projection step after that transaction. If its expected version is stale, the projector rereads the ordered session events, deterministically rebuilds or merges the conversation context, and retries the projection. It does not redispatch the tool and does not overwrite another accepted user message.
Until the projection catches up, clients can reconstruct the authoritative timeline from turn/action/session events. The UI may show a temporary history-sync state, but the accepted action and audit record remain visible. Projection lag or conflict is an operational error, not an action retry signal.
Portal-command memory writes are the production default because they preserve event, authorization, and audit boundaries. Direct PostgreSQL mode remains an explicitly enabled local/development compatibility profile. Longer term, memory recall should also use a scoped service API so the general agent pod does not require a database password.
Authentication And Session Security
- Authenticate before WebSocket upgrade or before accepting the first message.
- Derive tenant, host, user, and allowed agent definition from trusted claims and server-side mappings.
- Issue an opaque session handle or signed resume token with audience, expiry, principal, and agent binding.
- Never use caller-provided tenant, host, user, agent, or memory-bank IDs as authority.
- Rotate the session handle after privilege or policy changes.
- Revoke active sessions on user, agent, provider, or policy revocation.
- Serialize concurrent mutating prompts through the bounded durable per-session FIFO; only the active turn acquires the session version.
- Remove committed/default bearer tokens and rotate any token that may have been usable.
Credentials
The long-lived service should contain only credentials required for its service profile. It should not hold credentials for possible future tools.
- Prefer workload identity and brokered short-lived grants.
- Keep end-user authorization separate from the service’s portal identity.
- Exchange user authorization for a signed, short-lived, audience-restricted, turn/action-scoped light-gateway delegation token. Do not forward the unrestricted user token in production.
- Never forward SaaS model credentials into tenant runner sandboxes.
- Do not replace provider keys with a reusable proxy bearer token visible to the worker or generated payload. Use the protected, peer-bound runner broker channel and enforce budget and expiry at the broker.
- Never let a CLI child inherit the service environment.
- Project action credentials after policy approval and revoke them at terminal state or lease loss.
- Publish/sign/deploy credentials exist only in fixed actions.
- Do not place raw secrets in prompts, session history, memory, tool arguments, logs, artifacts, environment snapshots, or execution journals.
Failure And Recovery
Service Restart
The session and turn are reconstructed from durable state. An incomplete turn is not automatically replayed:
- RUNNING_MODEL with no action may be safely failed or retried by policy;
- RUNNING_ACTION queries the gateway idempotency record or runner attempt;
- an accepted ACTION_RESULT with stale history resumes projection/rebuild and never redispatches the action;
- UNKNOWN action outcome requires inspection or operator decision;
- COMPLETED result can be streamed again idempotently;
- WAITING_APPROVAL has no action compute or credentials; an explicitly retained session workspace remains paused/checkpointed under its independent bounded hold and expiry.
Client Disconnect
Policy decides whether the turn:
- cancels immediately;
- continues to a durable result for later resume; or
- continues only through the current non-effectful model call.
The decision is recorded at admission. A disconnect does not silently broaden the deadline.
Runner Or Backend Disconnect
Use the same lease, fencing, journal, inspection, watchdog, native TTL, and cleanup contract as workflow execution. light-agent accepts a result only for the current action attempt and fencing token. Result-ready notifications only wake reconciliation; startup and periodic scans of authoritative terminal attempts recover any notification lost while light-agent was disconnected.
Duplicate Client Message
The client supplies a message ID scoped to the session. A duplicate returns the existing turn or result. It does not create a second effectful action.
Limits And Admission
Every service profile and turn policy defines:
- maximum concurrent sessions and turns;
- maximum queued prompts per session/principal/agent and maximum queue wait;
- per-principal and per-agent quotas;
- maximum input/history/retrieval/tool-output tokens;
- maximum model calls and tool calls per turn;
- maximum correction actions and per-tool recovery attempts;
- model and total wall-clock deadlines;
- cost/token budget;
- maximum pending approval time;
- sandbox queue and runtime deadline;
- session idle and maximum lifetime;
- artifact and log limits.
Saturation is not a model or tool failure. Queue or reject before starting more work than the service, provider, gateway, or runner can support.
Audit And Observability
Record:
- authenticated session admission and resume;
- agent definition, model, catalog, memory, and execution-policy digests;
- turn state and idempotency decision;
- model provider/model, latency, token usage, and cost without hidden reasoning content;
- selected, hidden, and attempted tools;
- argument and result digests, effect class, approval, destination, and placement;
- runner/backend/lease/fencing identity for sandboxed work;
- cancellation, unknown outcome, retry, and reconciliation;
- memory recall/retain source classes;
- sandbox cleanup and artifact evidence.
Metrics include active sessions, turn latency/state, model/tool counts and budgets, per-session queue depth/wait, session activation conflicts, unauthorized resume attempts, rejected model tool names, malformed/schema-invalid arguments, downscoped-token issuance/rejection, recoverable action failures, unknown actions, history projection lag/conflict, approval wait, runner queue time, runtime-event lag/gaps, adapter failures, skill-package verification failures, channel duplicate/replay rejection, delivery latency/failure, scheduled-turn deferral, result-wakeup/catch-up latency, oldest unaccepted terminal attempt, model-broker denial/budget exhaustion, session-to-sandbox expiry skew, approval-held session count/age/cost, checkpoint/restore failures, and cleanup request latency/backlog.
Deployment Profiles
Shared Tenant Agent Service
One horizontally scalable deployment serves compatible enterprise and personal-assistant reasoning sessions for one tenant or strong tenant partition. It uses remote model APIs and gateway-only tools. It has no local workspace or external agent runtime.
Dedicated Agent Service
Use a separate long-lived deployment when an agent requires a distinct:
- tenant or legal boundary;
- model provider credential or regional endpoint;
- private network zone;
- service identity;
- latency/scaling profile;
- memory retention policy.
This is a deployment profile, not a requirement for every logical agent.
Coding Agent Worker Pool
light-agent submits agent-turn or agent-action execution subjects to
controller-rs. The tenant runner selects an approved backend, creates a task-
or session-scoped environment, and starts light-agent-worker with the pinned
runtime adapter. Worker pools are grouped by real backend, network, workspace,
model-proxy, and data-boundary compatibility, not by logical agent name.
Personal Assistant Channel Gateway
Deploy light-agent-channel separately from the reasoning service. Pool only
channels and users whose webhook exposure, credential store, retention,
regional, and delivery policies are compatible. Channel delivery can continue
while a reasoning replica restarts because accepted messages and responses are
durable.
Personal Edge Runner
Use a dedicated user- or tenant-owned runner when a personal assistant needs a local browser profile, desktop, filesystem, home network, or device access. The edge runner advertises explicit capabilities and receives short-lived leases; it is not a permanently authorized remote shell.
The implemented edge binding is principal-specific and names the exact runner,
backend, execution profile, action allowlist, required capabilities,
compatibility digest, and expiry. Light-Agent converts an allowed local effect
to a fixed light-edge-action structured command and Controller refuses to
reserve it on any other runner or backend. Revoked/expired bindings and action
or capability mismatches fail before scheduling.
Release Agent
Use a dedicated runner pool and strong sandbox/VM profile. The agent can inspect and patch a copy-on-write workspace, but trusted fixed actions apply the accepted patch, create a branch/PR, publish, sign, or deploy. Approval is bound to immutable subjects and waits without an active lease.
Implementation Plan
The detailed repository and pull-request sequence is maintained in
implementation/light-agent/2026-07-10-LightAgentRuntimeAndProfilesImplementationPlan.md.
Phase 0: Harden The Current Service
- Remove and rotate embedded default bearer tokens.
- Authenticate session admission and bind host/user/agent ownership.
- Populate memory-bank ownership and reject unauthorized resume.
- Revalidate model tool names against the accepted per-turn set.
- Strictly parse and schema-validate arguments.
- Exchange caller authorization for downscoped gateway delegation tokens.
- Bound/redact tool output and retrieved memory.
- Add overall turn timeout, cancellation, action/model-call limits, and a bounded per-session FIFO with one active mutating turn.
- Disable CLI providers in shared-service profiles.
Phase 1: Durable Sessions And Turns
- Add agent_session_t, agent_turn_t, agent_action_attempt_t, and the append-only agent session event stream/projection.
- Add message idempotency, durable FIFO sequencing, and optimistic session activation/versioning.
- Snapshot definition, catalog, model, memory, and execution policy.
- Persist model/action state and normalized bounded results; accept an action and append its ACTION_RESULT event before updating conversation history.
- Rebuild the history projection from ordered events after a version conflict.
- Add approval wait without compute.
- Distinguish action-lease release from a bounded execution-session
IDLE_APPROVAL_HOLD; task sandboxes clean immediately while an eligible non-secret session workspace may pause/checkpoint until its independent effective expiry. - Bind approvals to immutable action intents and consume approval into a new action/common execution attempt with fresh fencing; never reopen an earlier attempt.
- On session close, revocation, or expiry, atomically create an idempotent common execution-session cleanup request.
- Move production memory writes to portal-command mode.
Phase 2: Origin-Neutral Runner Contract
- Add execution subject and origin types to execution-runner-protocol.
- Make controller scheduling and common execution attempts origin-neutral.
- Reference shared execution_attempt_t from runner-backed agent_action_attempt_t rows while keeping agent-domain fields out of the common table.
- Authenticate allowed subject kinds per origin service.
- Keep workflow and agent domain transitions separate.
- Emit identifiers-only
execution_result_ready_v1PostgreSQL wakeups in the common terminal-result transaction. Add light-agent LISTEN handling plus startup and periodic indexed catch-up; notification delivery is never the source of truth. - Add durable execution-session cleanup dispatch and evidence reconciliation.
Phase 3: Agent Runtime Protocol And Worker
- Add agent-runtime-protocol, ordered runtime events, and capability documents.
- Add stable tool-source/placement identities and an immutable local
runtime-tool manifest. Gateway candidates intersect gateway
tools/list; local candidates intersect server compatibility, execution policy, leaseallowedTools, and live worker/local-MCP enumeration. - Add a deterministic mock adapter and the
light-agent-workerexecutable. - Route command, filesystem, browser, local MCP, and external agent runtimes through ExecutionBackend.
- Add a runner-owned model-broker transport contract with peer/attempt binding, no payload-visible reusable bearer, separate worker/payload identities, and broker-enforced model and budget policy.
- Reconcile worker events into agent-domain state without giving the worker direct database ownership.
Phase 4: Skill Packages And Materialization
- Add runtime compatibility and immutable skill-package records.
- Verify digest, provenance, scan result, entrypoint, and sandbox policy before read-only materialization.
- Make the trusted runner download, safely extract, verify, and stage selected packages before sandbox creation. The worker only revalidates mounted bytes and has no artifact-store credential or download path.
- Implement enterprise, workflow, coding, and personal-assistant materializers.
- Treat generated and repository-local skills as untrusted proposals/content.
Phase 5: Coding Agent Profile
- Enable Cube Sandbox for untrusted task-scoped coding turns.
- Add the first structured SDK/RPC coding adapter, preferably Pi.
- Add optional bounded workspace-session reuse.
- Add bounded approval-hold, pause/checkpoint, restore, expiry, cost, and cleanup semantics without renewing an action lease.
- Add protected runner-brokered model access, canonical patch export, protected paths, artifacts, watchdog, origin-driven session cleanup, native TTL backstop, and cleanup evidence.
- Migrate direct CLI model-provider implementations behind runtime adapters and disable their shared-service execution path.
Phase 6: Multi-Agent Service Pooling
- Resolve an authorized agent definition per session/turn.
- Cache immutable definitions/catalogs by host, ID, version, and digest.
- Route provider/data-boundary profiles without sharing incompatible secrets.
- Scale replicas with durable session admission rather than in-memory affinity.
Phase 7: Personal Assistant Profile
- Add light-agent-channel, channel/principal binding, webhook verification, delivery idempotency, and proactive trigger policy.
- Add typed connector tools and optional dedicated personal edge runners.
- Add quiet hours, notification policy, scheduled-turn admission, and connector credential brokering.
Phase 8: Workflow Bridge And Fixed High-Value Actions
-
Keep existing call.agent behavior as native-workflow by default.
-
Add an explicit agent-service job mode with schema, deadline, idempotency, cancellation, correlation, and delegation-depth controls.
-
Expose workflow start/status/cancel as typed agent tools.
-
Add accepted-patch, branch/PR, publish, sign, and deploy contracts.
-
Require immutable input, approval, provenance, and fresh action-scoped credentials.
-
Consume every approval into a new numbered domain action and common execution attempt with a fresh lease and fencing token.
-
Reuse the trusted fresh-checkout and protected-path design.
-
Rebuild releases from reviewed immutable commits.
Acceptance And Failure-Injection Tests
- A valid session cannot be resumed by another principal, agent, host, or tenant.
- Concurrent prompts from one or more replicas receive durable FIFO sequence; one turn becomes active and the rest remain queued in order.
- Queue-full admission is bounded and retryable; it does not drop or silently execute a prompt.
- Duplicate client delivery resolves to the existing turn and cannot create a second action.
- A model-returned hidden or unadvertised tool is rejected before gateway dispatch.
- A gateway-only tool absent from gateway
tools/listis hidden without removing an independently authorized runner tool; a runner tool absent from the lease/runtime/local manifest is hidden even if the gateway has the same name. - Model-facing alias collisions across placements fail closed, and a returned tool call cannot switch its snapshotted gateway/runner/workflow/fixed-service dispatch route.
- Malformed or schema-invalid arguments never become an empty object.
- Gateway authorization remains effective even when the catalog is stale.
- Tool and memory output cannot exceed context limits or become system instructions.
- A known recoverable command failure returns to the model within correction budgets; it does not automatically fail the turn.
- An unknown action is reconciled and cannot be repeated or described as a definite failure.
- A successful effect followed by a history-version conflict remains present in agent_action_attempt_t and the session event stream; projection recovery never redispatches it.
- Gateway calls use a signed token narrowed to the caller, agent, turn/action, tool, data boundary, policy digest, audience, and expiry; the original broad user token is not forwarded.
- Service restart during model-only work follows the configured retry policy.
- Restart after effectful dispatch reconciles instead of blindly repeating.
- Terminal result committed while light-agent is offline is found by indexed catch-up and accepted once; dropped, duplicate, and reordered notifications do not change correctness.
- WAITING_APPROVAL holds no active action lease, model-broker channel, or action credential. A task sandbox is cleaned; an eligible non-secret session workspace is retained only through a separately persisted bounded hold or verified checkpoint.
- Missing an action lease does not clean an approval-held session, while hold expiry, logical session close/revocation, or policy mismatch does. Approval cannot extend the session maximum lifetime.
- Approval creates a fresh post-approval action/common execution attempt and fencing token; the pre-approval attempt, handle, and grant remain unusable.
- CLI provider processes receive only an allowlisted environment inside a sandbox.
- Generated code cannot read a provider/proxy bearer, inspect or inherit the worker’s model-broker channel, impersonate another attempt, choose an unauthorized model, or exceed broker-enforced token/cost limits.
- A CLI/provider timeout kills the process tree and triggers cleanup.
- A shared light-agent process cannot directly start an external agent runtime.
- Worker runtime events are ordered, resumable, bounded, and cannot mutate an agent turn without origin-side acceptance.
- A runtime adapter cannot claim an unapproved capability or select a weaker sandbox than the turn policy requires.
- A coding turn receives only the selected immutable skill packages and workspace-local instructions cannot grant additional authority.
- Package download, digest/signature mismatch, unsafe archive entries, or staging failure occurs before sandbox start; sandbox code has no artifact-store credential or package-download egress.
- A generated skill remains inactive until it is scanned, reviewed, packaged, and assigned.
- A sandboxed turn survives controller and runner reconnect without accepting a stale fencing token.
- A session-scoped workspace cannot be reused across principal, agent, base, policy, backend, or expiry changes.
- Closing, revoking, or expiring an agent session durably fences active work and reclaims its physical sandbox without waiting for backend-native TTL; cleanup survives controller and runner restart.
- Toolbx and ordinary containers cannot satisfy a microVM requirement.
- A spoofed channel user, replayed webhook, duplicate delivery, or scheduled trigger cannot create an unauthorized or duplicate turn.
- A personal channel gateway has no model-provider key, unrestricted shell, or tenant-wide connector credential.
- A native-workflow agent call remains backward compatible and a service-mode call cannot cause an unbounded agent/workflow delegation cycle.
- Publish/sign/deploy cannot be invoked as a free-form agent tool.
- Secret scanning finds no credential in prompts, history, memory, logs, artifacts, journal, or child environment.
Open Decisions
- Whether the first production deployment remains one agent definition per service or introduces request-time multi-agent pooling immediately.
- Which session-scoped backends support safe checkpoint/restore.
- Which structured coding adapter is enabled first after the mock and native adapters, and which versioned SDK/RPC contract is pinned.
- Which runner-owned model broker and protected local transport are supported first on each backend: preconnected descriptor, peer-checked Unix-domain socket, vsock, or backend-native equivalent.
- Which object store, signer, scanner, and review service own immutable skill packages.
- Which channel bindings and personal connector grants remain in GenAI domain tables versus a dedicated channel/identity service.
- Which exact workflow syntax selects
agent-servicewhile preserving the existing native-workflow default. - Which model providers support useful request idempotency.
- Whether client disconnect defaults to cancel or durable continuation.
- Which memory read service replaces direct PostgreSQL recall.
- Which agent actions require workflow-owned approval versus standalone agent-owned approval.
Recommendation
Keep light-agent as the interactive session, durable turn, policy, memory, and model orchestration service for all profiles. Do not assign infrastructure per agent definition or fork separate enterprise, coding, and personal-assistant engines. Deploy service pools by real trust and data boundaries.
Use the shared controller/runner/ExecutionBackend path whenever an agent needs
local execution or stronger isolation. Host workspace-aware loops in the small
sandbox-side light-agent-worker; host messaging connections in the separate
light-agent-channel; keep external agent products behind runtime adapters.
Make the runner protocol origin-neutral so workflows and standalone agent turns
share capacity, fencing, backend lifecycle, watchdog, credentials, artifacts,
and cleanup without sharing domain ownership.
LLM Gateway
Status: design proposal, 2026-09-19. Nothing below is implemented. It records the intended direction for the LLM data plane as agent traffic grows beyond single request/response routing, and the boundary decisions that follow. The current implementation boundary is recorded at the end of this document; do not read any section here as a description of shipped behaviour.
This document covers the LLM data plane shared by every agent profile. The per-feature agent lifecycle lives in Light-Agent Execution, and durable business process state lives in Light-Workflow Runner.
Decision And Scope
Agents address one normalized inference interface. The gateway owns provider adaptation, routing, admission, usage accounting, and audit. It does not own the semantic content of an agent’s conversation.
Two deployment profiles are recognized as first-class, not as a temporary split. In the enterprise profile the gateway is inline on the inference path. In the personal/subscription profile the model call egresses directly to the provider and the gateway is a control plane only. The same agent binary must work in both; the only difference is whether the model call passes through the gateway.
Context caching, compression, and compaction are in scope as derived optimizations over content the agent supplied. Authoritative ownership of conversation state is out of scope for third-party agents, and optional for first-party agents.
Ownership
| Component | Responsibility |
|---|---|
| Agent (any profile) | Decide what enters its context, which tools to call, when to compact |
llm-gateway crate | Provider adaptation, alias routing, admission bounds, canonical replay, usage, PII, audit |
| Gateway compression layer | Deterministic, structural reduction of newly arriving content |
light-workflow | Durable process state: which step, which agent, which model, new-or-resume session |
| Model provider | Inference, prefix cache, provider-side reasoning state |
The gateway never decides what is semantically safe to drop from a conversation. It cannot see task intent, and a silent semantic edit surfaces later as an unexplained agent failure with no traceable cause.
What A Standard Interface Can And Cannot Do
A normalized wire interface is achievable and worth building. Model portability is not achievable at the wire level, and the design must not assume it.
When a coding agent degrades against a different provider’s model, the cause
is usually post-training, not protocol shape. Frontier models are trained
against specific tool schemas, specific edit formats (str_replace versus
whole-file versus unified diff), specific shell interaction patterns, and
specific reasoning-token semantics. Normalizing the envelope does not
normalize any of that away.
The useful consequence is that the per-model adaptation layer belongs in the
gateway: tool-schema dialects, prompt scaffolding, preferred edit format,
reasoning handling, and cache-breakpoint placement. The agent addresses one
interface; the gateway owns what each model needs in order to perform well.
crates/model-provider already carries per-provider adapters along these
lines, including CLI-shaped providers.
An agent must therefore be able to ask the gateway for its adaptation profile without issuing an inference request, because in the subscription profile the agent applies that profile itself before calling the provider directly.
Context Ownership
The Stateless API Reality
The Anthropic Messages API and OpenAI Chat Completions are stateless: the client sends the complete message array on every turn. An agent that “manages context” is deciding what to place in that array, not holding a server-side session. There is consequently no authoritative conversation on the gateway that could drift out of sync with the agent’s — the agent restates the whole conversation on every request.
OpenAI’s Responses API with previous_response_id is the exception. Treat any
provider-side conversation state as pass-through; do not attempt to mirror it.
Derived State, Not Synchronized State
Derived state cannot desynchronize; authoritative state can. Whatever the gateway retains about a conversation must be a pure function of what the client sent, keyed by content digest. If a retained entry ever disagrees with the incoming request, it is discarded and recomputed. No synchronization protocol is required, or permitted.
| Mode | Gateway holds | Works with | Cost |
|---|---|---|---|
| Derived (default) | Content-digest-keyed compression cache | Any client, including Codex CLI and Claude Code, unmodified | Client resends full context each turn |
| Authoritative (optional) | The conversation itself, per session | First-party agents only | Gateway becomes stateful; session affinity, failover, and blast radius follow |
The derived mode is the one that covers the fleet and is therefore the one that ships first. The authoritative mode buys simpler first-party agent code and fewer bytes on the agent-to-gateway hop; it does not buy anything the derived mode cannot already deliver on the inference path, and it converts a horizontally scalable stateless service into a sticky one. Adopt it only for first-party agents, and only after the derived path is qualified.
Compress On Entry, Then Freeze
Provider prefix caching requires the cached prefix to be byte-identical across requests. Rewriting earlier turns to save tokens is the single most cache-destructive operation available, so token reduction and cache-hit rate are in tension unless reduction happens exactly once, at the point new content first enters the context.
The rule is: compress a tool result the first time it arrives, then never touch it again. Tool results — file contents, test output, build logs, search dumps — carry nearly all of the reducible volume in an agent loop, and each one is appended exactly once.
Three requirements follow, and the second is a performance trap rather than a correctness one:
- The compressor is a pure deterministic function, versioned, and pinned per session. Upgrading a compressor mid-session invalidates every downstream cache entry and mutates history under the model.
- The mapping from original content digest to compressed bytes is cached. A naive implementation recompresses the entire accumulated history on every request; by turn eighty of a coding session that is significant latency for no benefit. This cache is what makes the derived mode viable, not merely correct.
- Anthropic
cache_controlbreakpoints are placed by the client. Rewriting content beneath them can strand a breakpoint mid-prefix and lose the cache the rewrite was meant to protect. Breakpoints must be re-placed as part of the same deterministic transformation.
Accept one consequence deliberately: after a rewrite, the agent’s record of what it sent no longer matches what the model saw. This is usually harmless, but it invalidates agent-side token accounting and any client-side caching the agent performs.
Sealing Instead Of Storing
Where the gateway must carry state across turns without becoming stateful,
prefer the sealed-envelope pattern already used by reasoning_seal.rs:
authenticated, encrypted state bound to tenant, alias, and route, returned to
the client and replayed on the next request. This keeps the data plane
horizontally scalable and gives the same recovery properties as a stateless
service. A compression dictionary version, a compaction generation, or a
context digest chain can travel this way.
Compaction Is Agent-Initiated
Deterministic, structural reduction is a gateway concern: it is task-agnostic, reversible in principle, debuggable, and safe. Semantic reduction — deciding that a file read forty turns ago no longer matters — is an agent concern, because only the agent knows the task.
A background process that rewrites live context with a cheap model is explicitly rejected. It is lossy, it is invisible at failure time, and it destroys the prefix cache it is nominally protecting.
The supported flow is a signal, not an action: the gateway reports context pressure and may propose a compaction; the agent decides and issues it. A compaction is then a deliberate new prefix whose cache cost is amortized intentionally, rather than a random invalidation.
Two Levels Of Context
Process context and conversation context are separate, with separate owners and lifetimes.
flowchart TD
W[light-workflow: durable process context] --> S1[Step: agent A, model M, resume session X]
W --> S2[Step: agent B, model N, new session Y]
S1 --> G[llm-gateway: per-session conversation]
S2 --> G
G --> P[Provider inference and prefix cache]
light-workflow owns which step is active, which agent and model serve it,
and whether the step resumes an existing session or starts a new one. It does
not observe individual model interactions.
Session identity is workflow-issued, not gateway-generated, so the workflow can resume, fence, and cancel deterministically, consistent with the claim and fencing semantics in Light-Workflow Runner.
A session must be reconstructable from durable inputs. If the gateway loses conversation state, the workflow replays accepted artifacts to rebuild it; a lost cache degrades cost and latency, never correctness. This is the same principle applied to phase handoffs in the development workflow design, where a new stage reconstructs from accepted artifacts rather than depending on a surviving conversation.
Transport
Retain request/response HTTP with a session identifier, and SSE for streaming. Do not adopt WebSocket for the inference path.
A persistent connection is not required for server-side context retention; a session header achieves the same thing, and HTTP/2 already keeps connections warm. WebSocket earns its cost only when the server must initiate messages.
The cost is specific to this platform. Authorization is per-request today — dual identity, scope tokens, and action authorization, qualified as part of the workflow authorization work. Moving the inference path onto long-lived connections means re-implementing authorization at the message level inside a connection, plus reconnect and resume semantics, and it complicates load balancing and per-request tracing. That is a meaningful security surface in exchange for a capability a header already provides.
Revisit only if server-initiated push becomes a real requirement.
Deployment Profiles
| Enterprise (API key) | Personal (subscription) | |
|---|---|---|
| Inference path | Through the gateway | Direct to provider |
| Compression, caching, PII, audit WAL | Gateway, inline | Not available inline |
| Routing decision | Gateway | Gateway |
| Adaptation profile | Gateway applies | Gateway supplies, agent applies |
| Usage accounting | Observed | Reconciled from provider telemetry |
| Authorization and policy | Gateway and workflow | Gateway and workflow |
Subscription credentials authenticate a consumer plan and are bound to the provider’s own endpoint. Routing that traffic through a proxy is a licensing and account-standing question, not merely a technical one, and it is not a posture this platform will adopt on a customer’s behalf. API-key traffic carries no such constraint.
This is not a concession. Unattended, automated workflow traffic is precisely what consumer subscriptions exclude, and the personal pilot design already records that personal subscriptions do not authorize Portal or API access. The enterprise product is the API-key path, where the gateway is inline by construction. Subscriptions serve the interactive personal profile.
The routing decision remains centralized even where the data path is not: the gateway and workflow decide that an interactive job runs on a developer’s own subscription seat while an unattended job runs API-key through the gateway. Centralized decision with a bypassing data path is a supported outcome, not a degraded one.
Packaging
Build a separate llm-gateway binary composing the shared crates. Do not
continue growing the LLM data plane as middleware inside light-gateway.
The LLM data plane and the API data plane have diverging runtime shapes. The API gateway is stateless and request-scoped. The LLM path already holds long-lived streams, large per-request bodies, replay buffers, an embedding admission lane with multi-gigabyte ingress bounds, and its own audit WAL; adding context retention makes it durably stateful. Capacity for the two scales on different axes — requests per second against concurrent sessions and token volume — and co-tenancy means planning for the worse of both. Memory pressure or an OOM on the LLM path should not remove API routing.
The split is cheap today because the crate boundary already exists and is
already thin: crates/llm-gateway carries the runtime, streaming, routing,
usage, and audit surfaces, while apps/light-gateway references
llm_gateway:: in a small number of places. A separate binary is largely a
new composition root over light-pingora, llm-gateway, and config-loader,
not a refactor. That apps/light-gateway/src/main.rs has grown past half a
megabyte is a further argument against adding responsibility to it.
The real risk of splitting is divergent security middleware, and it is manageable only if handled deliberately: the shared security and authorization chain must live in a crate that both binaries compose, never reimplemented in the new one. The documented handler order — correlation, unified-security, limit, access-control, then the LLM handler — must remain a single shared definition. If that chain cannot be factored cleanly, treat the difficulty as a signal to defer the split rather than to fork the middleware.
Trigger: split when the first durably stateful feature lands, which is the session context store described above. That is the point where the runtime profiles genuinely diverge.
Current Implementation Boundary
Checked 2026-09-19 against the working tree.
| Capability | Current boundary |
|---|---|
| LLM data plane | crates/llm-gateway implements provider dispatch, alias routing with multi-attempt fallback, canonical request replay, admission bounds, streaming, usage, PII, and an audit WAL |
| Deployment | One binary: apps/light-gateway composes the crate as an application handler behind the documented chain; llm-router.yml is inert unless explicitly enabled |
| Provider adaptation | crates/model-provider carries per-provider adapters, including CLI-shaped providers; mixed-format aliases use strictest-wins parsing with an OpenAI extension allowlist |
| Conversation state | None. The data plane is stateless; session_id appears only in receipt records. Context retention, compression, and compaction described above are unimplemented |
| Reasoning state | reasoning_seal.rs seals provider reasoning state into an authenticated encrypted envelope bound to tenant, alias, and route |
| Transport | HTTP request/response with SSE streaming; no WebSocket inference path |
| Profiles | Not modelled. The enterprise/subscription split above is a proposal, not a configuration surface |
| Separate binary | Not started |
Open Questions
- Whether the adaptation profile is served as a typed API for subscription agents, or compiled into the agent from the same immutable snapshot.
- Whether compression dictionaries are per-tenant or global, and how a dictionary version is pinned across a long session without a stateful store.
- How usage reconciliation from provider telemetry is attributed to a workflow step when the gateway never observed the call.
- Whether first-party authoritative sessions justify their operational cost at all, once derived-mode compression is qualified.
References
- Light-Agent Execution
- Agent Engine Pattern
- Light-Workflow Runner
- MCP Router
- Handler Chain
- PII Tokenization
- User, Application, And Workflow Authorization
- Personal Development Workflow Orchestration
Light CLI
Status: implemented in apps/light-cli (binary light). Revised 2026-09-22: the CLI is a public
client. It has no certificate and no enrollment: it signs the user in with the OAuth device grant
and calls the Gateway with the user’s own access token alone. It still carries the dev application
token (below), a public identifier that config-server and, later, controller-rs ask for. The sections that describe a per-install certificate, the bootstrap token,
the issuer, renewal and the dual-identity contract for the CLI (Identity Model, The Peer Pinning
Problem, Enrollment And Revocation, and the flow in Bootstrap And Connection Flow) record the design
this replaced; they still describe how a controlled workload (an agent, a runner) is enrolled, and
are kept for that and for history. See Revision 2026-09-22.
This document covers a first-party command-line client for workflow
interaction: starting stages, answering human-in-the-loop decisions, and
operator recovery. The web surface in portal-view remains the primary
interface for enterprise approvers; see Interfaces And Personas
for the boundary between them.
Revision 2026-09-22: A Public Client
The first design gave every install a certificate from light-identity-issuer, enrolled once with
an app token packaged in the download, and had the Gateway trust the CLI by CA. The reasoning was
that the application leg tells the Gateway which program is calling. That does not survive an
open-source, downloadable client:
- The enrolling credential ships in a download anyone can take, so anyone can enroll. A certificate proves that some program enrolled, not that it is the real CLI. It is friction, not a defence, and the same is true of the app token itself.
- OAuth’s guidance for native apps (RFC 8252) is the same: they cannot keep secrets, so they are
public clients acting for a user. That is how
portal-view’s browser already works: it holds no application credential either. - The certificate cost a CA, a bootstrap secret, Gateway trust configuration and a renewal protocol, for no security the user’s own token does not already provide.
So the CLI carries only the user’s login. What it may do is decided by that user’s roles and the route’s ACL, never by a claim about which program it is. Consequences:
- CLI: no
bootstraporrenew, and no key or certificate store. Exit code 4 (“not enrolled”) is retired. It keeps the long-lived dev application token instart-cli.sh(checked in on purpose, so a checkout works out of the box). That token identifies the program, is public by nature (it ships in a download) and authorises nothing on the user’s behalf. It is what the services that ask “which application is this” are given: the config server today (it reads the CLI’s settings with it), and controller-rs when the CLI registers. It is never sent to the Gateway or to light-oauth, and the tests check that. - Gateway:
/mcpwith workflow actions enabled used to require an app token, a verified peer and a user token from every caller.RoutePolicy.interactiveUserOnly(off by default) admits a caller that presents no application credential on its user token alone, as an interactive caller with no action reference. A caller that does present one is judged exactly as before, and a request that claims a workflow action without one is refused. Seedual_identity::admit. - light-identity-issuer stays, for agents and runners deployed in a controlled pipeline, where the delivery of the credential can be trusted.
Revision 2026-09-23: One Persistent Session
light is one long-lived terminal, like Claude Code, that stays open until /exit. The one-shot
subcommands (light auth login, light gateway check, …) are gone; every capability is a slash
command inside the session, and start-cli.sh just starts it. The reason is what the CLI is for: a
workflow reaches a human-in-the-loop step and the person should see it and answer it where they
already are, chat with an agent across many turns, and look at status and audit without a fresh
process (and a fresh login check) per question.
- Commands (implemented):
/login,/logout,/whoami,/agents,/chat [agent],/new,/disconnect,/tools,/help,/exit. Anything that is not a command is said to the agent you are chatting with;//xsays a message that begins with a slash. The prompt shows the agent. - Same engine three ways. The terminal (line editor, history, replies printed above the line
being typed), piped input, and
light -c '<line>'all drive oneShell, so a script behaves like a person and a test drives it line by line. In a script, a message waits for the agent’s reply before the next line runs, so/exitcannot cut it off. Exit code is that of the first failing line;/whoamiwhen signed out counts (6), solight -c /whoamistill answers “am I signed in”. - Not a TUI. No full-screen layout, panes or mouse. It is a scrolling terminal with a prompt; ordinary terminal features (scrollback, copy, pipes) keep working.
--jsonis gone with the subcommands. Structured output comes back per command when a command needs it for scripting; nothing asks for it yet.- Agent chat (
chat.rs) uses the Gateway’s/chatWebSocket as a native client: onlyAuthorization: Bearer <user token>, none of what the Gateway takes for a browser (Origin, the CSRF cookie or subprotocol), and never the token in the URL. The frames are the onesportal-view’s chat page uses. When the login’s access token expires (authentication_requiredor close code 4401) it fetches a fresh token and reconnects with the samesessionId. A message the server may have accepted is never sent again automatically: the person is told whether it was refused (send it again) or may have arrived (check before repeating). Reconnects are bounded. Text from an agent is untrusted and is stripped of terminal control sequences (ESC, C1, bidi and zero-width characters) before it is printed. - Approvals from the CLI are allowed, as the same audited operation the web UI performs, and each one will require explicit confirmation naming what is being approved. Not built yet.
- Agents come from
cli.agentServiceIds(comma-separated service ids; the config server can set it). Discovery from the Portal’s instance registry is later, and needs/portal/queryto accept a bearer token on the Gateway.
Next, in order: the human-task inbox (list, claim, choose an option, comment, complete, with confirmation), workflow status and audit views, log follow, then controller registration.
Decision And Scope
Provide a CLI, not a TUI. The CLI connects to light-gateway over the same
MCP/action surface portal-view already uses, never directly to
light-workflow. It authenticates as the signed-in user with a user access token
(see the revision above; the original design also had a per-install certificate for an
application leg, which was dropped).
The CLI carries no authority of its own. It acts as the signed-in user with that user’s existing grants. It never fabricates, auto-enrolls, or escalates a grant, consistent with the decision already recorded for the Portal enrollment prompt.
Out of scope: a full-screen TUI (the interactive session described in the 2026-09-23 revision is a scrolling prompt, not a TUI); replacing the worklist for enterprise approvers; direct access to Workflow’s internal routes; any long-lived static application secret distributed with the binary.
Interfaces And Personas
| Persona | Primary surface | Rationale |
|---|---|---|
| Enterprise approver | Web (portal-view) | SSO, audit trail, mobile, no install; episodic decisions |
| Developer in the coding loop | CLI | Already in a terminal on the pilot VM; context switch to a browser is the friction |
| Platform operator | CLI | Recovery and diagnosis must be scriptable and diffable |
| Automation and gates | CLI | Daily qualification gates cannot depend on an operator driving dialogs |
The CLI is additive. Where both surfaces expose an action they must call the same Gateway operation, so authorization and audit cannot drift between them.
Why Gateway, Not Workflow
flowchart LR
CLI[light CLI] -- mTLS + user token --> G[light-gateway]
WEB[portal-view] -- session + server ingress --> G
G -- A2 profile, mTLS --> W[light-workflow]
G -. CEL endpoint rules, Tool ACL, default deny .-> G
light-gateway is already the single authorization chokepoint: CEL endpoint
rules, default-deny access control, Tool ACL, and the A2 workflow-actions
profile that reaches Workflow over verified mTLS. Pointing the CLI there means
one authorization implementation, one audit path, and no second copy of the
policy chain to keep in sync.
Connecting directly to Workflow’s listener would bypass that chain, couple the CLI to internal route shapes, and require Workflow to grow a second caller contract. Both are rejected.
This also matches the commercial boundary: the CLI and Gateway are open source
in light-fabric, while the policy that governs them — tenancy, grants,
access-control snapshots, audit retention — is published by the commercial
control plane. The client is free; the governed backend is the product.
Identity Model
dual_identity::authenticate requires, and this design does not weaken:
- a verified TLS peer fingerprint taken from transport context only, never a header or forwarded assertion, failing closed when absent;
- a registered application service ID whose profile lists approved peer fingerprints;
- a user bearer token validated against the route policy issuer and audience.
The CLI therefore needs two credentials, and neither may be embedded in a distributed binary:
| Leg | Credential | Obtained by | Lifetime |
|---|---|---|---|
| Application (“what”) | Per-install client certificate | Enrollment, once per install | Medium, renewable, revocable |
| User (“who”) | Access token | Pairing or device grant, per user | Short, refreshed under proof of possession |
The install certificate identifies the installation, not the person. The user token identifies the person. Neither alone is sufficient, which is the property that makes a stolen token or a copied binary insufficient on its own.
Registered origin is Interactive, alongside the existing Portal ingress
identity — not Workflow, which additionally requires an action reference and
is reserved for service-origin callers.
The Peer Pinning Problem
This is the decision the CLI forces, and it should be settled before implementation.
AppProfile.peer_sha256 is a list of exact SHA-256 leaf fingerprints,
validated as 64 hex characters and required to be non-empty. That model is
sound for a small set of long-lived services. It does not scale to a CLI: every
developer installation produces a new leaf, so every install would append an
entry to a published access-control snapshot, and every revocation would
require republishing it.
Three options:
- CA-based peer trust (recommended). Extend
AppProfilewith a variant that trusts an enrollment CA and requires a verified certificate attribute binding the leaf to the CLI service ID, plus revocation checking. Peer identity stays transport-derived and fail-closed; only the matching rule changes from “this exact leaf” to “issued by this CA for this service ID.” - Enrollment-managed fingerprint list (interim). Keep leaf pinning, and have enrollment append the new fingerprint to a CLI-specific service ID with a bounded list size and expiry. Acceptable for the single-VM pilot, operationally unacceptable at fleet scale.
- Brokered ingress. A server-side component holds one application
credential on behalf of many CLIs, as
portal-view’s Vite ingress does today. This relocates the problem rather than solving it, adds a hop, and creates a trusted intermediary that can act for any user.
Recommend (1), with (2) as an explicitly time-boxed interim for the pilot. Do not adopt (3) for the CLI; the browser needed it because a browser cannot hold a client certificate, and a CLI can.
Obtaining The User Leg Without A Browser Redirect
Decided 2026-09-21: the RFC 8628 device authorization grant, served by
light-oauth. The full design, the security analysis and the protocol are in
light-portal-doc/src/design/light-oauth/device-authorization.md; this section
records only what the CLI does with it.
The earlier options are superseded. A loopback redirect does not work on a VM without a browser; a pasted cookie or token does not either, because the browser session’s tokens are short-lived and the CLI cannot renew them; and Portal-initiated pairing needs a new code type. The device grant is the standard answer, covers cloud and enterprise sign-in (including SSO and MFA, which happen in the user’s own browser), and needs no browser on the CLI’s machine.
The CLI is a public client (no secret) of light-oauth, reached through the Gateway
like everything else; the standard grant needs nothing else. It also presents its install
certificate on those requests, so a Gateway that wants only enrolled installs to be able to
start a sign-in can require it. The user approves on portal-view’s /device page. Signing
in lasts one day, or 90 days if the user ticks “remember me”; both lengths are light-oauth
server settings, the end is enforced by the server, and it does not slide. The CLI never prompts for a password, never opens a browser unless asked
(--open, and never over SSH), and never accepts a token typed by the user. light auth login shows a code; the user approves it in any browser.
Enrollment And Revocation
Enrollment mirrors the runner enrollment precedent in controller-rs, which
already issues a durable per-instance identity bound to an approved peer.
- The CLI generates a key pair locally and produces a CSR. The private key never leaves the host.
- Pairing (or device grant) authenticates the user leg and authorizes enrollment for that user and Host.
- The broker issues a certificate bound to an install ID, user, and Host, and registers it under the CLI service ID according to the peer-trust decision above.
- The CLI stores the install ID and credentials, and obtains short-lived access tokens thereafter.
Refresh is bound to the install certificate as proof of possession, so a captured refresh token alone cannot mint access tokens. The issuer records token fingerprints rather than bearer values, consistent with existing behaviour. Portal lists a user’s enrolled installs with last-seen data and can revoke one without disturbing the others; revocation must take effect at the Gateway without requiring a client to cooperate.
Credential storage uses the OS keychain where available and a 0600 file
fallback otherwise. Credentials are never logged, never placed in argv, and
never emitted in diagnostics.
Bootstrap And Connection Flow
Owner-specified 2026-09-20. The code facts below were checked by reading the tree on that date; none of the flow has been run end to end.
What the CLI connects to
Four services: config-server, controller-rs, light-identity-issuer and
light-gateway. It never connects to light-workflow or light-agent
directly; those are reached through the Gateway.
Distribution
The CLI is downloaded per environment and per version. A download from
dev.lightapi.net ships a default startup.yml with the dev env tag and that
domain; a download from lightapi.net ships the production env tag and
domain. Each release is a new download. The package carries the environment,
the endpoints, and a bootstrap CA bundle (the CAs that verify config-server,
the Controller, the issuer and the Gateway’s server certificate).
How the bootstrap credential reaches startup.yml is not settled: one token
shared by every download of a version, or a token minted per download for a
signed-in user. See Open Questions.
Sequence
- Install. The user downloads and unpacks the package described above.
- First start: configuration and registration. The CLI authenticates to
config-serverwith the bootstrap credential over TLS verified by the bootstrap CA bundle, and downloadsvalues.yml. That file carries the URL of thelight-identity-issuerfor this environment. The CLI also registers withcontroller-rs. Neither call needs a client certificate. - Enrollment. The CLI generates a key pair locally and sends the issuer a
CSR; the private key never leaves the host.
POST /v1/csrcarries the env tag, the bootstrap credential and the CSR. The issuer returns a leaf certificate and its issuing CA certificate, so the CLI can present the chain (see Gateway handshake). The issuer accepts the bootstrap credential once per install and never again for renewal. The issuer does not issue a “server certificate”: the CLI is a client, and verifies servers with its CA bundle. - User leg. The user signs in to Portal and gives the CLI a user access
token, pasted into the CLI. Production access tokens last 10–15 minutes, so
a pasted access token alone means re-pasting every quarter hour; production
needs the refresh token delivered as well, or a one-time code the CLI
redeems for both. Input is read without echo, and the refresh token is stored
in the OS keychain or a
0600file, never in argv, logs or diagnostics. - Connection. Every request to
light-gatewaytravels over mTLS with the enrolled certificate chain, and carries the application token and the user access token. This is the dual-identity contract, unchanged. - Renewal. Runs automatically; see Renewal.
- Recovery. If the key and certificate are both lost, download and reinstall. That only works if each download supplies a fresh bootstrap credential, because the issuer accepts a given one once.
Gateway handshake
The Gateway does not need the client certificate or its fingerprint. With CA trust it needs:
- The issuer’s CA certificate in its incoming client CA file
(
incomingClientCaFile). The Gateway’s TLS listener treats a client certificate as optional (allow_unauthenticated), but verifies any certificate that is presented against that file. A certificate from an unknown CA fails the handshake, and one that is absent is rejected later by the caller policy. - An app profile for the CLI service ID with origin
interactiveand acaTrustentry: the SHA-256 of the issuing CA certificate (issuerSha256), environment, and role. This replaces per-leafpeerSha256pinning for managed workloads. - A client that presents
[leaf, issuing CA]. The Gateway derivesissuer_digestfrom the second certificate in the presented chain (peer_certificates[1]). A leaf-only chain leaves it empty andcaTrustnever matches.
The identity in the certificate is the URI SAN
spiffe://lightapi.local/<role>/<service-id>/<install-id>. That trust domain is
hard-coded on both the Gateway (CaTrust::matches) and the issuer; they agree.
The environment is carried in the issuer-controlled subject O= attribute and must match the
profile’s caTrust.environment. Separate issuing CAs per environment remain recommended for
blast-radius isolation, but environment separation does not rely on that operational choice. The
Gateway’s server certificate is verified by the CLI’s bootstrap CA bundle and is
usually signed by a different CA from the issuer’s.
Renewal
Owner decision: after bootstrap, renewal is authenticated by the old key and certificate only, whether or not the certificate has expired. It replaces both the key and the certificate. It involves no token and no user; purely service-to-service callers renew the same way.
The renewal request carries:
- the env tag;
- the old certificate;
- a new CSR made with a newly generated key;
- a signature made with the old private key over the new CSR digest, the env tag and a fresh timestamp or nonce.
The issuer verifies that the old certificate chains to its CA (ignoring
notAfter), verifies the signature with the old certificate’s public key,
rejects a stale timestamp or reused nonce, takes the identity from the old
certificate (never from the request), and signs the new CSR. The proof is an
application-layer signature rather than mTLS, because an expired certificate
cannot complete a handshake.
CLI behaviour:
- Check on every start. Renew when
now >= renewAt; the issuer computes that deadline asnotAfter - renewLeadSeconds, which includes an already-expired certificate, and renew before connecting to the Gateway. - Keep the old key until the new certificate is safely stored, then swap atomically. A file lock prevents two parallel invocations (for example the daily gate’s scripts) from racing.
- If the certificate is still valid and renewal fails, continue with it. If it has expired and renewal fails, exit with a distinct exit code.
Implementation status. Implemented 2026-09-20 in light-identity-issuer, with
tests, but not deployed: the running issuer is still the older image, whose
/v1/renew accepts only the certificate and a CSR, verifies neither the CA
signature nor any possession signature, and rejects expired certificates. The
renewal request also now carries the proof as a proof object (timestamp,
nonce, base64 signature), and the issuance response carries chainPem,
caCertificatePem, installId, serviceId, role, notAfter, and the
policy-derived renewAt. The first-issuance
request carries no identity: the issuer derives it from the bootstrap credential.
Remaining gaps are tracked in the
issuance implementation plan
section 6.1.
Configuration files
config/startup.yml is the standard file, identical in shape for every
app (host, serviceId, envTag, acceptHeader, timeout, connectTimeout,
configServerUri, authorization, bootstrapCaCertPath). The CLI uses it to know which
environment it is in and which config server to ask. Which config server it asks decides which
instance it belongs to: a CLI downloaded from an instance carries that instance’s startup.yml,
and the settings below (its Gateway, its OAuth provider, and through the provider the sign-in
page) come from that instance’s config server. Everything specific to the CLI lives in
config/cli.yml, a template whose ${cli.<property>:<default>} placeholders resolve, highest
priority first, from an environment variable (CLI_OAUTHPROVIDERID), then the config server’s
values.yml (key cli.oauthProviderId), then the default. If the config server cannot be reached
the defaults apply; that is never fatal. To manage them centrally, create a config named cli in
Portal with these properties, add it to the com.networknt.light-cli-1.0.0 instance, and publish a
snapshot:
| Property | Type | Default | Meaning |
|---|---|---|---|
gatewayUri | string | https://localhost | Base URL of light-gateway; the CLI calls /mcp on it |
oauthUri | string | https://localhost | Base URL of light-oauth as the Gateway exposes it; must be https |
oauthProviderId | string | dev provider id | The provider segment of /oauth2/{providerId}/.... light-oauth puts it in the sign-in link (?provider=), which is how portal-view learns it; nothing in portal-view names a provider |
oauthClientId | string | 01a0bf82-e900-7739-93b0-f33c06db6edb | The device client: Client Profile cli, Client Type public or trusted (the Light CLI’s own client is trusted), set on Portal’s client page (all-in-lt/device-authorization/README.md) |
Implementation status of this flow
Superseded in part 2026-09-22 (see the revision above). Steps 3, 5 and 6 as written here (enrollment,
the certificate connection, renewal) were implemented and verified live on 2026-09-20, then
removed from the CLI. Step 2 (configuration lookup, with the application token) stays. The CLI enrolls nothing, holds no key or certificate, and
light gateway check calls /mcp with the user’s access token alone. Step 4 (the user token) is
light auth login | status | logout, below. Controller registration is not implemented.
light gateway check is tested against a stand-in Gateway that records what it received: the user
token in authorization, no x-scope-token, no client certificate, and nothing from the ignored
authorization line of startup.yml. The real Gateway needs interactiveUserOnly in its
workflow-actions.yml policy (workflow-actions/prepare.py sets it, and prepare-light-cli.py
applies it to a running Gateway); this needs a Gateway build that knows the field, and has not yet
been run live.
The user token: light auth
Since 2026-09-23 these are slash commands in the session: light auth login is /login, auth status
is /whoami, auth logout is /logout, and gateway check is /tools. The behaviour below is unchanged;
the approval code now prints on stdout, with the rest of the session’s output.
Implemented 2026-09-21 in apps/light-cli (auth.rs, oauth.rs, session.rs) and tested
against a stand-in for light-oauth behind a mutual-TLS front (like the Gateway); the server side
is in light-oauth (src/device.rs) and the approval page in portal-view
(src/pages/oauth/DeviceApproval.tsx). Revised 2026-09-22 from a separate mutual-TLS listener
to the standard grant on light-oauth’s one port; see the design in light-portal-doc.
light auth loginrequests a device code, prints the code and the approval address to stderr (stdout stays clean for--json), and polls at the server’s interval, honouringslow_down, until the code is approved, denied or expired. Three consecutive network failures while waiting end the attempt; one or two do not.light auth statusreads only the local session: who is signed in, when the access token and the login end. It exits 6 when nobody is signed in or the login has ended.light auth logoutrevokes the login on the server, then deletes the local tokens. If the server cannot be reached it deletes nothing and says so;--localdeletes the local tokens anyway.- Every Gateway call (
light gateway check) uses the session’s access token, refreshing first if it has under 60 seconds left.LIGHT_USER_ACCESS_TOKENstill wins when set. - Refresh rules. The rotated refresh token is saved (atomic replace, under the store
lock) before the new access token is used, so a crash or a parallel invocation cannot lose
the only valid one; parallel invocations refresh once. A refresh answered
invalid_grantdeletes the local session and exits 6 (“sign in again”). Any other failure, a network error included, leaves the session alone. The login end is absolute: refreshing never extends it. Refresh and logout go to the endpoint recorded in the session, not to whatevercli.ymlsays now. - Storage.
~/.light/<env>/user-session.json, mode0600. The device code is never stored; the CLI never prints a token. - Exit code 6 is new: sign in again. Codes 0, 1, 3, 4 and 5 are unchanged.
Not yet verified live end to end: it needs the Gateway routes and rate limits for the OAuth
paths (/oauth2/*/device_authorization, token, revoke, and the two approval routes) in the
Portal snapshot.
Differences from earlier sections
This section is the owner’s current intent. It conflicts with statements elsewhere in this document, which are left as written until the design is reconciled:
- Enrollment And Revocation has pairing or a device grant authenticate the user and authorize enrollment. In this flow enrollment is authorized by the bootstrap credential, and the user leg is obtained separately in step 4.
- Obtaining The User Leg says neither path may “accept a token typed by the user”. Step 4 has the user paste a token.
- Enrollment And Revocation binds refresh to the install certificate as proof of possession, and says revocation must take effect at the Gateway. Here, refresh handling of the user token is unspecified, and certificate revocation is deferred: user lock-out and refresh-token revocation are a separate layer on top of mTLS, unrelated to certificate renewal.
- Out of scope lists “any long-lived static application secret distributed with the binary”. A package carrying a token shared by every download of a version would contradict that.
Authorization
No new authorization surface. The CLI’s requests traverse the same CEL endpoint rules, default-deny access control, and Tool ACL as the web UI. The effective authority is the intersection of the user’s grants and the CLI service ID’s registration.
Two additional requirements:
- Confirmation and attribution for destructive operations. Cancel-and- release-VM, replan, and supersession require explicit confirmation with the expected feature version and reservation generation, and the audit record carries the install ID alongside the user.
- No implicit grant acquisition. If a required grant is absent the CLI reports what is missing and stops. It does not enroll one on the user’s behalf.
Command Surface
A deliberate subset, shaped by the routes and tools that exist rather than by what a UI can show. Illustrative, not final:
| Area | Commands |
|---|---|
| Auth | auth login, auth logout, auth status (implemented) |
| Features | feature list, feature show, feature start, feature accept, feature replan, feature cancel |
| Decisions | task list, task show, task approve, task reject, task request-changes |
| Runs | run status, run result, run cancel |
| Operator | vm list, vm holder, vm release |
| Evidence | findings list, artifact get, diff show |
Every command supports --json for machine consumption, because the daily
qualification gate is a first-class consumer, not an afterthought. Human
output is the default; JSON is the contract.
Historical implementation boundary (2026-09-19)
Historical boundary checked 2026-09-19 against the working tree; later changes retired the Workflow credential broker. See Workflow Invoke for the current path.
| Capability | Current boundary |
|---|---|
| CLI | None. No command-line framework dependency anywhere in light-fabric |
| Dual identity | crates/light-security/src/dual_identity.rs requires a transport-verified peer fingerprint, a registered service ID with non-empty exact-leaf peer_sha256, and a user bearer token |
| Peer trust | Exact leaf pinning only; no CA-based variant exists |
| Origins | Interactive, Workflow, Gateway, Receiver; Workflow-origin callers additionally require an action reference |
| Interactive caller precedent | portal-view registers a server-side ingress application identity with a pinned peer fingerprint; the browser holds no application credential |
| Credential broker (retired) | The former apps/light-workflow/src/credential_broker_api.rs exposed enroll, complete, and revoke; the routes and listener were removed. |
| Issuer (historical broker role) | portal-service/apps/light-oauth formerly supplied one-time-code acquisition over mTLS for the retired broker. |
| Device grant | Implemented 2026-09-21 in light-oauth (src/device.rs) and the CLI (light auth); see the design in light-portal-doc |
| Enrollment precedent | controller-rs runner enrollment issues durable per-instance identity bound to an approved peer |
Update 2026-09-20, verified by reading the tree, not by running it:
- Peer trust: the CA-based variant now exists as
AppProfile.ca_trust(issuer_sha256plusrole), alongside exact-leafpeer_sha256. No live Gateway policy uses it yet. - Issuer:
light-identity-issuer(crates/light-identity-issuer,apps/light-identity-issuer) exists and runs in the local stack, with several blocking gaps recorded in the issuance plan’s section 6.1.
Open Questions
- Which peer-trust option is adopted, and whether option 2 ships at all or the pilot waits for CA-based trust.
- Whether the CLI binary is named
light,lightctl, or something else, and whether it is a single binary or per-domain subcommands. - Whether approvals carrying legal or compliance weight are permitted from the CLI at all, or must be completed in the web UI for evidentiary reasons.
How install certificate renewal behaves on a host that has been offline past expiry.Resolved 2026-09-20: renewal works after expiry, authenticated by the old key and certificate (see Renewal).- How the bootstrap credential reaches
startup.yml. A token shared by every download of a version cannot satisfy a once-per-install guard, would be a public bootstrap credential, and, if it is also used atconfig-serverandcontroller-rs, would expose those services. A token minted per download for a signed-in user avoids all three and makes “download and reinstall” a working recovery. This also decides whether the Out of scope statement above stands. - How a Portal-issued user token, and in production its refresh token, reaches the CLI: pasted, or redeemed from a one-time code.
- What the CLI needs from
controller-rs. That decides the bootstrap credential’s scope. - How a retired release is cut off. A
sidthat carries the version (com.networknt.light-cli-1.0.0) would let a release be retired by removing its Gateway profile and refusing itssidat the issuer, without a revocation list. This depends on Gateway policy being distributed rather than hand-edited. - Whether
--jsonoutput is versioned as a stability contract once the daily gate depends on it.
References
- User, Application, And Workflow Authorization
- Access Control Handler
- Fine-Grained Authorization
- Unified Security Handler
- MCP Router
- Controller Registry Client
- LLM Gateway
- Personal Development Workflow Orchestration
Workload Identity Issuance
Update 2026-09-22: the Light CLI no longer uses this service. It is open source and downloadable anywhere, so any credential it could be enrolled with is as public as the download; it is a public client that signs the user in and calls the Gateway with the user’s token alone (see Light CLI). The issuer remains for workloads deployed in a controlled pipeline (agents, runners), where the delivery of the credential can be trusted.
Status: design proposal, 2026-09-19. Nothing below is implemented. Today,
per-workload mTLS credentials are produced by hand-run scripts
(prepare.py, prepare-portal-ingress.py in portal-config-loc/all-in-lt)
that generate a fixture CA, mint leaf certificates with multi-year validity,
and require an operator to re-run them and redeploy whenever a credential
expires or a new caller needs to be added. This document proposes replacing
that with an automated, SPIFFE/SPIRE-shaped issuance service. The current
implementation boundary is recorded at the end of this document.
Why This Is Needed Now, Not Later
Two failures in the same operating session motivated this: an ingress
credential minted with a 1-day validity expired and silently broke a caller
with no alert, and a separate A2 activation appeared inert because a
config-file mount and a compose overlay had drifted out of sync with no
automated check that they matched. Neither failure was a design flaw in
crates/light-security/src/dual_identity.rs — the authentication contract
held correctly in both cases, fail-closed as designed. The gap is entirely
operational: nothing renews a credential before it expires, and nothing
mints a new one without a human running a script against a live container.
This is the same problem the Light CLI design already
identified and deferred as “Slice 4: Multi-Install” — exact-leaf pinning in
AppProfile.peer_sha256 does not scale past a handful of long-lived
services. That deferral was correct for a single-install pilot. It stops
being correct as soon as any caller population grows past what a human can
re-mint by hand, which is already true for the current five-service fixture
and would be untenable at 1000 agents or 1000 CLI installs.
Decision And Scope
Introduce a dedicated issuance service that holds CA signing authority, issues short-lived leaf certificates to workloads on request, and expects every workload to renew automatically before expiry using its current, still-valid certificate as proof of possession. This is the SPIFFE/SPIRE model (workload API, rotating SVIDs), adopted in shape: purpose-built rather than an embedded SPIRE server, so VM and Kubernetes deployment are both handled through this issuer’s own bootstrap contract rather than SPIRE’s platform-specific attestation plugins (see Decisions).
Explicitly not config-server’s job. Config-server distributes policy —
RoutePolicy, AppProfile maps, endpoint rules, CA trust bundles — as a
broadcast snapshot every service already pulls at startup. Minting a
private key and signing a certificate is a narrower, per-instance,
authenticated, individually-audited operation with different sensitivity.
Keeping them separate means config-server’s blast radius stays what it is
today (wrong policy served) and does not grow to include “wrong key
signed.” Config-server’s role in this design is limited to distributing the
CA trust bundle and the issuer’s reachable address — both of which are
genuinely policy, not key material.
Out of scope: a general-purpose PKI product; hardware attestation; issuance for anything outside this system’s own service-to-service and CLI-to-Gateway boundary; replacing OAuth bearer tokens for the user leg, which this does not touch.
Architecture
flowchart LR
subgraph Issuer[Identity issuance service]
CA[CA signing key<br/>HSM/KMS-backed]
end
CS[config-server] -- CA trust bundle + issuer address --> W1
CS -- CA trust bundle + issuer address --> W2
W1[Workload: light-gateway] -- CSR + bootstrap token, once --> Issuer
W1 -- CSR + current cert as proof of possession, on renewal --> Issuer
Issuer -- short-lived leaf cert --> W1
W2[Workload: CLI install] -- CSR + pairing-derived token, once --> Issuer
Issuer -- short-lived leaf cert --> W2
W1 -- mTLS, CA-chain trust --> W2
Each workload:
- Generates its own key pair locally. The private key never leaves the
host, matching the existing enrollment precedent in
controller-rs’s runner admission and the CLI design’s enrollment section. - Authenticates to the issuer once, at first bootstrap, using whatever credential establishes “this install is allowed to exist” for its category — an enrollment token for a service, a pairing-derived grant for a CLI install (reusing the pairing flow already proposed in Light CLI).
- Receives a short-lived leaf certificate (hours, not years) chained to the CA, carrying a certificate attribute that binds the leaf to a specific install identity — the CA-based peer trust option already named as recommended in the CLI design.
- Renews automatically, well before expiry, by presenting its still-valid certificate back to the issuer as proof of possession and receiving a new one. No human, no redeploy, no script re-run.
Revocation of one install does not touch any other install’s credential. With short-enough lifetimes, non-renewal is sufficient; a compromised install can additionally be denied at the next renewal via a revocation list the issuer checks, scoped to one install ID.
One Cert Per Instance, Never Shared
At any fleet size — five services today, 1000 agents or 1000 CLI installs tomorrow — every instance holds its own unique leaf, never a cert shared across instances. Sharing collapses two properties this system depends on:
- Revocation granularity. Pulling one compromised install must not require rotating a credential every other install also presents.
- Audit attribution. An action’s audit record must be traceable to the one install that performed it, matching the CLI design’s requirement that “the audit record carries the install ID alongside the user.”
What scales is not the number of trusted leaves but the shape of the
trust rule. Today, AppProfile.peer_sha256 is an explicit, exact-match
list — appropriate for five long-lived services, unworkable for a fleet.
Under this design, the Gateway-side rule changes from “this exact leaf” to
“any certificate chaining to this CA, whose bound attribute matches this
route’s expected service/role” — a rule that does not grow with fleet size,
while the underlying leaf population still grows one-per-instance. This is
the CA-based AppProfile variant already proposed, not weakened, in the
Light CLI design.
What Config-Server Continues To Own
- The CA trust bundle (public certificates only — never a signing key).
- The issuer’s reachable address, so a workload with a live config connection can always find where to renew.
- The
RoutePolicy/AppProfilerole-matching rules that consume the identity this system issues.
This preserves the existing property that a live config-server snapshot is the single source of truth for “what does this service currently trust,” without asking it to also hold or use signing key material.
Relationship To Existing Manual Tooling
prepare.py and prepare-portal-ingress.py are not being thrown away
immediately; they are the fixture this design replaces once the issuer
exists, and remain the right tool for a from-scratch local CA bootstrap in
the interim. The distinction to preserve: those scripts mint credentials by
hand-signing with a locally generated fixture CA key and a DB-extracted
Portal signing key; this design’s issuer holds one CA consistently and
signs on authenticated request, with rotation as a first-class behavior
rather than an operator-run step.
Decisions
These were open questions during design and are now settled.
Build, not adopt SPIRE
Purpose-built, as one Rust crate/service reusing light-security’s existing
verification code. Driving reason: this must support both VM deployment
(the current pilot shape) and Kubernetes deployment on the same issuance
contract, and a purpose-built issuer keeps that a matter of how a workload
obtains its bootstrap credential (below), not a dependency on SPIRE’s own
Kubernetes-specific attestation plugins and operational model.
CA key custody
On disk for now, with the interface shaped so the signing operation can be swapped to AWS KMS or Google Cloud KMS/HSM without changing any workload- or Gateway-facing contract. Concretely: isolate “sign this CSR with the CA key” behind one narrow internal call in the issuer, so the only code that changes when moving to a managed HSM/KMS is that one call’s implementation, not the enrollment flow, the renewal flow, or any consuming service.
Bootstrap credential: reuse the existing long-lived Portal token, once
Every instance today — gateway, workflow, agent, or otherwise — already
obtains a long-lived token from Portal to authenticate to config-server and
Controller, plus a CA bundle (bootstrapCaCertPath in startup.yml) used
only to verify config-server’s own TLS server certificate. There is no
existing client-identity certificate to preserve; the token is the only
credential actually doing authentication today, which makes it a reasonable
“this install is authorized to exist” fact to reuse.
The refinement: use it for first issuance only, not for every renewal. The issuer accepts the long-lived Portal token exactly once per install to mint the first short-lived leaf certificate. Every renewal after that authenticates by proof-of-possession of the previously issued certificate’s private key, not by presenting the token again. This matters because a bearer token is portable — anyone holding it can replay it from anywhere — while a private key that must prove possession is not. Reusing the token repeatedly for renewal would just relocate the CLI’s original bearer-token weakness into this system; reusing it once, to bootstrap a possession-bound credential, does not.
The CLI does not have a pre-provisioned Portal service token the way a container workload does, so its bootstrap path stays what Light CLI already proposed — Portal-initiated pairing — while service workloads bootstrap from their existing long-lived token. Both terminate at the same issuer and the same first-issuance behavior.
Configurable lifetime and renewal lead time, per deployment environment
Made part of the issuer’s own configuration, distributed like any other
environment-scoped value (matching the existing envTag convention —
loc/dev/prod), for example:
portal-config-loc(pilot, VM): 10-day leaf lifetime, renew 1 day before expiry.- Production: 1-day leaf lifetime, renew 1 hour before expiry.
The environment is bound into the certificate. An issuer that serves several environments takes the
environment from the bootstrap token at first issuance and writes it into the certificate’s subject
(O=<envTag>; the URI SAN is unchanged). Gateway caTrust profiles name the expected environment
and match it against this verified subject attribute, in addition to the issuer digest, role, and
service ID. Renewal is refused with
EnvironmentMismatch (HTTP 403) unless the request names the environment in the presented certificate,
so a holder of a production certificate cannot renew into the development policy’s ten-day lifetime by
asking for it. A certificate with no environment in it (issued before this) cannot be renewed and must
enroll again.
The issuer digest is verified, not assumed. caTrust.issuerSha256 is matched against the certificate
after the leaf in the chain the client sends. The client chooses that order and content, so the Gateway
(the pingora-core patch) uses it only when that certificate really signed the leaf: the names match and
the leaf’s signature verifies under its key. Otherwise a leaf from one trusted CA could carry a different
trusted CA’s certificate to satisfy that CA’s profile.
Reconsidered from an initial 1-hour/5-minute proposal: the future production revocation path is specified separately in Service Identity and Revocation, but is not implemented yet. Until Portal and Config Server distribute that policy, expiry is the effective revocation bound. A 1-day lifetime bounds a leaked key’s usable window far tighter than today’s multi-year certs, at roughly 1/24th the renewal volume of an hourly cycle — a meaningful difference at 1000+ instances. The lead time should be read as “renewal must be complete by,” not “start trying at”; the renewal loop should begin attempting well before the 1-hour mark, with backoff and retry, so a transient issuer or network blip has room to resolve before anything actually expires. Tighten these figures later once the renewal loop has proven reliable at scale, rather than starting at the tightest setting.
Sequencing: build the issuer first, CLI is the first consumer
The CLI is a clean slate with no existing credential model to migrate, unlike
Gateway/Workflow/agents, which already have hand-mounted certs from
prepare.py in a working (if manual) state. Proving the issuer against a new
consumer first, then migrating the existing A2 fixture credentials to it
afterward, is lower risk than changing the load-bearing service-to-service
path first.
Open Questions
- Renewal loop robustness at the production 1-day lifetime. Retry policy, backoff, and alerting when a renewal attempt fails, given the smaller margin for error than the pilot’s 10-day/1-day figures.
- Revocation list distribution. Whether the per-install revocation list the issuer checks at renewal is itself distributed through config-server (consistent with “config-server owns policy”) or held only by the issuer.
- Relationship to the CLI’s Slice 4. Slice 4 becomes “adopt this issuer” rather than “build CA-based trust from scratch” — the Light CLI plan should be updated to reflect that once this service exists.
Current Implementation Boundary
Checked 2026-09-19 against the working tree.
| Capability | Current boundary |
|---|---|
| Certificate issuance | Manual, via prepare.py / prepare-portal-ingress.py in portal-config-loc/all-in-lt; no automated issuer exists |
| Peer trust | Exact leaf pinning only (AppProfile.peer_sha256); no CA-based attribute matching |
| Rotation | None; credential lifetime is fixed at mint time and expiry requires a manual re-run |
| Revocation | Republishing the policy snapshot with the entry removed; no per-install revocation list |
| Config-server’s role | Distributes RoutePolicy/AppProfile snapshots today; does not distribute a CA trust bundle as a distinct artifact |
| Bootstrap/enrollment | controller-rs runner admission is the closest existing precedent for issuing a durable per-instance identity bound to an approved peer |
References
- Light CLI
- LLM Gateway
- User, Application, And Workflow Authorization
- Controller Registry Client
- Config Loader
Database Design
The Light-Fabric utilizes a robust PostgreSQL schema to manage the entire lifecycle of agentic workflows, skills, agent execution, channels, and the biomimetic Hindsight memory system. The schema is organized into five logical layers:
This page catalogs the logical domain model and current table shapes. It does not require every table to remain in one physical database or in the Config Server schema. The target authority, database, schema, tenancy, and migration boundaries are defined in Control Plane And Operational Data.
1. Workflow Engine
These tables manage the definition and execution of long-running agentic workflows.
wf_definition_t
Stores the Agentic Workflow DSL (YAML) that defines the high-level orchestration logic.
process_info_t & task_info_t
Manage the runtime state of workflow instances (processes) and individual steps (tasks). They include input_data, context_data, and error_info to provide a resilient “scratchpad” for intermediate variables.
worklist_t & worklist_asst_t
Manage task assignments and visibility for human-in-the-loop interactions.
2. Agentic Core (The “Brain & Skills”)
These tables define the identity, expertise, and capabilities of individual agents.
agent_definition_t
Defines the agent’s persona, product profile, model policy, default execution profile, data boundary, and runtime limits. Enterprise, coding, and personal-assistant definitions share this table; a definition does not own a process or sandbox.
skill_t
Stores the “Expertise” of an agent in Markdown format. Skills are hierarchical and versioned.
tool_t & tool_param_t
The “Hands” of the agent. Defines governed REST/MCP capabilities and typed
execution metadata. A tool row or script field is not authority to execute
code; untrusted executable assets require an immutable reviewed package and an
approved runner sandbox. The target model adds or derives a stable internal
tool reference and records a server-owned gateway, runner, workflow, or
fixed-service placement, model alias, schema digest, effect class, and
dispatch-policy binding. Existing rows with proven current gateway-config
linkage migrate as gateway. Ambiguous legacy script rows remain undisclosed
until reviewed.
agent_skill_t & skill_tool_t
Maps agents to skills and skills to tools, implementing the Progressive Disclosure pattern where agents only see the tools required for their current skill context.
skill_package_t (proposed)
References immutable signed/scanned skill assets for coding, personal, or external runtime adapters. Package bytes live in object storage. The row binds digest, provenance, entrypoint, compatible profiles, required capabilities, review, revocation, and retention.
3. Hindsight Memory System
A biomimetic memory architecture that transitions from flat logs to structured “atoms of thought.”
agent_memory_bank_t
Profiles for memory banks, defining the “Personality and Disposition” (e.g., skepticism, empathy) of the memory layer.
agent_memory_unit_t
The individual “Atoms” of memory. Each unit contains content and a vector embedding (384-dim) for semantic retrieval.
agent_memory_entity_t & agent_memory_link_t
A Knowledge Graph layer that resolves entities and causal/semantic relationships between memory units.
4. Agent Session And Execution
agent_session_t, agent_turn_t, agent_action_attempt_t, and agent_approval_t (proposed)
Store authenticated session ownership, durable FIFO turns, effectful action
attempts, policy snapshots, budgets, bound single-use approvals,
reconciliation, and terminal results. Runner-backed actions reference the
shared execution_attempt_t, but agent-domain fields remain owned by
light-agent. Post-approval dispatch creates a new numbered agent action and a
new common execution attempt; it never reopens the pre-approval attempt.
Shared runner execution records (proposed)
runner_scheduling_request_t, execution_attempt_t, execution_session_t,
execution_session_cleanup_request_t, and execution_input_t store
origin-neutral capacity, fencing, normalized results, bounded session cleanup,
immutable staged inputs, and session state/version/fencing. Reused sessions can
enter a bounded IDLE_APPROVAL_HOLD with hold ID/expiry, cost policy, and
checkpoint/patch evidence while the action lease and credentials are gone.
Zero active attempts does not imply cleanup for a valid held session; the hold
cannot extend idle/max expiry or override close/revocation. Origin services
retain their own workflow or agent state. A terminal-attempt transaction emits
an identifiers-only wakeup;
origins query the authoritative row and use startup/periodic catch-up rather
than treating notification delivery as durable state.
agent_session_event_t (proposed)
An append-only authoritative ledger for accepted user messages, model results, actions, approvals, and terminal turn events. It is the source used to rebuild conversation context after a projection conflict.
agent_session_history_t
The materialized transcript used for active conversation context, linked to the session’s Hindsight memory bank. It is not the authoritative proof of an effectful tool action. Agent action attempts and append-only session events survive history-version conflicts and can rebuild this projection. See Light-Agent Execution.
5. Personal Assistant Channels And Triggers
agent_channel_binding_t and agent_channel_delivery_t (proposed)
Bind a verified messaging identity to a principal and agent, and deduplicate inbound events/outbound responses. They store credential references rather than channel secrets and do not replace agent turns.
agent_trigger_t (proposed)
Stores scheduled or connector-triggered turn policy, including timezone, quiet hours, rate/cost limits, notification destination, idempotency, and maximum delay.
See Light-Agent Execution and the dedicated Light-Agent runtime implementation plan for constraints and phased rollout.
DDL Specification
-- Workflow Definitions: Stores the Agentic Workflow JSON
CREATE TABLE wf_definition_t (
host_id UUID NOT NULL,
wf_def_id UUID NOT NULL,
namespace VARCHAR(126) NOT NULL,
name VARCHAR(126) NOT NULL,
version VARCHAR(20) NOT NULL,
definition TEXT NOT NULL, -- The Agentic Workflow DSL in YAML
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, wf_def_id),
UNIQUE(host_id, namespace, name, version)
);
CREATE TABLE worklist_t (
host_id UUID NOT NULL,
assignee_id VARCHAR(126) NOT NULL,
category_id VARCHAR(126) DEFAULT '(all)' NOT NULL,
status_code VARCHAR(10) DEFAULT 'Active' NOT NULL,
app_id VARCHAR(512) DEFAULT 'global' NOT NULL,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN NOT NULL DEFAULT TRUE,
update_user VARCHAR (255) DEFAULT SESSION_USER NOT NULL,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
PRIMARY KEY(host_id, assignee_id, category_id)
);
CREATE TABLE worklist_column_t (
host_id UUID NOT NULL,
assignee_id VARCHAR(126) NOT NULL,
category_id VARCHAR(126) DEFAULT '(all)' NOT NULL,
sequence_id INTEGER NOT NULL,
column_id VARCHAR(126) NOT NULL,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, assignee_id, category_id, sequence_id),
FOREIGN KEY(host_id, assignee_id, category_id) REFERENCES worklist_t(host_id, assignee_id, category_id) ON DELETE CASCADE
);
CREATE TABLE process_info_t (
host_id UUID NOT NULL,
process_id UUID NOT NULL, -- generated uuid
wf_def_id UUID NOT NULL, -- workflow definition id
wf_instance_id VARCHAR(126) NOT NULL, -- workflow intance id
app_id VARCHAR(512) NOT NULL, -- application id
process_type VARCHAR(126) NOT NULL,
status_code CHAR(1) NOT NULL, -- process status code 'A', 'C'
started_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
ex_trigger_ts TIMESTAMP WITH TIME ZONE NOT NULL,
custom_status_code VARCHAR(126),
completed_ts TIMESTAMP WITH TIME ZONE,
result_code VARCHAR(126),
source_id VARCHAR(126),
branch_code VARCHAR(126),
rr_code VARCHAR(126),
party_id VARCHAR(126),
party_name VARCHAR(126),
counter_party_id VARCHAR(126),
counter_party_name VARCHAR(126),
txn_id VARCHAR(126),
txn_name VARCHAR(126),
product_id VARCHAR(126),
product_name VARCHAR(126),
product_type VARCHAR(126),
group_name VARCHAR(126),
subgroup_name VARCHAR(126),
event_start_ts TIMESTAMP WITH TIME ZONE,
event_end_ts TIMESTAMP WITH TIME ZONE,
event_other_ts TIMESTAMP WITH TIME ZONE,
event_other VARCHAR(126),
risk NUMERIC,
risk_scale INTEGER,
price NUMERIC,
price_scale INTEGER, -- Scale (number of digits to the right of the decimal) of the risk column. NULL implies zero
product_qy NUMERIC,
currency_code CHAR(3),
ex_ref_id VARCHAR(126),
ex_ref_code VARCHAR(126),
product_qy_scale INTEGER,
parent_process_id VARCHAR(22),
deadline_ts TIMESTAMP WITH TIME ZONE,
parent_group_id NUMERIC,
process_subtype_code VARCHAR(126),
owning_group_name VARCHAR(126), -- Name of the group that owns the process
input_data JSONB, -- The initial data that triggered the workflow
context_data JSONB, -- The runtime "scratchpad" for intermediate variables
error_info TEXT, -- Detailed error or stack trace if the process fails
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, process_id),
FOREIGN KEY(host_id, wf_def_id) REFERENCES wf_definition_t(host_id, wf_def_id) ON DELETE CASCADE
);
CREATE TABLE task_info_t
(
host_id UUID NOT NULL,
task_id UUID NOT NULL,
task_type VARCHAR(126) NOT NULL,
process_id UUID NOT NULL,
wf_instance_id VARCHAR(126) NOT NULL,
wf_task_id VARCHAR(126) NOT NULL,
status_code CHAR(1) NOT NULL, -- U, A, C
started_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
locked CHAR(1) NOT NULL,
priority INTEGER NOT NULL,
completed_ts TIMESTAMP WITH TIME ZONE NULL,
completed_user VARCHAR(126) NULL,
result_code VARCHAR(126) NULL,
locking_user VARCHAR(126) NULL,
locking_role VARCHAR(126) NULL,
deadline_ts TIMESTAMP WITH TIME ZONE NULL,
lock_group VARCHAR(126) NULL,
task_input JSONB, -- Specific data passed to the task
task_output JSONB, -- Result returned by the task action
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, task_id),
FOREIGN KEY (host_id, process_id) REFERENCES process_info_t(host_id, process_id) ON DELETE CASCADE
);
CREATE TABLE task_asst_t
(
host_id UUID NOT NULL,
task_asst_id UUID NOT NULL,
task_id UUID NOT NULL,
assigned_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
assignee_id VARCHAR(126) NOT NULL,
reason_code VARCHAR(126) NOT NULL,
unassigned_ts TIMESTAMP WITH TIME ZONE NULL,
unassigned_reason VARCHAR(126) NULL,
category_code VARCHAR(126) NULL,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, task_asst_id),
FOREIGN KEY(host_id, task_id) REFERENCES task_info_t(host_id, task_id) ON DELETE CASCADE
);
CREATE TABLE audit_log_t
(
host_id UUID NOT NULL,
audit_log_id UUID NOT NULL,
source_type_id VARCHAR(126) NULL,
correlation_id VARCHAR(126) NULL,
user_id VARCHAR(126) NULL,
event_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
success CHAR(1) NULL,
message0 VARCHAR(126) NULL,
message1 VARCHAR(126) NULL,
message2 VARCHAR(126) NULL,
message3 VARCHAR(126) NULL,
message VARCHAR(500) NULL,
user_comment VARCHAR(500) NULL,
PRIMARY KEY(host_id, audit_log_id)
);
CREATE INDEX audit_log_idx1 ON audit_log_t (source_type_id, correlation_id, event_ts, user_id);
-- Agent Definitions: Stores the "Brain" configuration
CREATE TABLE agent_definition_t (
host_id UUID NOT NULL,
agent_def_id UUID NOT NULL,
agent_name VARCHAR(126) NOT NULL,
model_provider VARCHAR(64) NOT NULL, -- 'openai', 'anthropic', etc.
model_name VARCHAR(126) NOT NULL, -- 'gpt-4o', 'claude-3-5-sonnet'
api_key_ref VARCHAR(126), -- Reference to Secret Manager key
temperature NUMERIC(3,2) DEFAULT 0.7,
max_tokens INTEGER, -- max number of tokens can be used
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT TRUE,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, agent_def_id),
UNIQUE(host_id, agent_name)
);
-- Skills: Stores Instructions and Domain Knowledge (The "Expertise")
-- Note: Use entity_tag_t and entity_category_t with entity_type = 'skill'
-- for flat tagging and hierarchical folder structure of skills.
CREATE TABLE skill_t (
host_id UUID NOT NULL,
skill_id UUID NOT NULL,
parent_skill_id UUID, -- Self-reference for Hierarchy
name VARCHAR(126) NOT NULL,
description VARCHAR(500), -- High-level description for the initial LLM prompt
content_markdown TEXT NOT NULL, -- The actual instructions/prompts
description_embedding VECTOR(384), -- For semantic lookup/discovery
version VARCHAR(20) DEFAULT '1.0.0',
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, skill_id),
FOREIGN KEY(host_id, parent_skill_id) REFERENCES skill_t(host_id, skill_id)
);
CREATE INDEX idx_skill_active ON skill_t(active);
CREATE INDEX idx_skill_name ON skill_t(name);
-- Tools: Stores Executable Functions (The "Hands")
CREATE TABLE tool_t (
host_id UUID NOT NULL,
tool_id UUID NOT NULL,
name VARCHAR(126) NOT NULL,
description TEXT NOT NULL, -- Instructions for LLM on when/how to use this tool
-- Implementation specifics
implementation_type VARCHAR(50), -- 'java', 'mcp_server', 'rest', 'python', 'javascript'
implementation_class VARCHAR(500), -- FQCN if 'java'
mcp_server_name VARCHAR(126), -- MCP server name if 'mcp_server'
api_endpoint VARCHAR(1024), -- URL if 'rest'
api_method VARCHAR(10), -- HTTP Method if 'rest'
endpoint_id UUID, -- Reference to fine-grained auth endpoint
script_content TEXT, -- Source code if 'python'/'javascript'
response_schema JSONB, -- Strict output schema for tool results
description_embedding VECTOR(384), -- For semantic lookup/discovery
version VARCHAR(20) DEFAULT '1.0.0',
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, tool_id),
FOREIGN KEY(host_id, endpoint_id) REFERENCES api_endpoint_t(host_id, endpoint_id) ON DELETE CASCADE
);
CREATE INDEX idx_tool_host_endpoint ON tool_t(host_id, endpoint_id);
CREATE INDEX idx_tool_active ON tool_t(active);
CREATE INDEX idx_tool_name ON tool_t(name);
-- Tool Parameters: Defines the arguments for each tool
CREATE TABLE tool_param_t (
host_id UUID NOT NULL,
param_id UUID NOT NULL,
tool_id UUID NOT NULL,
name VARCHAR(255) NOT NULL,
param_type VARCHAR(50) NOT NULL, -- 'string', 'number', 'boolean', 'object', 'array'
required BOOLEAN DEFAULT true,
default_value JSONB,
description TEXT, -- Helps LLM understand what value to extract
validation_schema JSONB, -- JSON Schema for complex validation
order_index INTEGER DEFAULT 0,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, param_id),
FOREIGN KEY(host_id, tool_id) REFERENCES tool_t(host_id, tool_id) ON DELETE CASCADE
);
-- Skill Dependencies: Manages hierarchies where one skill requires another
CREATE TABLE skill_dependency_t (
host_id UUID NOT NULL,
skill_id UUID NOT NULL,
depends_on_skill_id UUID NOT NULL,
required BOOLEAN DEFAULT true,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY (host_id, skill_id, depends_on_skill_id),
FOREIGN KEY(host_id, skill_id) REFERENCES skill_t(host_id, skill_id),
FOREIGN KEY(host_id, depends_on_skill_id) REFERENCES skill_t(host_id, skill_id)
);
-- Agent-Skill Mapping: Links Agents to their Skills
CREATE TABLE agent_skill_t (
host_id UUID NOT NULL,
agent_def_id UUID NOT NULL,
skill_id UUID NOT NULL,
config JSONB DEFAULT '{}',
priority INTEGER DEFAULT 0,
sequence_id INTEGER DEFAULT 0, -- Order in which skills are concatenated
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, agent_def_id, skill_id),
FOREIGN KEY(host_id, agent_def_id) REFERENCES agent_definition_t(host_id, agent_def_id) ON DELETE CASCADE,
FOREIGN KEY(host_id, skill_id) REFERENCES skill_t(host_id, skill_id) ON DELETE CASCADE
);
CREATE INDEX idx_agent_skill_agent ON agent_skill_t(agent_def_id);
-- Skill-Tool Mapping: Implements Progressive Disclosure
CREATE TABLE skill_tool_t (
host_id UUID NOT NULL,
skill_id UUID NOT NULL,
tool_id UUID NOT NULL,
config JSONB DEFAULT '{}',
access_level VARCHAR(20) DEFAULT 'read', -- e.g., 'read', 'write', 'execute'
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, skill_id, tool_id),
FOREIGN KEY(host_id, skill_id) REFERENCES skill_t(host_id, skill_id) ON DELETE CASCADE,
FOREIGN KEY(host_id, tool_id) REFERENCES tool_t(host_id, tool_id) ON DELETE CASCADE
);
CREATE INDEX idx_skill_tool_skill ON skill_tool_t(skill_id);
-- -- Hindsight Advanced Memory System
-- Transitioned from flat logs to biomimetic memory banks (World, Experiences, Mental Models)
-- Memory bank profiles (Personality & Disposition)
CREATE TABLE agent_memory_bank_t (
host_id UUID NOT NULL,
bank_id UUID NOT NULL,
agent_def_id UUID, -- NULL if bank is shared across agents
user_id UUID, -- NULL if bank is global for the host/agent
bank_name VARCHAR(126) NOT NULL,
disposition JSONB NOT NULL DEFAULT '{"skepticism": 3, "literalism": 3, "empathy": 3}'::jsonb,
background TEXT,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, bank_id),
FOREIGN KEY(host_id) REFERENCES host_t(host_id) ON DELETE CASCADE,
FOREIGN KEY(host_id, agent_def_id) REFERENCES agent_definition_t(host_id, agent_def_id) ON DELETE CASCADE,
FOREIGN KEY(user_id) REFERENCES user_t(user_id) ON DELETE CASCADE
);
-- Source documents for memory units
CREATE TABLE agent_memory_doc_t (
host_id UUID NOT NULL,
doc_id UUID NOT NULL,
bank_id UUID NOT NULL,
original_text TEXT,
content_hash TEXT,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY (host_id, bank_id, doc_id),
FOREIGN KEY (host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE
);
-- Individual sentence-level memories (The "Atoms" of thought)
CREATE TABLE agent_memory_unit_t (
host_id UUID NOT NULL,
unit_id UUID NOT NULL,
bank_id UUID NOT NULL,
doc_id UUID,
content TEXT NOT NULL,
embedding vector(384),
context TEXT,
event_date TIMESTAMP WITH TIME ZONE NOT NULL DEFAULT now(),
occurred_start TIMESTAMP WITH TIME ZONE,
occurred_end TIMESTAMP WITH TIME ZONE,
mentioned_at TIMESTAMP WITH TIME ZONE,
fact_type VARCHAR(32) NOT NULL DEFAULT 'world' CHECK (fact_type IN ('world', 'experience', 'opinion', 'observation', 'mental_model')),
metadata JSONB DEFAULT '{}'::jsonb,
proof_count INT DEFAULT 1,
source_memory_ids UUID[] DEFAULT ARRAY[]::UUID[],
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, bank_id, unit_id),
FOREIGN KEY(host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE,
FOREIGN KEY(host_id, bank_id, doc_id) REFERENCES agent_memory_doc_t(host_id, bank_id, doc_id) ON DELETE CASCADE
);
CREATE INDEX idx_mem_unit_bank ON agent_memory_unit_t(bank_id);
CREATE INDEX idx_mem_unit_embedding ON agent_memory_unit_t USING hnsw (embedding vector_cosine_ops);
-- Resolved entities (Knowledge Graph Nodes)
CREATE TABLE agent_memory_entity_t (
host_id UUID NOT NULL,
entity_id UUID NOT NULL,
bank_id UUID NOT NULL,
user_id UUID, -- Link to user_t if this entity is a platform user
canonical_name TEXT NOT NULL,
mention_count INT DEFAULT 1,
metadata JSONB DEFAULT '{}'::jsonb,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY (host_id, bank_id, entity_id),
FOREIGN KEY (host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE,
FOREIGN KEY (user_id) REFERENCES user_t(user_id) ON DELETE CASCADE
);
-- Association between memory units and entities
CREATE TABLE agent_memory_unit_entity_t (
host_id UUID NOT NULL,
bank_id UUID NOT NULL,
unit_id UUID NOT NULL,
entity_id UUID NOT NULL,
PRIMARY KEY (host_id, bank_id, unit_id, entity_id),
FOREIGN KEY (host_id, bank_id, unit_id) REFERENCES agent_memory_unit_t(host_id, bank_id, unit_id) ON DELETE CASCADE,
FOREIGN KEY (host_id, bank_id, entity_id) REFERENCES agent_memory_entity_t(host_id, bank_id, entity_id) ON DELETE CASCADE
);
-- Cache of entity co-occurrences (Concept Relationship Graph)
CREATE TABLE agent_memory_entity_cooccur_t (
host_id UUID NOT NULL,
bank_id UUID NOT NULL,
entity_id_1 UUID NOT NULL,
entity_id_2 UUID NOT NULL,
cooccur_count INT DEFAULT 1,
last_cooccurred TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY (host_id, bank_id, entity_id_1, entity_id_2),
CONSTRAINT entity_cooccur_order_check CHECK (entity_id_1 < entity_id_2),
FOREIGN KEY (host_id, bank_id, entity_id_1) REFERENCES agent_memory_entity_t(host_id, bank_id, entity_id) ON DELETE CASCADE,
FOREIGN KEY (host_id, bank_id, entity_id_2) REFERENCES agent_memory_entity_t(host_id, bank_id, entity_id) ON DELETE CASCADE
);
CREATE INDEX idx_mem_cooccur_e1 ON agent_memory_entity_cooccur_t(host_id, entity_id_1);
CREATE INDEX idx_mem_cooccur_e2 ON agent_memory_entity_cooccur_t(host_id, entity_id_2);
-- Links between memory units (Semantic & Causal relationships)
CREATE TABLE agent_memory_link_t (
host_id UUID NOT NULL,
bank_id UUID NOT NULL,
from_unit_id UUID NOT NULL,
to_unit_id UUID NOT NULL,
link_type VARCHAR(32) NOT NULL,
weight FLOAT NOT NULL DEFAULT 1.0,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY (host_id, bank_id, from_unit_id, to_unit_id, link_type),
CONSTRAINT memory_links_type_check CHECK (link_type IN ('temporal', 'semantic', 'entity', 'causes', 'caused_by', 'enables', 'prevents')),
FOREIGN KEY (host_id, bank_id, from_unit_id) REFERENCES agent_memory_unit_t(host_id, bank_id, unit_id) ON DELETE CASCADE,
FOREIGN KEY (host_id, bank_id, to_unit_id) REFERENCES agent_memory_unit_t(host_id, bank_id, unit_id) ON DELETE CASCADE
);
-- Directives (Hard rules that override probabilistic learning)
CREATE TABLE agent_memory_directive_t (
host_id UUID NOT NULL,
directive_id UUID NOT NULL,
bank_id UUID NOT NULL,
name VARCHAR(256) NOT NULL,
content TEXT NOT NULL,
priority INT NOT NULL DEFAULT 0,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, bank_id, directive_id),
FOREIGN KEY(host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE
);
-- Reflections (Synthesized knowledge and high-level observations)
CREATE TABLE agent_memory_reflection_t (
host_id UUID NOT NULL,
reflection_id UUID NOT NULL,
bank_id UUID NOT NULL,
content TEXT NOT NULL,
embedding vector(384),
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, bank_id, reflection_id),
FOREIGN KEY(host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE
);
CREATE INDEX idx_mem_reflection_embedding ON agent_memory_reflection_t USING hnsw (embedding vector_cosine_ops);
-- Raw Session History (The source of Truth for active conversations)
CREATE TABLE agent_session_history_t (
host_id UUID NOT NULL,
session_id UUID NOT NULL,
bank_id UUID NOT NULL, -- Links the session to a Hindsight bank
messages JSONB NOT NULL DEFAULT '[]'::jsonb,
metadata JSONB DEFAULT '{}'::jsonb,
aggregate_version BIGINT DEFAULT 1 NOT NULL,
active BOOLEAN DEFAULT true,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP,
update_user VARCHAR(126) DEFAULT SESSION_USER,
PRIMARY KEY(host_id, bank_id, session_id),
FOREIGN KEY(host_id, bank_id) REFERENCES agent_memory_bank_t(host_id, bank_id) ON DELETE CASCADE
);
CREATE INDEX idx_session_bank ON agent_session_history_t(host_id, bank_id);
Control Plane And Operational Data
Status
Proposed.
This design defines the storage and authority boundary for Light-Fabric configuration, runtime state, memory, artifacts, audit evidence, and analytical data. It also defines the database work that must precede production A2A runtime persistence.
It complements:
- Database Design;
- Tenant Operational Store Registration;
- Hindsight Memory;
- Light-Agent Execution;
- Light-Workflow Runner;
- Tracing; and
- A2A Gateway.
Decision
Light-Fabric separates control-plane data from operational data by authority and lifecycle, not merely by whether a row was created through an event.
- Light Portal authoring uses Event Sourcing and CQRS. Its events, authoring projections, publication records, and Config Server snapshots remain in the Config Server database or schema.
- Config Server distributes bounded, immutable, audience-specific runtime instructions. It does not become a database for sessions, tasks, memories, artifacts, traffic records, or other facts created by runtime activity.
- Each tenant/host operational boundary uses an operational database. Services own separate schemas and database roles within that boundary; they do not share tables merely because they share one physical database.
- Runtime services write their authoritative operational state directly in local transactions. They may also use append-only domain ledgers and a transactional outbox for audit, integration, or projection work. Those operational events do not become Portal configuration events.
- Light Portal may manage operational data through open, authenticated runtime administration APIs. Portal View does not write operational tables directly and does not move operational content through Config Server.
light-knowledgekeeps its operational data in its Knowledge database. An organization may share that database across multiple authorized hosts when the Knowledge Base scope is explicitly organization-wide.light-gatewayremains stateless for application work except for bounded caches, retry state, and a durable telemetry or audit spool when required. It does not own agent sessions, A2A tasks, memories, or workflow state.
The default is not one database per agent and not one gateway-owned database
shared by every agent. The recommended enterprise topology is one operational
database per tenant/host and environment, with service-owned schemas such as
agent_ops, a2a_ops, workflow_ops, gateway_ops, and audit_ops. Physical
pooling is a deployment choice described below and does not change logical
ownership.
Why This Boundary Is Needed
The current repository contains a transitional mixture:
- Portal configuration is compiled into immutable
agent.ymland gateway projections; light-agentsessions, turns, actions, approvals, pinned policy evidence, quotas, and memory are operational state;- the current memory adapter can send some writes through Portal commands while reading session history and recall data directly from PostgreSQL; and
- several deployment examples still point runtime services at the
configserverdatabase.
That mixture is useful during bootstrapping but should not become the target architecture. It couples runtime availability and write throughput to the Portal command path, makes backup and retention boundaries ambiguous, and makes it difficult to run Light-Fabric without Light Portal.
The target keeps Light Portal valuable as the management plane without making it a required hop for a conversation, workflow transition, A2A task update, or memory enrichment.
Goals
- Preserve Event Sourcing and CQRS for Portal-managed authoring and publication.
- Give every operational record one authoritative service and store.
- Support enterprise dedicated databases, Portal cloud pooling, and customer-managed standalone deployments through the same contracts.
- Let every Light-Fabric service run without Light Portal using local configuration, secrets, and an operational store.
- Let Light Portal manage the same services more effectively through live configuration, status, operational administration, and audit views.
- Keep high-volume or high-churn data out of Config Server snapshots and authoring projections.
- Apply consistent host, principal, agent, workflow, task, and retention authorization to operational APIs.
- Define memory placement without confusing centrally managed content with control-plane configuration.
- Give A2A implementation phases a storage contract that does not create new operational tables in the Config Server schema.
- Permit independent backup, restore, retention, erasure, residency, and scale policies for configuration, operational state, Knowledge, artifacts, and analytics.
Non-Goals
- Do not require a separate physical PostgreSQL server for every service.
- Do not create one database for every logical agent definition.
- Do not make
light-gatewaythe owner of agent, A2A, or workflow state. - Do not put database credentials or other secrets in Config Server values.
- Do not copy mutable operational content into
values.yml. - Do not require all existing runtime tables to move before protocol-only A2A work can begin.
- Do not make Portal View a database client.
- Do not use the Portal event log as a high-volume traffic-log sink.
- Do not treat a log, metric, trace, artifact, chat message, or recalled memory as runtime authority.
Terminology
Control Plane
The control plane owns intended state: what an authorized administrator wants a service to do. Examples include an agent definition, effective skill set, route binding, access policy, memory policy, retention profile, model alias, workflow definition, and approved A2A publication.
Portal command events and their CQRS authoring projections belong here. A published generation is compiled into an immutable runtime projection and made available through Config Server.
Config Server Projection
A Config Server projection is the bounded, audience-specific input accepted by one runtime instance. It is selected by the registered host, service ID, environment tag, and instance binding. It contains policy, identifiers, digests, limits, and references. It must be usable without a live query to Portal authoring tables.
The word projection here does not mean a query view of operational activity. An operational read model belongs to the operational domain even if it is also built asynchronously.
Operational Or Data Plane
The operational plane owns facts created while work runs: what happened, what is happening, and the durable content required to continue or explain it. Examples include sessions, turns, workflow tasks, A2A correlation, memory units, action attempts, artifact metadata, quota consumption, idempotency records, audit evidence, and deletion tombstones.
Operational tables are normally written directly in the owning service’s transaction. An append-only domain event stream or outbox can accompany that transaction without changing the data into control-plane configuration.
Management Plane
Portal View is a management surface over both planes:
- configuration forms send commands to the Portal control plane; and
- operational screens call authenticated administration or query APIs exposed by the owning runtime service or a shared operational service.
The UI location does not determine data ownership. Editing a user memory in Portal View does not make that memory Config Server data.
Observability And Analytical Plane
Logs, metrics, traces, and analytical traffic records are derived operational evidence. A collector or durable audit publisher sends them to an approved log, telemetry, or analytics store. The control plane contains their collection, redaction, sampling, retention, and destination policy, not the collected records.
Classification Test
Use the following questions for every new field or table.
| Question | Control-plane signal | Operational signal |
|---|---|---|
| Who is authoritative? | An authorized administrator, publication workflow, or policy compiler. | A running service, authenticated user interaction, backend result, or reconciler. |
| What does the value describe? | Intended behavior and allowed authority. | Work performed, content learned, current state, or evidence. |
| What causes change? | Explicit authoring, review, approval, publication, revocation, or promotion. | Requests, conversations, task transitions, timers, callbacks, retries, or cleanup. |
| What is the write rate? | Low to moderate and versioned. | Potentially high-volume and high-churn. |
| Can a runtime snapshot reproduce it? | Yes; the published generation is the authority. | No; it can be recreated only by replaying operational history or not at all. |
| Does it require independent retention or erasure? | Usually tied to authoring and publication history. | Often tied to user, session, task, legal hold, artifact, or audit policy. |
| Does the runtime need to write it while Portal is unavailable? | Normally no. | Normally yes. |
When signals are mixed, split the concept into policy and instance state rather than storing one ambiguous aggregate. For example, a memory-retention profile is control-plane data; the expiry and deletion evidence for a particular memory unit are operational data.
Target Architecture
LIGHT PORTAL
Portal View configuration forms Portal View operations views
| |
v v
Portal command APIs Open operational APIs
| |
v v
Event store and CQRS Owning service
authoring projections |
| v
v Tenant/host operational DB
publication compiler and artifact storage
|
v
Config Server snapshots
|
+------------------+--------------------------+
|
v
gateway / agent / A2A / workflow
immutable runtime policy
Runtime logs, metrics, traces, and audit outboxes
|
v
collector / audit publisher -----> audit and analytics stores
Standalone deployment:
values.yml + secret files + the same operational APIs and operational stores
There are two management routes, not one overloaded configuration route:
- the configuration route authors and publishes intended state; and
- the operational route reads or mutates runtime-owned content under the same fine-grained authorization model used by the runtime.
An audit route collects evidence from both without becoming authoritative for either.
Data Ownership Matrix
| Domain | Control-plane authority | Operational authority | Recommended storage |
|---|---|---|---|
| Gateway routing and access | Routes, Instance API bindings, rule bodies, limits, redaction, and telemetry policy. | Bounded rate state, circuit state, accounting delivery, audit spool, and edge correlation. | Config Server projection plus gateway_ops or external telemetry systems. |
| Agent definition | Prompt, model, skills, tools, memory policy, knowledge bindings, execution policy, and public A2A publication. | None; the definition is intended state. | Portal event store, CQRS projection, and Config Server runtime projection. |
| Agent execution | Session and turn limits, approval rules, quota policy, retention, and data boundary. | Sessions, turns, action attempts, approvals, policy evidence, idempotency, quotas, and session events. | agent_ops in the tenant/host operational database. |
| Agent capacity and service pools | Pool definitions, agent assignments, compatibility dimensions and digests, enablement, and maximum_concurrency. | Accepted pool ID and digest, active occupancy, reservations, leases, and queue state. | Pool policy in the Agent projection; occupancy and reservation rows in agent_ops. |
| A2A external integration | Publications, backend bindings, protocol profiles, signing profiles, fine-grained policy, limits, and retention. | External adapter correlation, task facade state, idempotency, callbacks, cancellation, artifact metadata, and deletion evidence. | a2a_ops owned by light-a2a. |
| A2A native agent | Native A2A policy and publication. | Context/session and task/turn aliases plus native artifacts; the underlying session remains authoritative. | agent_ops owned by light-agent. |
| Workflow | Definitions, deployment policy, allowed callers, execution profiles, and retention policy. | Process/task state, worklists, attempts, leases, approvals, timers, outbox, and artifacts. | workflow_ops and shared runner-owned schemas. |
| Knowledge | Knowledge Base definitions, source bindings, ACL policy, embedding policy, and retention policy. | Ingested documents, chunks, embeddings, graph state, indexing jobs, and query evidence. | Dedicated Knowledge database and object storage. |
| Hindsight memory | Bank classes, sharing policy, retention, provider, recall limits, hard directives, and promotion rules. | Bank instances, memory units, links, entities, reflections, provenance, session history, erasure, and deletion evidence. | memory_ops, initially colocated with agent_ops but owned behind a Memory API. |
| Artifacts | Type, size, scan, visibility, retention, export, legal-hold, and promotion policy. | Bytes, immutable digest, owner, task linkage, scan result, expiry, holds, and tombstones. | Tenant object storage plus owning operational metadata schema. |
| Audit and traffic analysis | Required events, fields, redaction, sampling, retention, and sink policy. | Audit records, traffic observations, trace correlation, delivery cursor, and integrity evidence. | Tenant audit store and approved log/analytics platform; never Config Server. |
Service ownership must remain valid when schemas are later moved into separate physical databases. Cross-service relationships therefore use stable IDs, authenticated APIs, and integration events rather than cross-schema foreign keys that require permanent colocation.
The service-pool row is a concrete example of a split that the current code has
not yet made. light-agent selects a pool by joining agent_pool_assignment_t
and agent_service_pool_t inside its admission transaction and takes row locks
on both, while agentPolicy.execution.servicePools already carries the same
definitions in the immutable Agent projection. After the split, pool
definitions, assignments, compatibility digests, and concurrency ceilings are
read from the accepted projection, and only occupancy and reservation rows may
be locked, because a runtime cannot hold a database lock on control-plane
content it no longer stores.
Tenant And Host Storage Topology
Recommended Logical Boundary
The isolation key is the Portal tenant/host and environment. A production runtime must bind its operational store to the same host and environment as its accepted Config Server projection. A connection that resolves to a different boundary fails startup or reload validation.
Within the boundary:
- every service has a distinct database role;
- every service owns its migrations and schema version;
- roles cannot write another service’s schema;
- tenant/host identifiers remain on authoritative rows as defense in depth;
- backup, restore, erasure, and legal-hold operations preserve service ownership; and
- cross-service reporting uses APIs, outbox events, or read-only analytical projections rather than shared write access.
Physical Deployment Profiles
| Profile | Physical layout | Intended use | Required invariants |
|---|---|---|---|
DEDICATED_HOST | One operational database per tenant/host and environment, with service-owned schemas. | Enterprise and regulated deployments. | Dedicated credentials, host binding, independent backup and residency. |
POOLED_TENANT | Multiple tenants share a managed cluster or database; every row and partition is tenant-scoped. | Light Portal cloud for small and intermediate customers. | Enforced tenant key, row-level or equivalent isolation, per-tenant encryption context, noisy-neighbor limits, and export/erasure support. |
CUSTOMER_MANAGED | Customer supplies the operational database and object store. | Standalone open-source or hybrid deployment. | Published migrations, readiness validation, least-privilege roles, and no Portal dependency. |
The product contract is logical isolation, not a promise that every small tenant receives a separate PostgreSQL server. A customer can promote from a pooled profile to a dedicated profile without changing service APIs or data semantics.
Knowledge Exception
Knowledge is often organization-scoped rather than runtime-host-scoped. A Knowledge database may therefore serve multiple hosts in the same organization when all of the following are explicit:
- an organization and Knowledge Base scope;
- per-source and per-principal authorization;
- host-to-Knowledge-Base bindings in control-plane policy;
- data residency and retention compatibility; and
- query-time delegation that cannot broaden the caller’s authority.
This exception does not authorize agent sessions, user memory, workflow tasks, or A2A task state to cross host boundaries.
Memory Boundary
Memory is the clearest example of why “managed in Portal View” does not mean “stored in Config Server.” Portal administrators and end users can manage memory centrally while the memory remains operational data.
| Memory concern | Plane | Reason |
|---|---|---|
| Allowed bank types and scopes | Control | Defines which sharing models may exist. |
| Default bank selection and creation policy | Control | Constrains runtime behavior. |
| Recall, retention, reflection, promotion, and erasure policy | Control | Defines authority and lifecycle. |
| Provider, embedding model, limits, and residency | Control | Selects governed runtime dependencies. |
| System prompt and hard authorization directives | Control | Must be reviewed, versioned, and immutable for an accepted publication. |
| Concrete user, agent, shared, or session bank | Operational | It is a runtime instance with an owner and lifecycle. |
| Session transcript and history projection | Operational | It is created and enriched by conversation. |
| User profile fact or preference | Operational | It can be edited centrally but represents mutable user content. |
| Learned fact, experience, link, entity, or reflection | Operational | It is derived from runtime activity. |
| Provenance, expiry, legal hold, erasure status, and deletion evidence | Operational | It applies to concrete content and must survive independently of configuration. |
Content called a “directive” requires careful classification. A reviewed hard rule that can change agent behavior or authority is control-plane content and must be projected with a digest. A remembered preference or observation is operational, untrusted model context and cannot grant tools, credentials, network access, or policy exceptions.
Memory Service Boundary
The first implementation may embed the Memory API and repository inside
light-agent while storing memory tables in memory_ops or agent_ops. The
API contract should nevertheless be independent of the repository so that a
shared open-source light-memory service can be introduced later for:
- user memory shared by several agents;
- organization-managed retention and erasure;
- centralized reflection and embedding work;
- Portal View administration; and
- scale or residency boundaries different from agent execution.
light-agent depends on the Memory API abstraction, not on Portal commands or
Portal table layouts. Portal View uses the same authenticated administration
API for search, correction, export, retention hold, and deletion. Every request
is checked against host, principal, agent, bank, operation, and fine-grained
policy.
The current Portal-command memory write mode is a migration adapter. It should not be the final production authority because it makes Portal availability part of the conversation write path while reads still depend on operational PostgreSQL state.
Configuration And Store Binding
Each runtime projection contains immutable store-binding policy, but no credential and no operational content. The runtime combines that projection with a deployment-owned secret file.
An illustrative shared contract is:
operationalStore:
contractVersion: ${operationalStore.contractVersion:1}
profileId: ${operationalStore.profileId:}
deploymentProfile: ${operationalStore.deploymentProfile:DEDICATED_HOST}
hostId: ${operationalStore.hostId:}
environment: ${operationalStore.environment:}
serviceOwner: ${operationalStore.serviceOwner:}
schema: ${operationalStore.schema:}
minimumSchemaVersion: ${operationalStore.minimumSchemaVersion:1}
expectedDatabase: ${operationalStore.expectedDatabase:operations}
databaseUrlFile: ${operationalStore.databaseUrlFile:/run/secrets/operational-database-url}
objectStoreProfileId: ${operationalStore.objectStoreProfileId:}
auditSinkProfileId: ${operationalStore.auditSinkProfileId:}
This block is a common semantic contract, not necessarily one shared Rust
configuration struct. A service may namespace it under agent, a2a,
workflow, or gateway while retaining the same validation rules.
The accepted binding must satisfy:
- projection audience, host, environment, service, and instance identity;
- an allowlisted deployment profile;
- exact service-owned schema and role;
- compatible migration and schema versions;
- an expected database identity check rather than only a syntactically valid connection string;
- secret-file availability and least-privilege connectivity;
- object and audit-store binding where required; and
- last-known-good reload semantics for mutable control-plane policy.
Config Server may distribute profileId, schema, version, expected database,
and secret reference. The actual connection URL, password, encryption key, and
object-store credential remain deployment secrets. Portal cloud may
materialize them through its secret-management integration; standalone users
mount the same files themselves.
The binding should eventually be reusable in agent.yml, the light-a2a
audience projection template, and workflow configuration rather than copying a
database URL into every product policy. light-knowledge retains its dedicated knowledge.yml connection and
database identity because its organization-shared topology is intentionally
different.
The Host Admin lifecycle, managed and customer-managed deployment profiles, secret handling, provisioning state machine, and decommission workflow are defined in Tenant Operational Store Registration.
Portal And Standalone Operation
With Light Portal
Portal provides:
- structured authoring, review, publication, revocation, and history;
- instance and store-profile binding;
- live Config Server generation and activation;
- runtime registration, health, schema compatibility, and drift visibility;
- operational views backed by service APIs;
- fine-grained administration of sessions, memories, tasks, artifacts, holds, export, and erasure; and
- audit and analytical dashboards backed by approved operational sinks.
Portal does not become a synchronous proxy for normal runtime database writes.
Without Light Portal
Every open-source service must accept the equivalent intended state from a
local values.yml or other supported configuration source, load secrets from
files or a customer secret manager, run its own migrations or validation under
an explicit startup policy, and expose the same operational APIs.
Configuration export and operational export are deliberately separate:
- a configuration bundle contains templates, values, public artifacts, digests, and secret references; and
- an operational backup or export contains runtime data under its own authorization, encryption, retention, and privacy policy.
Downloading values.yml must never silently download user memories, sessions,
task payloads, artifacts, traffic logs, or database credentials.
Operational API Requirements
Every service-owned operational domain exposes open contracts suitable for Portal and standalone administration. The APIs must:
- derive host and principal scope from authenticated server-side identity;
- authorize the exact operation and resource rather than trusting an ID;
- support pagination, filtering, and bounded exports;
- use idempotency and optimistic concurrency for mutations;
- emit normal audit evidence for read, correction, export, hold, and deletion;
- avoid returning secret material or raw content without explicit content-read authority;
- preserve provenance and deletion tombstones where policy requires them; and
- remain available independently of Portal authoring services.
Portal may cache a display projection, but the operational service remains authoritative. The UI must show source, freshness, and failures instead of presenting a stale Portal copy as current state.
Audit, Traffic, And Analysis
Traffic records are operational/analytical data, not Config Server data. The control plane specifies which events are required and which fields must be redacted. The runtime produces the records.
Use three complementary paths:
- Rust
tracingemits structured, policy-safe logs and trace correlation to stdout or an approved collector. - Authoritative security, accounting, approval, artifact, and deletion events are committed with the operational transaction through an outbox or an equivalent durable local handoff.
- High-volume traffic analysis flows to a tenant-approved log or analytical store. A bounded operational index may retain correlation, status, latency, policy digest, and evidence digest without duplicating request content.
Do not perform a synchronous remote audit-database write on every gateway request. A gateway that must survive collector failure uses a bounded, encrypted, backpressured spool with explicit fail-open or fail-closed policy by event class. Dropping a debug trace and losing a required authorization audit record are not the same failure.
Prompts, messages, model output, memory, tool arguments, task payloads, and artifact bytes are excluded from ordinary traffic logs by default. Fine-grained content access applies through the owning operational API; there is no separate implicit Portal-administrator bypass.
Availability And Failure Semantics
| Failure | Required behavior |
|---|---|
| Portal unavailable | Published runtimes continue with accepted Config Server or local configuration and their operational stores. Portal authoring and operational UI are unavailable, but the runtime data path does not stop solely for that reason. |
| Config Server temporarily unavailable | A runtime may continue with a valid last-known-good generation until its expiry and revocation policy requires failure. It does not query Portal tables as a fallback. |
| Operational database unavailable | The owning service fails or queues work according to its durability contract. It never falls back to writing operational rows into Config Server. |
| Audit collector unavailable | Required events use the configured durable spool or fail policy; optional telemetry may be sampled or dropped with metrics. |
| Knowledge database unavailable | Knowledge-dependent operations fail or degrade according to the published policy; agent sessions and memory do not silently move into the Knowledge database. |
| Portal operational API proxy unavailable | Direct runtime administration remains possible for authorized standalone operators; normal runtime work is unaffected. |
Migration From The Current Layout
Migration is organized by authority rather than by copying every table at once.
Step 0: Freeze Contracts
Before new runtime schemas are implemented:
- classify every existing and proposed table;
- define the tenant/host operational-store binding;
- define service schema ownership and least-privilege roles;
- define migration, readiness, backup, restore, and rollback contracts;
- pin operational API and audit/outbox envelopes; and
- prohibit new runtime-written tables in the Config Server schema.
Step 1: Bootstrap Tenant Operational Stores
- create repeatable database and schema provisioning;
- publish service-owned migrations;
- validate expected database, host, environment, role, and schema version;
- support dedicated, pooled, and customer-managed profiles; and
- add readiness and Portal diagnostics without exposing credentials.
Step 2: Decouple Cross-Boundary Constraints
No table moves in this step. It removes the referential dependencies that would otherwise make a physical move impossible, and it runs entirely inside the current database so that every change is reversible.
The executable Phase 0 inventory identifies 24 removable foreign keys plus one temporarily retained Agent-to-memory invariant:
| Crossing | Constraints | Replacement |
|---|---|---|
| Operational to control plane (12) | Agent memory, policy evidence, quota usage, sessions, runner requests, and runner sessions reference Host, user, Agent definition, quota policy, service-pool, or runner-binding authoring tables. | Local scope root plus pinned ID, version, publication, digest, validity, and revocation evidence validated at admission. |
| Operational to later Workflow service (6) | Execution attempts and scheduling requests reference Workflow process, task, or approval state; fixed actions reference Workflow approvals. | Stable Workflow reference plus authenticated API/event validation and reconciliation. |
| Agent to execution service (5) | Agent sessions, turns, action attempts, and approval consumption reference execution sessions, attempts, or scheduling requests. | Stable execution reference plus authenticated result/status events and reconciliation. |
| Control plane to operational semantic target (1) | agent_memory_directive_t references a concrete runtime memory bank. | Versioned Agent policy targets a bank profile or scope selector rather than a runtime bank ID. |
| Within the agent boundary, pending the later memory split | agent_session_t references agent_memory_bank_t. | Retained until the Memory API owns the invariant. |
The exact names, source and target columns, replacement contract, and test
owner are frozen in
implementation/light-portal/development-database-topology/phase0/foreign-key-boundary-v1.json.
Constraints wholly inside one service transaction boundary stay in place;
agent_session_t_host_id_bank_id_fkey is the named retained invariant.
- add a local operational-scope root in each operational schema that records the
expected host, environment, and, where applicable, organization scope, so that
scope validation no longer depends on an FK into
host_t; - replace control-plane foreign keys with pinned identifiers, versions, and digests that are validated at admission against the accepted projection instead of enforced by the database;
- replace cross-service foreign keys with stable references plus API, outbox, or reconciliation checks that restore the invariant the constraint used to guarantee;
- keep the constraint and its replacement active together long enough to prove the replacement rejects the same violations; and
- drop each constraint only after its replacement has test coverage.
Step 3: Move Shared Execution Foundations
Shared execution state moves before Agent cutover because Agent tables point at
it. Moving Agent first would strand execution_session_t,
execution_attempt_t, and runner_scheduling_request_t references across a
boundary that does not yet exist.
- migrate execution session, attempt, and runner scheduling state to its owning runner schema;
- establish the fencing, lease, replay, and reconciliation contracts that Agent and Workflow will both depend on; and
- verify that no Agent or Workflow table still requires a database-enforced reference into these tables.
Step 4: Move Agent Execution And Embedded Memory
- migrate
agent_session_t,agent_turn_t, action, approval, event, idempotency, quota, pool occupancy, and related operational tables toagent_ops; - colocate concrete Hindsight banks, memory content, graph data, reflections,
session history, and deletion evidence in
agent_opsinitially, becauseagent_session_tandagent_memory_bank_tstill share integrity expectations that no service yet owns; - keep memory policy, digests, and store bindings in the immutable Agent projection;
- migrate
agent_memory_directive_tas a semantic exception rather than a table move, as described below; - replace the Portal-command write authority with the Memory API and operational repository;
- make Portal View memory management call the Memory administration API; and
- backfill, compare, cut over, and remove the old write path without an unbounded dual-write period.
Split memory into a separate memory_ops schema only after the Memory API owns
the integrity that the agent_session_t to agent_memory_bank_t constraint
enforces today. Splitting the schema before the API owns the invariant converts
a database-enforced relationship into an unchecked one.
Hard Directives Are A Semantic Migration
agent_memory_directive_t is the one memory table that does not move to an
operational schema. A hard directive can change agent behavior or authority, so
it is control-plane content: it becomes versioned, reviewed, digest-bound
authoring data compiled into the Agent projection.
That reclassification also changes its shape. A directive currently references a concrete operational bank; a published directive instead targets a bank profile or scope selector, because control-plane content cannot depend on the existence of one runtime bank instance. Remembered user preferences and observations stay ordinary operational memory and are not promoted by this change.
Step 5: Move Remaining Workflow State
- migrate process, task, worklist, timer, approval, artifact, and outbox state under their owning services;
- remove cross-service write access and replace remaining cross-schema dependencies with stable contracts; and
- preserve replay, fencing, idempotency, and reconciliation evidence.
Step 6: Complete Gateway, Audit, And Analytics Separation
- move any durable accounting, retry, circuit, and audit-spool state to
gateway_opsor a purpose-built open service; - publish structured audit records to the tenant audit store;
- connect logs and traces to approved collectors; and
- verify that Config Server stores only policy and publication history.
Each cutover must define source-of-truth time, read routing, write fencing, backfill watermark, validation, rollback limit, and deletion of stale credentials. Indefinite bidirectional dual write is not an acceptable steady state.
Relationship To A2A Gateway Delivery
The complete platform-wide database migration is not a prerequisite for starting A2A work. The storage contract is a prerequisite, and separation must be implemented for every operational record that the A2A release itself creates.
| Work stream | May proceed before physical migration? | Storage dependency |
|---|---|---|
| A2A Phase 0 protocol, threat model, canonical operations, errors, and conformance fixtures | Yes. | Must adopt the ownership, retention, store-binding, artifact, and audit contracts from this design. |
| Shared parsing, version negotiation, card handling, and stateless gateway routing | Yes. | No durable task state in light-gateway; bounded caches only. |
| Portal A2A authoring, Instance API binding, card publication, policy compilation, and Config Server projections | Yes, in parallel with operational-store bootstrap. | These are control-plane records and immutable runtime projections. |
light-a2a external sidecar correlation, task facade, cancellation, artifact, and restart reconciliation | No for production. | a2a_ops, artifact storage, audit outbox, migrations, backup, and readiness must exist first. |
Native light-agent A2A task/turn mapping and artifacts | No for production. | Agent sessions, turns, idempotency, memory, and artifact metadata must be authoritative outside Config Server first. |
| Governed outbound A2A with durable task ownership and retry | No for production. | The calling runtime’s operational store and audit handoff must be ready. |
The recommended sequence is:
- complete Step 0 and freeze the cross-service storage contracts;
- start tenant operational-store bootstrap and A2A protocol/control-plane work in parallel;
- decouple cross-boundary constraints and move shared execution foundations;
- migrate Agent and embedded Memory operational state and create
a2a_ops; - implement production external-sidecar and native-agent A2A persistence on those stores;
- complete governed outbound A2A; and
- migrate remaining Workflow, Gateway, and analytical state incrementally.
This avoids two undesirable extremes: blocking useful A2A protocol and Portal
work on a platform-wide database move, or shipping A2A quickly by adding more
operational tables to configserver and making the eventual migration harder.
For the first production A2A release, the hard gate is:
No A2A task, context, idempotency record, callback state, artifact metadata, session/turn alias, retry record, or runtime audit evidence is authoritative in the Config Server database or schema.
Security And Isolation Requirements
- A runtime accepts an operational-store binding only when host, environment, audience, service, and schema identity match its accepted policy.
- Database credentials are service-specific and loaded from secret files or a secret manager.
- Migration credentials are distinct from runtime credentials.
- Runtime roles have no write access to Portal event, authoring projection, or Config Server snapshot tables.
- Portal services have no direct write access to runtime-owned schemas.
- Pooled deployments enforce tenant isolation in the database and in every API; an application predicate alone is insufficient.
- Object keys are opaque and tenant-scoped. Possession of an object URL or artifact ID is not access authority.
- Operational exports are encrypted, audited, bounded, and authorized independently of configuration export.
- Backup, restore, replication, and analytical pipelines preserve tenant, residency, retention, and deletion requirements.
- Recalled memory and analytical data remain untrusted input and cannot modify configuration or grant authority.
Verification And Exit Gates
Contract Gates
- every table in the implementation plan has one plane, owning service, schema, retention authority, and migration owner;
- configuration and operational exports have separate schemas and endpoints;
- Config Server projections contain store bindings and policy but no operational records or credentials;
- every operational API binds host and principal from authenticated state; and
- no service requires cross-schema write access.
Isolation Gates
- store-scope validation is storage-profile aware: an ordinary operational store validates the (host, environment) pair and refuses to start against a database bound to a different pair, while an organization-shared Knowledge store validates the (organization, knowledge-store binding, residency) tuple and additionally confirms that the requesting host is authorized for that Knowledge Base;
- a service role cannot write another service’s schema;
- pooled-tenant tests prove that guessed IDs and missing tenant predicates cannot cross the boundary;
- backup, restore, export, erasure, and legal-hold tests preserve tenant scope; and
- Portal View operations use authenticated APIs rather than direct database access.
Availability Gates
- Agent, Gateway, A2A, Workflow, and Knowledge runtimes continue their allowed work while Portal is unavailable;
- a valid last-known-good Config Server generation can be used according to its expiry and revocation contract;
- operational database failure never redirects writes to Config Server;
- required audit delivery survives a sink outage within the configured spool bounds; and
- standalone deployments pass the same operational API and migration tests as Portal-managed deployments.
Migration Gates
- zero cross-schema or cross-database foreign keys remain, and every constraint removed during decoupling has a documented replacement with test coverage proving it rejects the violations the constraint used to reject;
- backfill counts, ownership keys, digests, and representative business queries match before cutover;
- the cutover fences the old writer and has a bounded rollback point;
- no indefinite dual write remains;
- obsolete database credentials and grants are revoked; and
- the old operational tables are archived or removed only after retention and rollback obligations are satisfied.
A2A Gates
light-gatewayrestart loses no authoritative A2A task state because it owns none;light-a2arestart reconciles sidecar tasks froma2a_opswithout Portal or Config Server table queries;light-agentrestart preserves native context/session and task/turn mapping fromagent_ops;- A2A artifacts use tenant-scoped metadata and object storage with independent retention from chat and Hindsight memory;
- inbound and outbound retry/idempotency state remains within the selected runtime’s operational boundary; and
- A2A runtime and audit writes continue when Portal is unavailable.
Resolved Decisions
- Plane classification follows authority and lifecycle, not UI location or the mere presence of an event.
- Portal Event Sourcing and CQRS remain the control-plane authoring model.
- Config Server stores immutable runtime policy projections, not operational content.
- The default enterprise boundary is one operational database per tenant/host and environment with service-owned schemas and roles.
- There is no database per logical agent by default.
- A gateway does not own the operational database of agents it routes.
- Knowledge operational data stays in the Knowledge database and may be organization-shared under explicit policy.
- Memory policy and hard directives are control-plane data; concrete banks, session/user/agent memory, history, reflections, and deletion evidence are operational data.
- Portal View manages operational content through open authenticated APIs, not direct tables or Config Server publication.
- Standalone and Portal-managed deployments use the same runtime and operational contracts.
- Database credentials remain deployment secrets; Config Server carries bindings and references only.
- A2A storage contracts and tenant-store bootstrap precede durable A2A implementation, but a complete platform-wide migration does not block protocol, gateway-routing, or Portal-publication work.
- Cross-boundary constraints are decoupled, and shared execution foundations are moved, before any Agent or Memory table changes schema. Referential integrity is replaced deliberately rather than dropped as a side effect of a move.
- Hard memory directives are control-plane content targeting a bank profile or scope, not operational rows bound to a concrete bank instance.
- Service-pool definitions and concurrency ceilings are control-plane policy read from the Agent projection; only occupancy and reservation state is operational.
Open Questions
- Should the first shared Memory API ship embedded in
light-agentonly, or shouldlight-memorybe extracted immediately for user memory shared across agents? - Which operational administration APIs should be routed through
light-gateway, and which should remain on a private management network? - Which audit records require synchronous local durability, and which may use best-effort collector delivery?
- Should Portal cloud begin with database-per-tenant or pooled schemas with enforced tenant isolation, and what threshold promotes a tenant to a dedicated database?
- Which existing Portal command events represent true configuration and which operational commands must migrate first?
- Should artifact metadata begin in service-owned schemas or in a shared open-source artifact service with service-specific ownership tables?
Tenant Operational Store Registration
Operational storage is customer-owned data-plane infrastructure. The Portal control plane stores a Host-scoped, non-secret registration that tells runtime services where their organization has made a database available. The Portal does not connect to that database and does not create, rotate, stop, or delete it.
flowchart LR
A[Host administrator] -->|register metadata| P[Light Portal]
P -->|non-secret properties| C[Config Server]
C --> R[Gateway, Workflow, Agent, Deployer]
S[Deployment secret file] --> R
R --> D[(Customer operational database)]
The version-2 contract is scoped to HOST; an Environment field is not part
of registration. Runtime instances still have environments, and the Portal
projects the same Host registration to each eligible instance while preserving
that instance’s environment as routing metadata.
The registration contains the database engine, DNS name, port, database name,
TLS mode, runtime username, schema generation, credential generation, and a
logical credential reference. It must never contain a password or database
URL. For MOUNTED_FILE, the reference is an absolute file path such as
/run/secrets/operational-database-url.
The only active lifecycle operations are register, update, deactivate, and unregister. Version-1 provisioning events remain replayable to rebuild historical audit state, but live submission is rejected and the historical job and provider-profile tables are write-guarded.
Local development uses three databases in the existing PostgreSQL container:
operations, operations_networknt, and operations_taiji. Production
customers register databases operated inside their own organizations.
Light-Deployer Design
light-deployer is the cluster-local Kubernetes deployment executor in
Light Fabric.
This document focuses only on the deployer service that lives in
apps/light-deployer. The broader Light Portal deployment workflow, approval
flow, deployment history model, controller routing, and portal UI are covered
outside this repository.
Purpose
light-deployer receives a deployment command, fetches Kubernetes templates,
renders them with deployment values, validates the resulting resources, applies
or deletes resources in the target Kubernetes cluster, and returns safe status
details.
It is intentionally narrow. It does not decide whether a user is allowed to deploy an instance, does not own portal deployment history, and does not create tenant business workflows. Those decisions belong to Light Portal, Light Controller, and the workflow engine.
Service Boundary
light-deployer owns:
- local deployment policy enforcement
- template repository fetch
- YAML template rendering
- manifest parsing and resource summary generation
- Kubernetes dry-run, apply, delete, status, and pruning
- safe event and error reporting
- direct local/MicroK8s deployment endpoints
light-deployer does not own:
- tenant authorization
- instance metadata
- deployment approval
- deployment history persistence
- config snapshot creation
- long-running human workflow decisions
The deployer should reject commands outside its local policy even if an upstream service sends them.
Runtime Model
The service follows the same runtime pattern as light-agent.
main.rs builds the domain service and starts it through:
#![allow(unused)]
fn main() {
LightRuntimeBuilder::new(AxumTransport::new(app))
}
The HTTP listener is owned by light-runtime and light-axum, not by
service-specific socket code. Bind address, HTTP/HTTPS ports, service identity,
and registry settings live in runtime config files.
Default config files:
config/server.ymlconfig/deployer.ymlconfig/portal-registry.yml
Local cargo run resolves config from apps/light-deployer/config when run
from the workspace root. The container image runs from /app and uses
/app/config.
Public Endpoints
Phase 1 exposes a direct HTTP surface for local and MicroK8s testing:
GET /health
GET /ready
POST /mcp
GET /mcp/tools
GET /mcp/tools/list
GET /mcp/tools/{tool}
POST /deployments
POST /mcp/tools/{tool}
GET /events?request_id=...
POST /mcp is the MCP JSON-RPC 2.0 endpoint. It supports tools/list,
tools/call, and a minimal initialize response. This is the endpoint that
MCP clients, Light Portal, and AI agents should use.
/deployments accepts the canonical deployment request directly.
/mcp/tools/{tool} maps tool names onto the same internal service functions as
a REST-style local debugging convenience. The convenience tool-list endpoints
return metadata with name, description, inputSchema, endpoint, and
method, but they are not the MCP protocol endpoint.
Supported tool names:
deployment.renderdeployment.dryRundeployment.diffdeployment.applydeployment.deletedeployment.statusdeployment.rollback
The direct HTTP mode is useful for development and managed environments. The same internal command handling should later be reused by controller-mediated WebSocket/MCP routing.
Request Model
A deployment request is explicit and auditable.
{
"requestId": "01964b05-0000-7000-8000-000000000001",
"hostId": "01964b05-552a-7c4b-9184-6857e7f3dc5f",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "petstore-dev",
"action": "deploy",
"values": {
"name": "petstore",
"image": {
"repository": "networknt/openapi-petstore",
"tag": "latest"
}
},
"template": {
"repoUrl": "https://github.com/networknt/openapi-petstore.git",
"ref": "master",
"path": "k8s"
},
"options": {
"dryRun": false,
"waitForRollout": true,
"timeoutSeconds": 300,
"pruneOverride": false
}
}
The current implementation supports inline values. The request model also
contains fields for future values references and immutable snapshot metadata so
it can align with the full portal deployment workflow.
When invoking a specific /mcp/tools/{tool} endpoint, callers do not need to
send action. The deployer derives the action from the tool name. The generic
/deployments endpoint still expects an explicit action in the request body.
For the MCP endpoint, callers use JSON-RPC:
{
"jsonrpc": "2.0",
"id": "tools-list-1",
"method": "tools/list",
"params": {}
}
Tool invocation uses tools/call:
{
"jsonrpc": "2.0",
"id": "render-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "local",
"namespace": "light-deployer",
"values": {},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}
tools/call derives the deployment action from params.name; callers should
not provide an action field in arguments.
Actions
render- Fetch templates, render manifests, add namespaces and management labels, and return resource summaries plus a manifest hash.
dryRun- Render manifests and validate them against Kubernetes using server-side dry-run.
diff- Render manifests, fetch current managed resources, calculate additions, modifications, and pruned resources, and return a redacted diff summary.
deploy- Accept the request, run the deployment in the background, apply manifests, prune removed managed resources, and stream events.
undeploy- Delete resources associated with the deployment.
status- Return current managed resource status.
rollback- Reserved for redeploying a previous immutable portal snapshot. Native Kubernetes rollout undo is not the target rollback model because it does not restore ConfigMaps, Secrets, or values snapshots.
Template Fetching
Templates are loaded through the TemplateSource trait.
The current source supports two modes:
- local template root through
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR - remote HTTPS Git clone through
gix
For remote repositories, the deployment request provides:
{
"template": {
"repoUrl": "https://github.com/networknt/openapi-petstore.git",
"ref": "master",
"path": "k8s"
}
}
Private HTTPS Git access is controlled by environment variables:
LIGHT_DEPLOYER_GIT_TOKEN: token or app passwordLIGHT_DEPLOYER_GIT_USERNAME: optional username override
Defaults:
- GitHub uses
x-access-token - Bitbucket Cloud uses
x-token-auth
SSH authentication is intentionally deferred because it requires private key
handling and strict known_hosts validation.
Template Format
The built-in renderer uses simple placeholders:
image: ${image.repository}:${image.tag:latest}
Supported behavior:
- nested paths such as
image.repository - default values after
: - render failure when a required value is missing
- placeholder replacement only inside YAML string scalar values
The renderer parses YAML into serde_yaml::Value, traverses the AST, replaces
placeholders, and serializes or applies structured YAML values afterward. This
avoids the most common raw string replacement bugs around quoting,
indentation, certificates, and multi-line values.
Because placeholders currently produce strings, templates should avoid
placeholders in numeric-only Kubernetes fields unless Kubernetes accepts a
string value there. For example, containerPort should be fixed or rendered by
a future typed placeholder extension.
Resource Metadata
After rendering, the deployer ensures every resource has the target namespace and adds management labels:
app.kubernetes.io/managed-by=light-deployerlightapi.net/host-idlightapi.net/instance-idlightapi.net/request-id
These labels are used for status lookup and pruning.
Kubernetes Execution
Kubernetes execution is behind the KubeExecutor trait.
Current implementations:
KubeRsExecutor: real Kubernetes API execution throughkube-rsNoopKubeExecutor: local render/test mode
Execution mode:
LIGHT_DEPLOYER_KUBE_MODE=real: force real Kubernetes modeLIGHT_DEPLOYER_KUBE_MODE=noop: force no-op mode- default: real mode when
KUBERNETES_SERVICE_HOSTis present, otherwise no-op
The production path uses kube-rs, not kubectl.
Kubernetes operations should use:
- in-cluster ServiceAccount auth when running as a pod
- server-side dry-run for validation
- server-side apply with field manager
light-deployer - structured status and error handling
Pruning
The deployer is declarative. If a previously managed resource is no longer rendered from the template, it should be considered for pruning.
Pruning is calculated by comparing:
- current resources in the namespace with
lightapi.net/instance-id - resources rendered from the new template
The policy layer enforces blast-radius protection:
- maximum delete percentage
- sensitive kinds requiring override
- explicit
pruneOverridein deployment options
This prevents stale resources while still protecting against accidental large-scale deletion.
Policy
The local deployer.yml policy constrains what a deployer is allowed to do.
Policy dimensions:
- allowed namespaces
- allowed repository hosts
- allowed repository URL prefixes
- allowed image registries
- allowed actions
- allowed Kubernetes kinds
- blocked Kubernetes kinds
- prune settings
- development insecure mode
Version 1 allows application-level resource kinds by default:
DeploymentServiceIngressConfigMapSecret
Cluster-scoped and control-plane resources are blocked by default:
NamespaceClusterRoleClusterRoleBindingCustomResourceDefinition- admission webhooks
Security
The deployer can mutate a Kubernetes cluster, so its default posture must be conservative.
Required practices:
- run in Kubernetes with a dedicated ServiceAccount
- prefer namespace-scoped
RoleandRoleBinding - restrict allowed namespaces and resource kinds
- restrict template repository hosts or prefixes in production
- restrict image registries in production
- never log raw rendered Secret manifests
- never log raw Kubernetes patch/apply payloads containing Secret data
- return redacted summaries and diffs
Secret values in rendered manifests are redacted before being included in responses or diffs. Kubernetes Secret values are base64 encoded, not encrypted, so they must be treated as plaintext for logging purposes.
Response Model
Responses include enough detail for callers to understand what happened without exposing secrets.
Important fields:
requestIdactionstatusdeployerIdclusterIdnamespacemanifestHashtemplateCommitSharesourcesdiffeventserror
Resource summaries contain kind, namespace, name, apiVersion, and action. Full rendered manifests should not be returned or persisted by default.
Event Model
Long-running operations return quickly and continue in the background.
Clients can subscribe to:
GET /events?request_id=...
Events contain:
- request ID
- timestamp
- status
- message
- optional resource identity
The event stream is currently direct SSE. Controller-mediated mode can forward the same event shape later.
Installation
The app includes Kubernetes install manifests under apps/light-deployer/k8s:
- namespace
- RBAC
- deployment
- service
The deployment runs the container with LIGHT_DEPLOYER_KUBE_MODE=real. The
image contains /app/config, and server.yml defaults the HTTP port to 7088.
For MicroK8s testing:
./apps/light-deployer/build.sh latest
docker save networknt/light-deployer:latest | microk8s ctr image import -
microk8s kubectl apply -f apps/light-deployer/k8s/namespace.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/rbac.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/deployment.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/service.yaml
Current Limitations
- Direct HTTP/MCP-style mode is implemented first; controller-mediated WebSocket routing is a later integration step.
- Inline values are implemented; config-server
valuesReffetching is still a future integration point. - Rollback is represented in the model but needs portal snapshot integration.
- Helm and Kustomize are not implemented yet.
- Typed placeholders are not implemented yet.
- Rollout watch depth is intentionally basic in the first phase.
Design Direction
Keep light-deployer small and cluster-local.
The deployer should execute precise deployment commands, enforce local safety policy, and report structured results. It should not grow into a portal, workflow engine, or deployment database. That separation keeps the service easy to install inside customer clusters and reduces the security blast radius.
Module Registry
Status: Phase 4 implemented for light-gateway/gateway; additional module
reloaders remain planned.
Purpose
Light Fabric needs a runtime module registry equivalent to the ModuleRegistry
feature in light-4j.
In light-4j, each active component registers its runtime configuration when
the component loads. Older integrations exposed this through the
/adm/server/info REST endpoint, but the current control-plane path uses MCP
tools through portal-registry. The same registry is also used by the
config-reload operation to decide which modules can reload configuration from
the config server.
Light Fabric already has structured config files and a shared runtime startup flow, but it does not yet have a central registry that answers these operational questions:
- which modules are active in this running instance
- which config file each module loaded
- what masked runtime config is currently active
- which modules can be reloaded without restarting the process
- what happened during the last reload attempt
This document proposes a registry in light-runtime so every Light Fabric
application can expose the same control-plane behavior.
Goals
- Register built-in runtime configs such as
startup,server,client, andportal-registry. - Register application configs such as
gateway,deployer,ollama, andmcp-client. - Store only masked config snapshots in the registry.
- Expose a Java-compatible server-info payload through the
get_service_infoMCP tool. - Expose a module list through the
get_modulesMCP tool for config reload selection. - Support control-plane reload requests for one module, several modules, or all
modules through the
reload_modulesMCP tool. Phase 3 reports non-reloadable modules as skipped. Phase 4 adds real hot reload forlight-gateway/gateway. - Keep the feature transport-neutral by routing management requests through
portal-registry, not through framework-specific REST routes.
Non-Goals
- Do not make every config hot-reloadable in the first phase.
- Do not rebind server ports or TLS listeners unless a transport explicitly supports it.
- Do not expose decrypted secrets through diagnostics.
- Do not make Rust type names part of the public control-plane contract.
- Do not add
/adm/...REST endpoints for Light Fabric.
Current Light Fabric Runtime Shape
The natural home for this feature is crates/light-runtime.
LightRuntimeBuilder already owns the startup sequence:
- load local bootstrap config
- optionally fetch remote config from config server
- build
RuntimeConfig - call registered runtime modules
- bind the transport
- register the running instance with the controller
- mark the runtime ready
RuntimeConfig already carries the merged resolved_values, config_dir, and
external_config_dir. Application code can use those fields to load resolved
application config without reparsing values.yml.
The config registry should build on that runtime boundary instead of creating a separate app-local registry per product.
Registry Model
Add a shared registry type in light-runtime.
#![allow(unused)]
fn main() {
pub struct ModuleRegistry {
entries: RwLock<BTreeMap<String, ModuleEntry>>,
reloaders: RwLock<BTreeMap<String, Arc<dyn ReloadableModule>>>,
}
pub struct ModuleEntry {
pub module_id: String,
pub config_name: String,
pub kind: ModuleKind,
pub active: bool,
pub enabled: Option<bool>,
pub reloadable: bool,
pub config: serde_json::Value,
pub masks: Vec<MaskSpec>,
pub loaded_at: DateTime<Utc>,
pub last_reload: Option<ReloadStatus>,
}
pub enum ModuleKind {
Core,
Framework,
Application,
Plugin,
}
}
Use stable module IDs instead of Rust type names. Java uses class names because they are stable operational identifiers in the JVM. Rust type names are not a good public API and can change during refactoring.
Example module IDs:
light-runtime/startuplight-runtime/serverlight-client/clientlight-runtime/portal-registrylight-gateway/gatewaylight-deployer/deployerlight-agent/ollamalight-agent/mcp-client
The registry key should be module_id. Each entry also carries config_name
so the server-info response can preserve the Java-style component map keyed by
config name.
Registered Config Loading
Add a small registered-loader API around the existing ConfigLoader behavior.
#![allow(unused)]
fn main() {
let gateway_config: GatewayConfig = context
.config()
.load_registered(
"gateway",
"light-gateway/gateway",
[MaskSpec::key("password")],
)?;
}
The helper should:
- merge the base file from
config_dir - overlay the external file from
external_config_dir - resolve variables from
RuntimeConfig.resolved_values - deserialize the typed config
- serialize the resolved config to
serde_json::Value - apply masks to the serialized copy
- store only the masked copy in
ModuleRegistry - return the typed config to the caller
This keeps the app code simple and prevents accidental registry entries that contain raw secrets.
Phase 2 added this shared registered-loader path in ModuleRegistry and
attached the registry to RuntimeConfig so apps that load after runtime
bootstrap can register resolved config through the same runtime-owned registry.
Apps that load before runtime startup can create the registry first, register
their application configs, and pass that registry into LightRuntimeBuilder.
For modules that must validate typed config before changing the registry
snapshot, the same loader is also available as load_config(...) followed by
register_loaded_config(...) after validation succeeds.
Masking
Masking must happen at registration time. The registry should not store raw config and then mask it later.
Support two mask forms:
#![allow(unused)]
fn main() {
pub enum MaskSpec {
Key(String),
Path(String),
}
}
MaskSpec::Key("password") masks every matching key recursively, matching the
current light-4j behavior.
MaskSpec::Path("oauth.clientSecret") masks a precise path for configs where a
generic key would be too broad.
Suggested default masks:
authorizationpasswordsecretclientSecretapiKeytokenportalTokencontrollerDiscoveryTokenprivateKeytlsKeyPathbootstrapKeyPath
Add a runtime flag such as server.maskConfigProperties or
admin.maskConfigProperties, defaulting to true, for parity with the Java
server.maskConfigProperties behavior. Even if this flag is disabled, the
control-plane documentation should treat unmasked output as a local debugging
mode only.
Server Info MCP Response
The get_service_info MCP tool response should preserve the same logical shape
that portal-view already understands from Java instances.
{
"deployment": {
"apiVersion": "0.1.0",
"frameworkVersion": "0.1.0"
},
"environment": {
"host": {
"ip": "127.0.0.1",
"hostname": "light-gateway-0"
},
"runtime": {},
"system": {}
},
"security": {},
"component": {
"server": {},
"gateway": {}
},
"plugin": {},
"plugins": [],
"modules": []
}
component should remain keyed by config_name for compatibility.
modules should provide richer Rust metadata:
[
{
"moduleId": "light-gateway/gateway",
"configName": "gateway",
"kind": "application",
"active": true,
"enabled": true,
"reloadable": true,
"loadedAt": "2026-05-07T14:30:00Z",
"lastReload": {
"status": "success",
"message": "reloaded from config server",
"completedAt": "2026-05-07T14:45:00Z"
}
}
]
MCP Access
Expose the registry only through MCP tools served by the runtime’s
portal-registry connection.
MCP tools:
get_service_info
get_modules
reload_modules
These are invoked through standard MCP JSON-RPC calls:
{
"jsonrpc": "2.0",
"id": "info-1",
"method": "tools/call",
"params": {
"name": "get_service_info",
"arguments": {}
}
}
The controller remains the management channel. portal-registry receives the
MCP request from the controller, dispatches it to the local runtime registry,
and returns the result through the same websocket session. Light Fabric should
not expose a parallel REST admin surface for this feature.
For compatibility with the existing Java and portal-view workflow,
get_modules returns a string list of module IDs:
{
"modules": [
"light-runtime/server",
"light-gateway/gateway"
]
}
The richer module metadata remains available in the modules field of
get_service_info.
Reload Request
The reload_modules tool should accept omitted arguments, ALL, or explicit
module IDs.
{
"modules": [
"light-gateway/gateway",
"light-runtime/portal-registry"
]
}
An omitted modules value, an empty array, or ["ALL"] targets all registered
modules. Registered modules without concrete reload implementations are
reported as skipped instead of being marked as reloaded.
The response should be explicit about what happened:
{
"modules": ["light-gateway/gateway"],
"reloaded": ["light-gateway/gateway"],
"skipped": [
{
"moduleId": "light-runtime/server",
"reason": "requiresRestart"
}
],
"failed": [
{
"moduleId": "light-agent/ollama",
"message": "missing ollama.yml"
}
]
}
modules is a Java-compatible alias for the successfully reloaded module IDs
and is the field portal-view reads today. reloaded, skipped, and failed
carry the more explicit Rust result details.
Reload Implementation
Phase 4 adds a reload trait for modules that can safely swap runtime config.
#![allow(unused)]
fn main() {
#[async_trait]
pub trait ReloadableModule: Send + Sync {
async fn reload(&self, ctx: ReloadContext) -> Result<ReloadOutcome, RuntimeError>;
}
}
ReloadContext includes:
- a refreshed
RuntimeConfig - updated
resolved_values - the existing
config_dir - the existing
external_config_dir - the shared
ModuleRegistry
Reload flow:
- Re-fetch
values.yml, certs, and files from the config server intoexternal_config_dir. - Rebuild the merged
resolved_values. - Resolve requested module IDs.
- For each reloadable module, call its
reloadimplementation. - Each module validates the new typed config before swapping it into live state.
- Update the registry entry and
last_reloadstatus. - Return a detailed reload result.
Use ConfigManager<T> or another ArcSwap-backed holder for modules that need
hot reload. This avoids locking the request path while still allowing atomic
config replacement.
Phase 4 implements this with ConfigManager<T> in light-runtime. It stores an
Arc<T> behind a short-lived RwLock, so request handlers clone the current
config quickly and reloaders replace the entire typed config only after the new
config has loaded and validated.
Reloadability Rules
Classify configs by reload safety.
Reloadable candidates:
light-gateway/gatewaylight-deployer/deployerlight-agent/ollamalight-agent/mcp-client- route, policy, provider, or rule configs that are already read through swappable state
Requires restart by default:
- bind IP
- HTTP/HTTPS port
- protocol enablement
- TLS certificate path used by the listener
- runtime config directory
- config-server bootstrap identity
- controller registration identity
Some server.yml fields can still be reloadable later, such as
shutdownGracefulPeriod, but listener-affecting fields should stay
requiresRestart until each transport supports safe rebinding.
Framework Integration
The registry should not require each framework to expose admin routes.
light-runtime should attach an MCP-capable RegistryHandler to the
portal-registry client. When the controller invokes tools/list or
tools/call, the handler can advertise and execute the local management tools
without involving light-axum or light-pingora request routing.
This keeps light-axum and light-pingora focused on application traffic. It
also avoids adding service ports, Kubernetes routes, or Pingora request filters
only for control-plane operations.
Application Integration
light-gateway is integrated first because it already loads gateway.yml from
RuntimeConfig.resolved_values, config_dir, and external_config_dir. It
loads the resolved typed config, validates upstreams, and then stores the
masked registry snapshot. In Phase 4, light-gateway/gateway also registers a
ReloadableModule that reloads and validates gateway.yml, updates the masked
registry snapshot, and swaps the live GatewayConfig through ConfigManager.
light-deployer loads deployer.yml before the runtime is started, so it
creates a ModuleRegistry before loading its config, registers the final
env-overridden deployer config, and passes the same registry to
LightRuntimeBuilder.
light-agent also loads application configs before runtime startup. It now
registers ollama.yml and mcp-client.yml in the pre-runtime registry and
passes that registry into LightRuntimeBuilder. The existing manual
PortalRegistryClient setup is unchanged so the registry feature does not
reintroduce duplicate controller registration.
Current Registered Modules
Phase 4 registers these modules:
| Module ID | Config name | Kind | Reloadable |
|---|---|---|---|
light-runtime/startup | startup | core | no |
light-runtime/server | server | core | no |
light-client/client | client | core | no |
light-runtime/portal-registry | portal-registry | core | no |
light-gateway/gateway | gateway | application | yes |
light-deployer/deployer | deployer | application | no |
light-agent/ollama | ollama | application | no |
light-agent/mcp-client | mcp-client | application | no |
The application modules are visible in get_service_info once their owning
application loads them. get_modules returns the corresponding module ID
strings for portal-view selection. light-gateway/gateway can reload without a
restart. Other application modules keep reloadable=false until their runtime
state is moved behind swappable holders.
Rollout Plan
Phase 1: Registry and Masked Info
- Implemented:
ModuleRegistry,ModuleEntry, and mask utilities inlight-runtime. - Implemented: built-in runtime config registration.
- Implemented: tests proving raw secrets are not stored in registry entries.
- Implemented: Java-compatible server-info response assembly.
- Implemented: module-list response.
- Implemented: a
portal-registryMCP handler that exposesget_service_infoandget_modules.
Phase 2: Application Registration
- Implemented: convert
light-gateway/gatewayto registered config loading. - Implemented: convert
light-deployer/deployer. - Implemented: convert
light-agent/ollamaandlight-agent/mcp-client. - Implemented: add docs showing module IDs and reloadability.
Phase 3: Controller Operations
- Implemented: add MCP
tools/listandtools/callsupport forreload_modules. - Implemented: align portal-view calls so Java and Rust instances can be managed with the same control-plane workflow.
- Implemented: return Java-compatible
modulesstring lists while preserving detailedreloaded,skipped, andfailedreload result fields.
Phase 4: Hot Reload
- Implemented: add
ReloadableModule,ReloadContext, andReloadOutcome. - Implemented: add
ConfigManager<T>for swappable typed configs. - Implemented: implement reload for
light-gateway/gateway. - Implemented: add reload result tracking in the registry.
- Implemented: add tests for registry reload results, gateway live config swapping, and config-server-backed reload context refresh.
Open Questions
- Should module IDs be centrally reserved in
light-runtime, or should each application own its ID namespace? - Should the Java-compatible
componentmap include only active modules, whilemodulesincludes inactive-but-known modules? - Should MCP tool execution be enabled whenever
portal-registryis enabled, or guarded by a separate admin-tools flag? - Should
server.maskConfigProperties=falsebe allowed in production builds, or should Rust always mask known dangerous keys?
Implementation Sequence
Phase 1 implemented registry and masked server info first, without hot reload.
Phase 2 added application registration, so portal-view can display Rust
application modules next to Java modules once it calls the MCP tools through
portal-registry.
Phase 3 added the controller-facing reload_modules tool and Java-compatible
module ID lists.
Phase 4 added the first real hot reload implementation for
light-gateway/gateway. The next implementation step is to move additional
application configs, such as light-deployer/deployer,
light-agent/ollama, and light-agent/mcp-client, behind swappable runtime
state before marking them reloadable.
Module Hot Reload
This document describes the design and implementation of the hot reload mechanism in Light Fabric, explaining how modules reload configuration at runtime without requiring a full process restart.
Overview
In Light Fabric, certain configurations can be updated dynamically at runtime to support continuous delivery and quick configuration tuning (e.g., routing changes, CORS policies, security settings, or service discovery URLs). The system provides a unified Module Registry and Reloadable Modules architecture that allows the control plane (via MCP tools) to trigger config reloads.
Reload Flow
When the reload_modules MCP tool is invoked, the control plane initiates the following sequence:
sequenceDiagram
participant ControlPlane as Control Plane / Portal Registry
participant Handler as Runtime MCP Handler
participant Config as RuntimeConfig
participant Registry as Module Registry
participant Reloader as Module Reloaders
ControlPlane->>Handler: Call reload_modules
Handler->>Config: reload_context()
Note over Config: Re-fetch remote files,<br/>re-read local config yml,<br/>and build ReloadContext
Config-->>Handler: Return ReloadContext
Handler->>Registry: reload_modules(context, target_modules)
Note over Registry: Update direct-registry & client configs
loop For each reloadable module
Registry->>Reloader: reload(context)
Reloader->>Config: Load new file config
Reloader->>Registry: register_loaded_config(...)
Note over Reloader: Store fresh config in ConfigManager
end
Registry-->>Handler: Return ReloadModulesResult
Handler-->>ControlPlane: Return result JSON
- Build Reload Context: The runtime constructs a
ReloadContextcontaining a freshRuntimeConfigby parsing the updated config files (local or fetched from the config server) and merging dynamicvalues.ymlparameters. - Pre-update Built-in Configs: The core configurations stored inside
ModuleRegistry(such aslight-client/clientandlight-runtime/direct-registry) are updated in-memory using the reloaded config. - Dispatch to Module Reloaders: The registry iterates over the target modules and invokes the corresponding
ReloadableModule::reloadimplementation. - Atomic State Swap: Inside each reloader, the new configuration is parsed, validated, registered in the registry, and swapped atomically using
ConfigManager<T>.
ConfigManager and Thread Safety
To prevent request latency during reloads, Light Fabric uses a thread-safe ConfigManager<T> to manage dynamic configurations.
ConfigManager wraps an Arc<T> with a short-lived RwLock. Request handlers clone the Arc instantly (a simple reference count increment) without blocking, while the reloader replaces the entire Arc<T> atomically after the new configuration is successfully parsed and validated.
Core Hot Reload Implementations
Direct Registry Reload (light-runtime/direct-registry)
The direct registry maps service IDs to direct URLs for service discovery.
- Reload Process: The direct URLs are updated in
values.yml(either locally or on a remote config server). On reload,ReloadContextparses the new URLs, andreload_modulesupdates the registered config for"light-runtime/direct-registry"in theModuleRegistry. - Propagation: Runtimes such as the
McpRouterRuntimeor theTokenRuntimere-read the updateddirect_registryconfig from the freshRuntimeConfigwhen they are reloaded.
Client Configuration Reload (light-client/client)
The client configuration contains TLS settings and OAuth token provider configurations.
- Reloadable Flag: The client module is registered with
reloadable: trueat startup. - Reload Process: The
ReloadContextre-loadsclient.ymlfrom disk, applying new TLS properties or OAuth credentials. Thereload_modulesfunction updates the registered client configuration (applying proper masks to client secrets and certificates). - Reloader: A registered
ClientReloadermarks the transition success. Dependent modules (likelight-pingora/mcp-routerandlight-pingora/token) query the new client configuration from the context upon reload.
Reloadable vs. Non-Reloadable Configs
| Module ID | Config File | Type | Reloadable | Description |
|---|---|---|---|---|
light-runtime/startup | startup.yml | Core | No | Core server boot credentials |
light-runtime/server | server.yml | Core | No | Server host, IP, and listeners |
light-runtime/portal-registry | portal-registry.yml | Core | No | Connection to portal registry |
light-runtime/direct-registry | values.yml | Core | Yes | Service discovery direct URL overrides |
light-client/client | client.yml | Core | Yes | Outbound TLS and OAuth client credentials |
light-pingora/handler | handler.yml | Framework | Yes | Active handler chains and route mappings |
light-pingora/correlation | correlation.yml | Framework | Yes | Traceability and MDC logging settings |
light-pingora/cors | cors.yml | Framework | Yes | CORS origin and header limits |
light-pingora/mcp-router | mcp-router.yml | Framework | Yes | MCP server configurations and upstream rules |
light-pingora/token | token.yml | Framework | Yes | OAuth client credentials token handlers |
Note
Modifying non-reloadable configurations requires a full restart of the gateway process to bind new server listeners or configure registry websocket connections securely.
Verification & Testing
Module hot-reloading can be verified using the following automated test suites:
- Direct Registry Test:
reload_modules_updates_direct_registry_config(defined incrates/light-runtime/src/module_registry.rs) asserts that updated direct discovery URLs are correctly reflected in the registry. - Client Config Test:
gateway_client_config_reload(defined inapps/light-gateway/src/main.rs) asserts that updated TLS verification settings inclient.ymlare loaded and reflected in the registry.
Controller Registry Client
The Controller Registry Client (portal-registry) manages the connection between a gateway (or agent) instance and the Light Portal control plane. It enables runtime instance registration, service discovery queries, and dynamic configuration synchronization over a secure WebSocket connection.
Architecture Overview
The registry client operates as a background service inside the runtime. It establishes a persistent connection to the controller and handles bidirectional communication.
sequenceDiagram
participant Instance as Gateway / Agent
participant Client as Portal Registry Client
participant Controller as Control Plane / Portal
Instance->>Client: Initialize and run()
loop Connection Loop
Client->>Controller: WebSocket Handshake (WSS)
Note over Client,Controller: Negotiate TLS & Custom Certificates
alt Connection Succeeded
Client->>Controller: service/register (JSON-RPC Request)
Controller-->>Client: Registration Response (Instance ID)
Note over Client: State = Registered
loop Active Connection (tokio::select!)
alt Heartbeat interval (30s)
Client->>Controller: Ping
Controller-->>Client: Pong
end
alt Server Request
Controller->>Client: JSON-RPC Request / Notification
Client->>Controller: Response
end
end
else Connection Failed / Severed
Note over Client: State = Disconnected
Note over Client: Calculate backoff + jitter
Client->>Client: Sleep before retry
end
end
Core Features
1. WebSocket Protocol & Handshake
The connection is established over standard WebSocket (secured via TLS: wss://). Once connected, the client performs an initial JSON-RPC handshake:
- Method:
service/register - Parameters:
ServiceRegistrationParams(containingserviceId,version, hostaddress, listeningport,tags,envTag, and a verificationjwttoken). - Result:
RegistrationResponsereturning a uniqueruntimeInstanceIdassigned by the control plane.
2. Heartbeat (Ping/Pong)
To prevent network firewalls from dropping inactive connections and to detect silent TCP half-open connection drops, the client sends a WebSocket Ping frame every 30 seconds.
- If the control plane fails to reply, or the socket write fails, the connection loop is terminated immediately to initiate reconnection.
- The client also responds immediately with a
Pongto any inboundPingframes received from the controller.
3. Exponential Backoff with Jitter
When a connection is lost, terminated, or fails to initialize, the client retries using an exponential backoff strategy:
- Base delay starts at
1 secondand doubles on subsequent retries up to a maximum of60 seconds. - Random Jitter of
0-1000 millisecondsis added to each sleep duration. - Thundering Herd Prevention: Jitter prevents synchronized gateway instances (e.g., in a Kubernetes cluster) from flooding the control plane with connection requests at the exact same moment when it restarts.
4. TLS & Certificate Verification
The client supports establishing WSS connections with two cert verification modes:
- Standard Verification (
verifyHostname: true): Validates the server certificate chain against loaded CA certificates and verifies that the certificate hostname matches the controller domain. - No-Hostname Verification (
verifyHostname: false): Useful in local development or custom routing networks. It validates the certificate chain against the trusted CA bundle but bypasses hostname verification.
Component Configuration
Registry settings are loaded from portal-registry.yml or mapped in startup configuration:
| Configuration Property | Type | Default | Description |
|---|---|---|---|
portalUrl | String | The API endpoint of the Portal Registry controller | |
portalToken | String | JWT verification token used for handshakes | |
controllerDiscoveryToken | String | Token utilized for discovery lookups | |
bootstrapCaCertPath | Path | Optional path to CA certificate bundle |
Verification & Testing
The registry client behaves predictably under connection drops and can be verified via:
- Handshake Verification:
registration_and_metadata_update_match_controller_protocol(defined incrates/portal-registry/src/client.rs) asserts correct JSON-RPC registration format and success handling. - WebSocket Gateway integration:
websocket_gateway_proxies_text_binary_close_subprotocol_and_headers(defined inapps/light-gateway/src/main.rs) tests end-to-end WebSocket proxying alongside a mock controller registry. - Reconnect Loop Verification:
test_registry_client_reconnects_and_reregisters_on_run_level(defined incrates/portal-registry/src/client.rs) verifies that client terminates connection on socket drop, calculates backoff delay, reconnects, and re-registers automatically. - Heartbeat Timeout Verification:
test_heartbeat_timeout_detects_silent_controller_loss(defined incrates/portal-registry/src/client.rs) tests that the client detects silent connection drops by terminating and transitioning toDisconnectedif the controller does not respond to Ping within the configured heartbeat timeout window.
Cache Control Plane
Status: Proposed
Purpose
Light Fabric should expose the same cache operations through the portal control
plane that Java services expose through light-4j and portal-registry.
Today, portal-view can list caches and inspect cache entries for a running
service instance. The next required operation is clearing a cache so cached data
can be reloaded from its source of truth after operational data changes. A
common case is clearing the reference-data cache in portal-service after
reference tables are changed from light-portal.
The feature should be generic. It should not be a portal-service only endpoint.
Any Java or Rust service that registers with the controller and has named local
caches should be manageable through the same MCP tool contract.
Current Shape
The Java implementation already has most of the control-plane pieces:
light-4j/cache-managerdefines the genericCacheManagerAPI.light-4j/caffeine-cacheprovides the Caffeine-backed implementation.light-4j/portal-registryexposes MCP tools such aslist_cachesandget_cache_entries.controller-rsand the Java controller forward instance-specific MCP tool calls byruntimeInstanceId.portal-viewcalls the controller MCP websocket and passesruntimeInstanceIdfor cache exploration.
The main semantic gap is that CacheManager.removeCache(name) removes the cache
from the manager in the Caffeine implementation. For a control-plane clear
operation, the desired behavior is different: invalidate all entries while
keeping the configured cache alive so the next application read repopulates it.
Goals
- Add a generic whole-cache clear operation.
- Keep the control-plane contract compatible between Java services and Light Fabric services.
- Expose cache operations through
portal-registryand controller MCP routing, not through service-specific REST endpoints. - Let portal-view clear a selected cache from the existing Cache Explorer page.
- Use the same feature for
portal-servicereference data caching. - Preserve existing cache inspection behavior.
Non-Goals
- Do not remove or unregister a configured cache when clearing entries.
- Do not require every service to use the same cache backend.
- Do not expose raw secrets or unsafe object internals through cache inspection.
- Do not build event-driven cross-service cache invalidation in the first phase.
- Do not confuse runtime data caches with the
config-cachedirectory used for remote configuration files.
MCP Tool Contract
Add a new generic tool:
{
"name": "clear_cache",
"description": "Clear all entries from a named cache on a live runtime instance.",
"inputSchema": {
"type": "object",
"required": ["runtimeInstanceId", "name"],
"properties": {
"runtimeInstanceId": { "type": "string", "format": "uuid" },
"name": { "type": "string" }
}
}
}
The controller accepts runtimeInstanceId, removes it from the forwarded
arguments, and sends this to the target runtime:
{
"name": "clear_cache",
"arguments": {
"name": "reference-data"
}
}
Recommended success response:
{
"supported": true,
"status": "success",
"name": "reference-data",
"beforeSize": 42,
"afterSize": 0
}
Recommended unsupported response:
{
"supported": false,
"status": "unsupported",
"name": "reference-data",
"message": "Cache support is not available on this service."
}
Key-level invalidation can be added later as a separate
invalidate_cache_entry tool with { "name": "...", "key": "..." }.
Whole-cache clear should be implemented first because it solves the reference
data reload case without introducing cache-key UX and serialization questions.
Java Compatibility Work
In light-4j, add an explicit clear operation to the generic cache API:
void clear(String cacheName);
The Caffeine implementation should call cache.invalidateAll() and keep the
cache in the manager. It may call cache.cleanUp() before returning size data.
removeCache(name) should keep its existing unregister/remove semantics.
portal-registry should advertise clear_cache in tools/list and handle it
in tools/call by using CacheManager.getInstance(). The handler should
return supported: false when cache classes or a cache manager are not
available, matching the current list_caches and get_cache_entries behavior.
The controller catalogs need the same tool so portal-view can call it through the normal controller websocket:
controller-rstool catalog and command serialization- Java
light-controllertool catalog and routed-call handling, if it remains a supported control-plane runtime
Light Fabric Runtime Design
Light Fabric should provide a small cache abstraction at the runtime layer so applications do not each define a different operational surface.
A practical shape is:
#![allow(unused)]
fn main() {
#[async_trait::async_trait]
pub trait RuntimeCache: Send + Sync {
async fn len(&self) -> usize;
async fn entries_summary(&self) -> serde_json::Value;
async fn clear(&self);
}
#[derive(Default)]
pub struct CacheRegistry {
caches: RwLock<BTreeMap<String, Arc<dyn RuntimeCache>>>,
}
}
The registry should support:
- register named cache
- list cache names
- get summarized entries
- clear a named cache
moka is the preferred default backend for async Rust services because it maps
well to the Caffeine use case. Applications should still be free to register
custom cache wrappers as long as they implement the runtime trait.
RuntimeMcpHandler in light-runtime should expose the same tools as Java:
list_cachesget_cache_entriesclear_cache
If a runtime has no cache registry, these tools should return supported: false rather than failing the request.
Portal Service Reference Data Cache
portal-service can use the generic Light Fabric cache for /r/data.
Suggested cache names:
reference-datareference-data-relation
Suggested keys:
host:{hostId|global}:lang:{lang}:table:{name}host:{hostId|global}:lang:{lang}:table:{name}:rela:{rela}:from:{from}
The request flow becomes:
/r/datareceives a reference-data request.ReferenceServicebuilds a stable cache key from host, language, table, relation, and source value.- On cache hit, return cached reference data.
- On cache miss, query Postgres, cache the result, and return it.
- When reference data changes in
light-portal, an operator clearsreference-dataorreference-data-relationfor the targetportal-serviceruntime instance from portal-view. - The next
/r/datacall reloads from Postgres.
This keeps the first implementation manual and deterministic. A later phase can subscribe to reference-table change events and clear matching caches automatically.
Portal View UX
The existing Cache Explorer page should stay the main UI.
Add a clear action for the selected cache:
- show the selected cache name
- require confirmation before clearing
- disable the button while the request is running
- call
clear_cachewith{ runtimeInstanceId, name } - show success or error status
- refetch cache entries after a successful clear
The UI should not require users to know whether the target service is Java or Rust. Unsupported runtimes should show the returned unsupported message.
Implementation Phases
Phase 1: Java clear support
- Add
CacheManager.clear(cacheName). - Implement it in
caffeine-cache. - Add
clear_cachetoportal-registryMCP tools. - Add targeted tests for clearing while preserving the configured cache.
Phase 2: Controller and portal-view
- Add
clear_cacheto controller tool catalogs and command routing. - Add the Cache Explorer clear button and confirmation.
- Verify the existing
runtimeInstanceIdforwarding path is reused.
Phase 3: Light Fabric generic cache
- Add a runtime cache registry and trait.
- Add
mokabacked cache support. - Expose
list_caches,get_cache_entries, andclear_cachefromRuntimeMcpHandler. - Add focused
light-runtimetests for supported and unsupported cache cases.
Phase 4: Portal service reference data
- Register
reference-dataandreference-data-relationcaches. - Cache
/r/dataquery results. - Clear the cache from portal-view and verify the next request reloads from Postgres.
Verification
Recommended targeted checks:
mvn -q -pl cache-manager,caffeine-cache,portal-registry test
cargo test -p light-runtime
cargo check --workspace
yarn build
Use the Maven command in light-4j, the Cargo commands in light-fabric and
portal-service as appropriate, and the frontend build in portal-view.
Client Configuration And Modules
Status
Brainstorming proposal for standardizing client.yml across Light Fabric
runtime, framework modules, and products.
The immediate trigger is that different Rust modules currently interpret
client.yml differently. For example, light-runtime reads a small top-level
verifyHostname field for controller and config-server clients, while
light-pingora token and SPA modules read a Java-style nested tls section.
That split makes a single client.verifyHostname: false value unreliable.
This document proposes a common contract so every Rust module uses the same
client.yml file and the same typed configuration model.
Purpose
client.yml should describe outbound client behavior for a running service:
- TLS trust, hostname verification, and optional client identity.
- HTTP request timeout, retry, circuit breaker, connection pool, and HTTP/2 behavior.
- OAuth 2.0 token, key, sign, dereference, and provider-selection behavior.
- Path-prefix-to-service mapping used when different downstream services use different OAuth providers.
The file should be loaded once through the runtime configuration system, registered once in the module registry with secrets masked, then shared by all modules that make outbound calls.
Compatibility Contract
The Java light-4j client.yml remains the compatibility baseline. Rust can
clean up the internal model, but it should not remove behavior that Java
http-client and client-config expose.
Important Java sections:
tls:
verifyHostname: ${client.verifyHostname:true}
loadDefaultTrustStore: ${client.loadDefaultTrustStore:true}
loadTrustStore: ${client.loadTrustStore:true}
trustStore: ${client.trustStore:client.truststore}
trustStorePass: ${client.trustStorePass:password}
loadKeyStore: ${client.loadKeyStore:false}
keyStore: ${client.keyStore:client.keystore}
keyStorePass: ${client.keyStorePass:password}
keyPass: ${client.keyPass:password}
defaultCertPassword: ${client.defaultCertPassword:changeit}
tlsVersion: ${client.tlsVersion:TLSv1.3}
oauth:
multipleAuthServers: ${client.multipleAuthServers:false}
token:
cache:
capacity: ${client.tokenCacheCapacity:200}
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
expiredRefreshRetryDelay: ${client.expiredRefreshRetryDelay:2000}
earlyRefreshRetryDelay: ${client.earlyRefreshRetryDelay:4000}
server_url: ${client.tokenServerUrl:}
serviceId: ${client.tokenServiceId:com.networknt.oauth2-token-1.0.0}
proxyHost: ${client.tokenProxyHost:}
proxyPort: ${client.tokenProxyPort:}
enableHttp2: ${client.tokenEnableHttp2:true}
authorization_code: {}
client_credentials: {}
refresh_token: {}
token_exchange: {}
key: {}
sign: {}
deref: {}
pathPrefixServices: ${client.pathPrefixServices:}
request:
errorThreshold: ${client.errorThreshold:2}
connectTimeout: ${client.connectTimeout:2000}
timeout: ${client.timeout:3000}
resetTimeout: ${client.resetTimeout:7000}
injectOpenTracing: ${client.injectOpenTracing:false}
injectCallerId: ${client.injectCallerId:false}
enableHttp2: ${client.enableHttp2:true}
connectionPoolSize: ${client.connectionPoolSize:1000}
connectionExpireTime: ${client.connectionExpireTime:1800000}
maxReqPerConn: ${client.maxReqPerConn:1000000}
maxConnectionNumPerHost: ${client.maxConnectionNumPerHost:1000}
minConnectionNumPerHost: ${client.minConnectionNumPerHost:250}
maxRequestRetry: ${client.maxRequestRetry:3}
requestRetryDelay: ${client.requestRetryDelay:1000}
poolMetricsEnabled: ${client.poolMetricsEnabled:false}
poolWarmUpEnabled: ${client.poolWarmUpEnabled:false}
poolWarmUpSize: ${client.poolWarmUpSize:1}
healthCheckEnabled: ${client.healthCheckEnabled:true}
healthCheckIntervalMs: ${client.healthCheckIntervalMs:30000}
Rust should add fields such as tls.caCertPath, tls.clientCertPath, and
tls.clientKeyPath because PEM files are the native Rust deployment shape.
Rust does not need to support Java-specific JKS/JCEKS truststore or keystore
formats. If those Java-only fields appear in a Rust client.yml, they can be
ignored because config-server should control which fields it injects for Rust
services.
Initial Rust Gaps
At the start of this migration, the Rust implementation had three separate interpretations of client configuration:
| Area | Current behavior | Problem |
|---|---|---|
light-runtime config-server and portal-registry clients | Read ClientConfig { verify_hostname } from top-level client.yml | Did not understand the Java nested tls.verifyHostname shape |
light-pingora token, security JWKS, stateless auth, and MSAL exchange | Read ClientTokenConfig with tls, oauth, pathPrefixServices, and request | Was closer to Java, but framework-local and did not drive runtime clients |
light-gateway upstream proxy | Read the resolved flat value client.verifyHostname directly from values.yml | Bypassed typed client.yml and could disagree with other modules |
Before this design, Rust support was also partial compared with Java:
| Java capability | Initial Rust status |
|---|---|
tls.verifyHostname | Supported by Pingora token/SPAs, not by runtime controller/config-server clients |
| CA trust | Supported through Rust caCertPath; Java truststore fields are not modeled |
| Client certificate and key for mTLS | Not yet modeled for outbound clients |
| TLS version | Not yet modeled |
| Request connect and total timeout | Supported for token/SPAs |
| Retries, circuit breaker, pool sizing, pool health | Not yet modeled as shared client behavior |
OAuth authorization_code | Supported by SPA auth |
OAuth client_credentials | Supported by token handler |
OAuth refresh_token | Supported by SPA auth |
OAuth token_exchange | Supported by MSAL exchange and SPA auth |
OAuth token key / JWKS | Partially supported by security runtime |
token.key.serviceIdAuthServers and audience | Not fully modeled in Rust |
OAuth sign | Not yet modeled |
OAuth sign.key / sign JWKS | Not yet modeled |
OAuth deref | Not yet modeled |
| Multiple auth providers by service id | Supported for client credentials, but should become a shared resolver |
pathPrefixServices | Supported in token handler, but should become shared resolver logic |
Goals
- Keep
client.ymlas the only config file for outbound client behavior. - Make the Java nested shape canonical:
tls.verifyHostname, not top-levelverifyHostname. - Load and register the resolved
client.ymlonce throughlight-runtime. - Share one typed
ClientConfigacross runtime, Pingora, gateway, agent, deployer, MCP clients, model-provider clients, and future products. - Preserve Java-compatible field names and config-server placeholder names.
- Support direct URL, direct registry, and portal registry service discovery consistently for token, key, sign, deref, and generic outbound calls.
- Keep secrets masked in module registry snapshots and logs.
- Make invalid active client config fail startup or reject reload before it changes live runtime behavior.
- Allow Rust-native PEM fields without forcing Java keystore names into every Rust deployment.
Non-Goals
- Do not move handler activation into
client.yml. Handler-specific files such astoken.yml,statelessAuth.yml, andmsal-exchange.ymlstill decide whether a handler runs. - Do not implement every Java-only low-level connection-pool behavior in the first phase. The shared schema should include the fields so config is not lost, but unsupported fields can be ignored deliberately until the transport supports them.
- Do not expose decrypted client secrets, tokens, or legacy Java password fields through module registry, MCP tools, logs, metrics, or cache output.
- Do not require every module to use OAuth. The shared config must support simple TLS-only clients too.
Resolved Decisions
- Create a separate
light-clientcrate now so the shared config, HTTP client factory, OAuth client, and provider resolver can be reused without coupling every consumer tolight-runtime. - Standardize Rust outbound TLS material on PEM paths. Java truststore and keystore formats are not required for Rust services.
client.ymlreload should not force an immediate portal-registry reconnect. Reload is primarily for newly onboarded JWKS/JWT access and future outbound requests. Existing long-lived controller connections can keep running until their normal reconnect or service restart.- Unsupported Java fields can be ignored by Rust. Config-server should avoid injecting unsupported fields into Rust service config.
- Ignored Java-only fields should be ignored silently. Rust startup does not need to warn about fields that config-server may omit for Rust services.
oauth.multipleAuthServersremains accepted for Java compatibility, but Rust should infer multi-provider mode whenserviceIdAuthServersis configured.pathPrefixServicesstays inclient.yml. It is outbound-client provider selection and is different from inbound path routing to downstream services.- Circuit breaker behavior is only needed by Pingora. Shared request config can carry the Java-compatible fields, but non-Pingora clients do not need to own circuit breaker state.
- SAML bearer is not required for Light Fabric and should remain out of scope unless a future product explicitly needs it.
Proposed Canonical Shape
The canonical Rust client.yml should stay close to Java:
tls:
verifyHostname: ${client.verifyHostname:true}
caCertPath: ${client.caCertPath:}
clientCertPath: ${client.clientCertPath:}
clientKeyPath: ${client.clientKeyPath:}
tlsVersion: ${client.tlsVersion:TLSv1.3}
request:
connectTimeout: ${client.connectTimeout:2000}
timeout: ${client.timeout:3000}
maxRequestRetry: ${client.maxRequestRetry:3}
requestRetryDelay: ${client.requestRetryDelay:1000}
errorThreshold: ${client.errorThreshold:2}
resetTimeout: ${client.resetTimeout:7000}
injectCallerId: ${client.injectCallerId:false}
enableHttp2: ${client.enableHttp2:true}
connectionPoolSize: ${client.connectionPoolSize:1000}
connectionExpireTime: ${client.connectionExpireTime:1800000}
maxReqPerConn: ${client.maxReqPerConn:1000000}
maxConnectionNumPerHost: ${client.maxConnectionNumPerHost:1000}
minConnectionNumPerHost: ${client.minConnectionNumPerHost:250}
poolMetricsEnabled: ${client.poolMetricsEnabled:false}
poolWarmUpEnabled: ${client.poolWarmUpEnabled:false}
poolWarmUpSize: ${client.poolWarmUpSize:1}
healthCheckEnabled: ${client.healthCheckEnabled:true}
healthCheckIntervalMs: ${client.healthCheckIntervalMs:30000}
oauth:
multipleAuthServers: ${client.multipleAuthServers:false}
token:
cache:
capacity: ${client.tokenCacheCapacity:200}
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
expiredRefreshRetryDelay: ${client.expiredRefreshRetryDelay:2000}
earlyRefreshRetryDelay: ${client.earlyRefreshRetryDelay:4000}
server_url: ${client.tokenServerUrl:}
serviceId: ${client.tokenServiceId:com.networknt.oauth2-token-1.0.0}
proxyHost: ${client.tokenProxyHost:}
proxyPort: ${client.tokenProxyPort:}
enableHttp2: ${client.tokenEnableHttp2:true}
authorization_code:
uri: ${client.tokenAcUri:/oauth2/token}
client_id: ${client.tokenAcClientId:}
client_secret: ${client.tokenAcClientSecret:}
redirect_uri: ${client.tokenAcRedirectUri:}
scope: ${client.tokenAcScope:}
client_credentials:
uri: ${client.tokenCcUri:/oauth2/token}
client_id: ${client.tokenCcClientId:}
client_secret: ${client.tokenCcClientSecret:}
scope: ${client.tokenCcScope:}
serviceIdAuthServers: ${client.tokenCcServiceIdAuthServers:}
refresh_token:
uri: ${client.tokenRtUri:/oauth2/token}
client_id: ${client.tokenRtClientId:}
client_secret: ${client.tokenRtClientSecret:}
scope: ${client.tokenRtScope:}
token_exchange:
uri: ${client.tokenExUri:/oauth2/token}
client_id: ${client.tokenExClientId:}
client_secret: ${client.tokenExClientSecret:}
scope: ${client.tokenExScope:}
subjectToken: ${client.subjectToken:}
subjectTokenType: ${client.subjectTokenType:urn:ietf:params:oauth:token-type:jwt}
requestedTokenType: ${client.requestedTokenType:}
audience: ${client.tokenExAudience:}
key:
server_url: ${client.tokenKeyServerUrl:}
serviceId: ${client.tokenKeyServiceId:com.networknt.oauth2-key-1.0.0}
uri: ${client.tokenKeyUri:/oauth2/key}
client_id: ${client.tokenKeyClientId:}
client_secret: ${client.tokenKeyClientSecret:}
enableHttp2: ${client.tokenKeyEnableHttp2:true}
serviceIdAuthServers: ${client.tokenKeyServiceIdAuthServers:}
audience: ${client.tokenKeyAudience:}
sign:
server_url: ${client.signServerUrl:}
serviceId: ${client.signServiceId:com.networknt.oauth2-token-1.0.0}
uri: ${client.signUri:/oauth2/sign}
timeout: ${client.signTimeout:2000}
client_id: ${client.signClientId:}
client_secret: ${client.signClientSecret:}
proxyHost: ${client.signProxyHost:}
proxyPort: ${client.signProxyPort:}
enableHttp2: ${client.signEnableHttp2:true}
key:
server_url: ${client.signKeyServerUrl:}
serviceId: ${client.signKeyServiceId:com.networknt.oauth2-key-1.0.0}
uri: ${client.signKeyUri:/oauth2/key}
client_id: ${client.signKeyClientId:}
client_secret: ${client.signKeyClientSecret:}
enableHttp2: ${client.signKeyEnableHttp2:true}
audience: ${client.signKeyAudience:}
deref:
server_url: ${client.derefServerUrl:}
serviceId: ${client.derefServiceId:com.networknt.oauth2-token-1.0.0}
uri: ${client.derefUri:/oauth2/deref}
client_id: ${client.derefClientId:}
client_secret: ${client.derefClientSecret:}
proxyHost: ${client.derefProxyHost:}
proxyPort: ${client.derefProxyPort:}
enableHttp2: ${client.derefEnableHttp2:true}
pathPrefixServices: ${client.pathPrefixServices:}
Compatibility aliases:
- Accept
serverUrlin addition to Javaserver_urlfor Rust callers. - Accept
clientIdandclientSecretin addition to Javaclient_idandclient_secretonly as aliases. The emitted template should keep Java names. - Temporarily accept top-level
verifyHostnameonly as a migration fallback, but register a warning and normalize it intotls.verifyHostname.
Serde strategy for the top-level verifyHostname fallback:
- The shared
ClientConfigshould deserialize into a struct that has atls.verifyHostnamefield and a separate#[serde(default)]top-levelverify_hostnamefield. - After deserialization, a post-parse normalization step should check whether
the top-level field was explicitly set. If so, it logs a deprecation warning
and copies the value into
tls.verify_hostnameonly when the nested field was not also explicitly set. - When both the top-level and nested fields are present, the nested
tls.verifyHostnamevalue wins. The top-level value is ignored after the warning. - Do not rely on two competing
#[serde(default)]fields resolving the conflict. Use a customDeserializeimpl or an explicit post-parse step.
Serde strategy for Java-compatible but unimplemented sections:
- Do not use
#[serde(deny_unknown_fields)]for the top-levelClientConfigor OAuth section during Phase 1. - Known but not-yet-implemented Java sections such as
oauth.signandoauth.derefshould deserialize into typed structs orserde_json::Valueplaceholders so representative Java fixtures load successfully. - Demand-driven validation decides whether a section is required. If no active
module consumes
oauth.signoroauth.deref, those sections can be present and ignored silently.
Proposed Rust Modules
Shared Config Model
Create one shared typed config model outside light-pingora and
light-runtime:
crates/light-client/src/lib.rs
crates/light-client/src/config.rs
crates/light-client/src/http.rs
crates/light-client/src/oauth.rs
crates/light-client/src/provider.rs
light-runtime should use light-client for loading, validating, and building
outbound clients, but the reusable client model should not live inside the
runtime crate.
Core types:
#![allow(unused)]
fn main() {
pub struct ClientConfig {
pub tls: ClientTlsConfig,
pub request: ClientRequestConfig,
pub oauth: ClientOauthConfig,
pub path_prefix_services: BTreeMap<String, String>,
}
pub struct ClientTlsConfig {
pub verify_hostname: bool,
pub ca_cert_path: Option<PathBuf>,
pub client_cert_path: Option<PathBuf>,
pub client_key_path: Option<PathBuf>,
pub tls_version: Option<TlsVersion>,
}
pub struct ClientRequestConfig {
pub connect_timeout_ms: u64,
pub timeout_ms: u64,
pub max_request_retry: u32,
pub request_retry_delay_ms: u64,
pub error_threshold: u32,
pub reset_timeout_ms: u64,
pub inject_caller_id: bool,
pub enable_http2: bool,
pub pool: ClientPoolConfig,
}
}
TlsVersion should be an enum with serde names for Java-compatible strings
such as TLSv1.2 and TLSv1.3, rather than a raw string in runtime code.
Secrets should use a type that serializes as masked data for registry output, or the registry masks should cover every secret field recursively.
Runtime Loader
light-runtime should own the startup lifecycle for client.yml loading, but
delegate parsing and validation to light-client:
- Load local
values.yml. - Load local
startup.yml. - Load local
client.ymlwith resolved values for config-server bootstrap. - Fetch remote config if configured.
- Rebuild the final
RuntimeConfigwith the remoteclient.ymloverlay. - Register masked
light-client/clientinModuleRegistry.
Every runtime client should use this shared config:
- config-server fetch client
- portal-registry WebSocket client
- MCP client
- future model-provider outbound clients
- framework/application clients through
RuntimeConfig.client
For the earlier hostname-verification bug, the controller client should read:
runtime_config.client.tls.verify_hostname
not a separate top-level ClientConfig.verify_hostname.
HTTP Client Factory
Add a small factory that converts ClientConfig plus optional per-endpoint
overrides into concrete clients:
#![allow(unused)]
fn main() {
pub struct ClientFactory {
config: Arc<ClientConfig>,
direct_registry: DirectRegistryConfig,
registry_client: Option<Arc<PortalRegistryClient>>,
}
pub struct EndpointOptions {
pub server_url: Option<String>,
pub service_id: Option<String>,
pub proxy_host: Option<String>,
pub proxy_port: Option<u16>,
pub enable_http2: Option<bool>,
pub timeout_ms: Option<u64>,
}
}
Responsibilities:
- Build
reqwest::Clientwith consistent TLS, timeout, proxy, HTTP/2, retry, and pool settings for non-Pingora consumers. - Build Pingora
HttpPeeroptions from the same TLS config for gateway upstream proxying. - Resolve endpoint base URL by priority:
- direct
server_url direct-registry.yml- portal-registry discovery by
serviceId
- direct
- Apply per-service
AuthServerConfigoverrides without duplicating resolver logic in each handler.
The config-server bootstrap path still starts from BootstrapConfig because it
needs enough client settings before remote client.yml has been fetched. To
keep light-client independent from light-runtime, the factory should not
take a BootstrapConfig type directly. Instead, light-runtime should adapt
BootstrapConfig.connect_timeout, BootstrapConfig.timeout, authorization,
and bootstrap CA path into EndpointOptions or a small bootstrap options type
owned by light-client.
OAuth Client
Add a shared OAuth client module that implements Java http-client behavior:
oauth/client_credentials
oauth/authorization_code
oauth/refresh_token
oauth/token_exchange
oauth/key
oauth/sign
oauth/deref
The existing light-pingora SpaTokenClient, token handler client
credentials code, and security JWKS fetcher should delegate to this shared
module. Handler modules still own request-path decisions, cookies, headers,
and rejection mapping.
OAuth provider selection should be one reusable resolver:
#![allow(unused)]
fn main() {
pub struct OAuthProviderResolver {
client: Arc<ClientConfig>,
}
impl OAuthProviderResolver {
pub fn service_for_path(&self, path: &str) -> Option<&str>;
pub fn client_credentials_provider(&self, service_id: Option<&str>) -> Result<AuthServerConfig>;
pub fn key_provider(&self, service_id: Option<&str>) -> Result<AuthServerConfig>;
}
}
Rules:
- Single-provider mode uses global
oauth.token.*defaults. - Multi-provider mode is enabled when
oauth.multipleAuthServers: trueor when relevantserviceIdAuthServersmaps are non-empty. - Multi-provider mode selects the service id from an explicit request header
first, then outbound
pathPrefixServices. client_credentials.serviceIdAuthServers[serviceId]selects the token provider.key.serviceIdAuthServers[serviceId]selects the JWKS/key provider.- Per-service config inherits unset values from global
oauth.tokendefaults. - Path-prefix matching should be boundary-aware in Rust. Java uses
startsWith; the Rust implementation can be stricter as an intentional improvement. Exact rule: a prefix matches when the request path equals the prefix or starts withprefix + "/". Therefore/apimatches/apiand/api/orders, but does not match/api-v2. pathPrefixServicesis not an inbound routing table. It maps outbound request paths to service ids only for client-side OAuth provider selection.
Consumer Modules
All modules should consume the same shared config:
| Module | Uses |
|---|---|
light-runtime/config-server | light-client tls, request |
light-runtime/portal-registry | light-client tls, request |
light-pingora/security | oauth.token.key, tls, request, provider resolver |
light-pingora/token | oauth.token.client_credentials, token cache settings, provider resolver |
light-pingora/stateless-auth | authorization_code, refresh_token, token client |
light-pingora/msal-exchange | token_exchange, token client |
light-gateway/proxy | tls.verifyHostname, PEM mTLS, request timeout, retry, circuit breaker, and pool settings where Pingora supports them |
light-agent | controller/MCP outbound clients |
light-deployer | controller/MCP/outbound clients as needed |
Reload Behavior
client.yml should be reloadable as a module, but reload must be conservative:
- Load and validate the new config into a fresh
ClientConfig. - Build new shared client factories and OAuth clients.
- Swap the config atomically for future requests.
- Clear OAuth token caches because client credentials, scopes, providers, or trust settings may have changed.
- Keep old in-flight requests on their existing client instances.
- Reject the reload if active modules cannot build required clients from the new config.
Reload atomicity: all runtimes that consume client.yml must be swapped
together in the same reload callback. Today, the gateway TokenReloader
already rebuilds token_runtime, stateless_auth, and msal_exchange as a
unit. This must remain a hard requirement. A reload that updates the client
config without also rebuilding dependent runtimes would leave stale TLS or
OAuth state in the old runtime instances.
Controller registration is long-lived. Reloading client.yml should not force
an immediate portal-registry reconnect. New TLS and request settings should
apply to future outbound clients and the next normal controller reconnect, but
the active controller WebSocket can remain open.
Validation Rules
Base validation:
tls.verifyHostname: falserequires explicit trust material unless the transport has a clear dev-only mode.- If Rust-native mTLS is configured, both client certificate and client key paths are required.
request.connectTimeoutandrequest.timeoutmust be positive.proxyPortmust be 0 to 65535.pathPrefixServiceskeys must start with/.- Secret fields may be empty only when the consuming active module does not need that grant.
OAuth validation should be demand-driven:
- If
tokenhandler is active and enabled, validateclient_credentials. - If
stateless-authis active, validateauthorization_codeandrefresh_token. - If
msal-exchangeis active, validatetoken_exchange. - If
security.ymlenables JWKS bootstrap from key service, validateoauth.token.key. - If a future sign module is active, validate
oauth.sign. - If a future deref module is active, validate
oauth.deref.
This avoids forcing every service to configure every Java OAuth section.
Validation failure behavior:
- At startup, validation failures are fatal. The process must exit with a clear error message identifying which active module requires which missing or invalid client config section.
- On reload, validation failures are non-fatal. The reload is rejected, the old config stays live, and the rejection reason is logged and reported through the module registry reload outcome.
Masking
Mask these fields recursively in registry output:
client_secretclientSecrettrustStorePasskeyStorePasskeyPassdefaultCertPasswordsubjectTokenaccess_tokenrefresh_tokenid_tokenauthorization- any field ending in
Tokenwhose value is a scalar string (not a nested object, list, or URN-typed field likesubjectTokenTypeorrequestedTokenType) - any field ending in
Secret
Explicit exclusions from suffix matching:
subjectTokenType- a URN string, not a secret.requestedTokenType- a URN string, not a secret.
The registry should store only the masked snapshot. It should not store raw config and mask later.
Migration Plan
Phase 0: Deprecation Logging
- Add a
tracing::warn!inlight-gatewaywhere it readsresolved_values["client.verifyHostname"]to alert operators that this path is deprecated and will be replaced byruntime_config.client.tls.verify_hostname. - This gives operators visibility into the migration before behavior changes.
Phase 1: Unify The Schema
- Add the
light-clientcrate with the full sharedClientConfigtype. - Make
light-runtimeload nestedtls.verifyHostname. - Keep top-level
verifyHostnameas a temporary compatibility fallback. - Update Rust config templates to include only the canonical nested shape.
- Add tests proving
client.verifyHostname: falsereaches config-server, portal-registry, token, security JWKS, SPA auth, and gateway proxy clients.
Phase 2: Move Consumers To Shared Config
- Replace
light-pingora::token::ClientTokenConfigwith thelight-clientshared type or a type alias. - Replace gateway direct
resolved_values["client.verifyHostname"]lookup withruntime_config.client.tls.verify_hostname. - Move JWKS, token, and SPA token HTTP client construction behind the shared client factory.
- Register one masked
light-client/clientmodule instead of separate partial client registry entries.
Phase 3: Shared OAuth Provider Resolver
- Extract provider selection from the token handler.
- Support
token.key.serviceIdAuthServersandaudience. - Use the same resolver for token injection and JWT key lookup.
- Keep Java field names and config-server placeholders.
Phase 4: Java Feature Completion
- Implemented sign client support in
light-client. - Implemented deref client support in
light-client. - Implemented Rust-native PEM mTLS for reqwest clients and Pingora upstreams.
- Implemented retry, circuit breaker, and pool behavior where the Rust transport supports them.
Open Questions
None at this stage.
Test Plan
Unit tests:
- Parse the Java
client.ymltemplate into the shared Rust config. - Parse the current Rust
client.ymltemplate into the shared Rust config. - Resolve
client.verifyHostnameintotls.verifyHostname. - Accept top-level
verifyHostnameonly as a fallback and prefer nested TLS when both are set. - Mask every secret field in the module registry snapshot.
- Validate provider selection by service id and path prefix.
- Validate per-service override inheritance for token and key providers.
Runtime tests:
- Config-server bootstrap uses
tls.verifyHostname. - Portal-registry controller WebSocket uses
tls.verifyHostname. - Gateway upstream proxy uses
tls.verifyHostname. - Token handler, stateless auth, MSAL exchange, and security JWKS all receive
the same
ClientConfiginstance or snapshot. - Client reload clears token caches and rejects invalid active grant config.
- Reload round-trip: verify that reloading from config A to config B swaps the
ClientConfig, creates fresh token caches, and that in-flight requests on the old config are not affected. Verify that a reload from valid config to invalid config is rejected and the old config stays live.
Compatibility tests:
- Reuse representative Java
client.ymlfixtures for single provider, multiple providers, proxy, token key, sign, and deref sections. - Confirm Java-compatible form bodies for
authorization_code,client_credentials,refresh_token, andtoken_exchange. - Confirm config-server injected YAML strings and structured YAML maps both
deserialize for
serviceIdAuthServersandpathPrefixServices.
Embedded Configuration Templates
Status
Initial implementation completed. Rust applications in light-fabric and
related portal-service applications keep template configuration files under
each app’s config directory. Container images may copy those files into
/app/config-defaults, then runtime overlays local config, downloaded
config-cache, remote values.yml, and environment variables.
That works well for container deployments. It is awkward for native binary
deployments on a VM because the operator must copy a full template directory
beside the binary even when they only want to provide values.yml, certs, or a
small local override.
This design embeds the template files into the Rust binary while keeping the
app config directories in source control as the readable template source.
Purpose
Embedded configuration templates should make the Rust deployment model match the Java module model more closely:
- The application binary carries its default template files.
- Operators provide only overrides, usually
values.yml,startup.yml, certs, keys, or environment variables. - Config-server can still return
values.ymlafter bootstrap, plus external files for explicit migration or operational exceptions. - Developers and operators can still inspect the app’s
configdirectory in source control to learn supported properties.
The embedded files are defaults. They are not runtime state and should not be written out automatically unless an explicit diagnostic/export command is added later.
Current Model
The current runtime model has these filesystem layers:
| Layer | Example | Purpose |
|---|---|---|
| Default templates | config-defaults/server.yml | App-provided templates copied into the container image |
| Local config | config/values.yml, config/startup.yml | Operator overrides and bootstrap inputs |
| External/cache config | config-cache/values.yml | Files downloaded from config-server |
| Remote values | config-server response body | Runtime values fetched during bootstrap |
| Environment variables | CLIENT_VERIFYHOSTNAME=false | Last-mile process overrides during placeholder expansion |
For light-fabric runtime applications, LightRuntimeBuilder passes
default_config_dir, config_dir, and external_config_dir into
light-runtime. load_bootstrap_config() reads bootstrap-time values.yml,
startup.yml, and client.yml before remote config-server bootstrap. After
remote bootstrap, runtime config loads server.yml, client.yml,
portal-registry.yml, and framework/application module files through the same
merged configuration path.
Some portal-service apps share the light-runtime path, while standalone apps
such as config-server and light-oauth have local helper functions that merge
config-defaults and config.
Goals
- Allow a native binary deployment to start with embedded templates and a small
external
config/values.yml. - Keep
apps/<app>/config/*.ymlas the source of truth for template content. - Keep container deployment behavior compatible with the current
/app/config-defaultscopy. - Preserve the existing overlay order and placeholder expansion behavior.
- Support bootstrap-time files such as
startup.ymlandclient.yml. - Support runtime module files such as
handler.yml,proxy.yml,model-provider.yml, provider configs, and product-specific files. - Provide one reusable loading abstraction for
light-fabricandportal-serviceinstead of app-specific parsing logic. - Avoid writing embedded templates to disk during normal startup.
Non-Goals
- Do not embed secrets, certificates, private keys, trust bundles, static web assets, or downloaded config-server files.
- Do not remove the source
configdirectories. They remain the reviewable, documented template source. - Do not make
values.ymlmandatory. Apps should keep current defaults where they are already valid. - Do not make config-server responsible for delivering template files that are already part of the binary.
- Do not change the meaning of
values.ymlplaceholders or environment variable expansion.
Proposed Layer Order
The new effective source order should be:
- Embedded template file from the binary.
- Filesystem default template from
config-defaults, if present. - Local operator file from
config. - External/cache file from
config-cache, when runtime loading supports it. - Remote
values.ymlpayload from config-server. - Environment variables during placeholder resolution.
This keeps existing container images compatible. If config-defaults exists, it
can override the embedded template. That gives operators and image builders a
transition path and a deliberate escape hatch for patched images.
For native binary deployment, config-defaults is simply absent and the binary
falls back to embedded templates.
Structured config files and values.yml should use different overlay
semantics:
| File type | Semantics | Reason |
|---|---|---|
Structured config files such as server.yml, handler.yml, proxy.yml, and model-provider.yml | Source-level override. The highest-priority source that contains the file supplies the whole template. | Avoids surprising hybrid files assembled from embedded, image, local, and cache layers. Operators should use values.yml for partial property overrides. |
values.yml | Key-level overlay in source order, followed by remote values and environment variables. | values.yml is explicitly the property override surface. Partial overlays are expected and useful. |
After the structured file source is selected, placeholders in that file are resolved from the merged values map and environment variables.
Embedded Template Representation
include_dir is a possible embedding mechanism. It embeds the entire app
config directory at compile time and avoids custom directory-scanning build
scripts in every application crate:
#![allow(unused)]
fn main() {
use include_dir::{include_dir, Dir};
pub static EMBEDDED_CONFIG: Dir<'_> = include_dir!("$CARGO_MANIFEST_DIR/config");
}
The runtime should hide the concrete embedding mechanism behind a small config source abstraction. A typed file representation is still useful as the stable runtime boundary:
#![allow(unused)]
fn main() {
pub struct EmbeddedConfigFile {
pub name: &'static str,
pub content: &'static str,
}
}
Application code should pass a flattened static file list into the runtime:
#![allow(unused)]
fn main() {
LightRuntimeBuilder::new(transport)
.with_embedded_config(embedded_config::FILES)
.build();
}
include_str! is still acceptable for one or two files, but application
main.rs files should not accumulate hand-maintained include_str! lists.
include_bytes! is not preferred for YAML templates because configuration
templates should be valid UTF-8 before they are parsed.
The initial implementation uses a shared build-time generator instead of adding
an external embedding dependency. Each app has a small build.rs that calls
config-embed-build, which scans the committed config directory and produces
a manifest like this under OUT_DIR:
#![allow(unused)]
fn main() {
pub const FILES: &[config_loader::EmbeddedConfigFile] = &[
config_loader::EmbeddedConfigFile {
name: "server.yml",
content: include_str!(concat!(env!("CARGO_MANIFEST_DIR"), "/config/server.yml")),
},
config_loader::EmbeddedConfigFile {
name: "startup.yml",
content: include_str!(concat!(env!("CARGO_MANIFEST_DIR"), "/config/startup.yml")),
},
];
}
Build-Time Generation Fallback
The project currently uses the build-time manifest path. Each app uses a shared
build.rs helper to scan its config directory and generate the embedded
manifest. The generator lives in one reusable crate so apps do not carry
duplicated build logic.
The generated manifest should:
- Include only known text config extensions, initially
.yml,.yaml,.json, and.toml. - Preserve the file name relative to the app
configdirectory. - Emit
cargo:rerun-if-changed=config. - Fail the build if a template file cannot be read as UTF-8.
Nested config paths are not needed for current app templates, but the manifest
should allow names such as oauth/server.yml if a future product needs them.
Runtime API
Add embedded defaults to LightRuntimeBuilder:
#![allow(unused)]
fn main() {
LightRuntimeBuilder::new(transport)
.with_embedded_config(embedded_config::FILES)
.with_default_config_dir(DEFAULT_CONFIG_DIR)
.with_config_dir(CONFIG_DIR)
.with_external_config_dir(EXTERNAL_CONFIG_DIR)
.build();
}
RuntimeConfig should carry the embedded source as skipped runtime state, the
same way it carries default_config_dir and registries today:
#![allow(unused)]
fn main() {
pub struct RuntimeConfig {
// existing fields
#[serde(skip, default)]
pub embedded_config: &'static [EmbeddedConfigFile],
}
}
The stable contract is lookup by relative file name and iteration for diagnostics or dumping. The concrete representation can remain a static file slice or later move behind a provider abstraction if needed.
The low-level loader should accept named in-memory content as another config source:
#![allow(unused)]
fn main() {
pub enum ConfigSource {
Embedded { name: &'static str, content: &'static str },
File(PathBuf),
}
}
ConfigLoader can then parse embedded and filesystem sources with the same
YAML/JSON/TOML parser. Structured config loading should select the highest
priority source for the requested file. values.yml loading should continue to
merge maps in source order.
Bootstrap Behavior
Bootstrap must support embedded templates because this is the path that native deployments need most.
load_bootstrap_values() should merge:
- Embedded
values.yml, if present. config-defaults/values.yml, if present.config/values.yml, if present.
load_bootstrap_config() should load startup.yml and client.yml from:
- Embedded templates.
config-defaults.config.
For startup.yml and client.yml, the highest-priority source that contains
the file should be used as the full template. Placeholder resolution still uses
the merged bootstrap values.
After bootstrap fetches remote values, load_values_map() should merge embedded
values.yml before the existing file and remote layers. This allows remote
values to override embedded placeholders exactly as they override copied
template files today.
Application Integration
Light-Gateway
light-gateway should be the first light-fabric application to adopt the
runtime API because it has the richest template set:
- bootstrap and server files
- client and portal registry files
- handler chain files
- proxy, resource, MCP, websocket, auth, token, metrics, and rule-related files
After integration, a native gateway deployment can run with the binary plus a
small config/values.yml and any required cert/key files.
Light-Agent
light-agent uses the same runtime API for agent.yml and mcp-client.yml.
agent.yml is the typed immutable Agent-audience projection whose placeholders
are resolved only after Config Server values have been merged. Its model is an
llm-gateway Alias; Light-Agent does not embed or load provider-specific
templates such as openai.yml, codex.yml, or anthropic.yml.
Light-Deployer
light-deployer currently has a separate app-level config load for
deployer.yml. It should either move to the shared embedded-source helper or
set embedded defaults on LightRuntimeBuilder and use the same merged source
logic for its application config.
Portal-Service App
portal-service/apps/portal-service already uses LightRuntimeBuilder, but it
loads portal-service.yml before runtime startup to create the database pool.
That pre-runtime load should use the same shared embedded-source helper.
The portal-service.yml config remains non-reloadable because dbUrl and
hostId feed process-owned state.
Portal-Service Config-Server And Light-OAuth
portal-service/apps/config-server and apps/light-oauth do not bootstrap from
config-server. They should still embed their server.yml templates so native
deployment does not require a copied config-defaults directory.
Because these apps have local merge helpers today, they should consume a shared
config-loader helper that can merge:
- Embedded defaults.
- Filesystem defaults.
- Local config.
This keeps their behavior aligned with light-runtime without requiring them
to become runtime-bootstrap applications.
Operator Model
Shared client template
crates/light-client/config/client.yml is the canonical Rust client template.
It projects TLS, request, OAuth token/JWK, signing, dereference, and routing
settings into the shared ClientConfig model. Product and mounted deployment
copies are registered in scripts/client-config-targets.json.
From the light-fabric checkout, synchronize sibling repositories with:
python3 scripts/sync-client-config.py
python3 scripts/sync-client-config.py --check
cargo test -p light-client --test config_template
Use --workspace /path/to/workspace for a different sibling-checkout root.
Missing repositories are reported as not checked; missing files in a present
repository and unregistered product templates fail the check. CI checks the
copies available in its checkout. Before a cross-repository release, run the
check with all repositories in the manifest present.
The manifest preserves existing product/environment fallback differences (local TLS hostname verification, request timeout, retry delay, and OAuth/CA locations). All copies expose the same keys. Deployment values and Config Server values continue to override these fallbacks. Prefer values.yml for new deployment settings instead of introducing another fallback override.
A mounted client.yml replaces the entire embedded template. A TLS-only copy
therefore hides OAuth/JWK settings even when Config Server supplies
client.tokenKeyServerUrl and client.tokenKeyUri. Deploy the complete synced
template, then restart the affected process to reload it. Embedded copies need
an application rebuild. Synchronizing source files does not restart running
containers or verify authenticated chat end to end.
The obsolete A2A top-level connectTimeout, requestTimeout, and
maxIdlePerHost fields were ignored by the shared client model. Its synced
template uses the supported nested request fields and shared defaults.
Standalone services that do not load ClientConfig, and Java client.yml
templates, are outside this manifest.
For a native deployment, the recommended layout becomes:
/opt/light-gateway/
light-gateway
config/
values.yml
startup.yml # optional, only when values/env defaults are not enough
cert.pem # optional external asset
key.pem # optional external asset
The operator no longer needs to copy every template file beside the binary. They only provide files that are deployment-specific.
For a container deployment, the current layout continues to work:
/app/light-gateway
/app/config-defaults/*.yml
/config/values.yml
/app/config-cache/values.yml
In the long term, the /app/config-defaults copy can become optional. Keeping it
during migration is useful because it lets operators inspect templates inside
the image and provides a familiar override layer.
After embedded templates are stable across production deployments, Docker images
should deprecate and then remove the unconditional /app/config-defaults copy.
Template inspectability should move to explicit dump/print commands rather than
extra image layers.
Diagnostics
The runtime should expose enough information to make source precedence clear:
- Log whether embedded templates were registered for the application.
- When a required config file is missing, include the searched source names:
embedded,
config-defaults,config, andconfig-cache. - Module registry snapshots should show the resolved config, not the raw embedded template.
- Module registry metadata should include config source provenance when
available, for example
embedded,file:/app/config-defaults/server.yml, orfile:/config/server.yml. - Normal startup should not write embedded templates to disk.
Native operators should have explicit inspection commands:
light-gateway --print-default-config server.yml
light-gateway --dump-default-configs ./config-defaults
The print command writes one embedded template to stdout. The dump command writes all embedded templates to a target directory so operators can inspect, copy, and customize them.
Controller Server Info Compatibility
Rust services register with the controller, and the controller can call the runtime MCP service-info path to inspect runtime configuration. This behavior must continue to work with embedded templates.
The service-info response should expose resolved runtime configuration, not raw templates. The implementation contract is:
- Select the effective structured config source, such as embedded
server.yml, filesystemconfig/server.yml, or cachedconfig-cacheserver.yml. - Build the merged values map from embedded, filesystem, cached, remote
values.yml, and environment variables. - Resolve placeholders in the selected config source.
- Deserialize the resolved config into the typed runtime or module config.
- Register that typed config in
ModuleRegistry. - Return
ModuleRegistrycomponent configs from the controller service-info MCP call.
With that flow, the controller still sees every registered config file with defaults and overrides applied. Embedded templates only replace the missing filesystem default-template layer. They should not bypass typed config loading, masking, module registration, reload validation, or service-info reporting.
Source provenance can be added as metadata beside each registered config, but it must not replace the resolved config payload that operators and the controller depend on.
Testing Strategy
Add unit tests at the shared loader boundary:
- Embedded-only
server.ymlloads successfully. - Local
config/server.ymlreplaces embeddedserver.ymlrather than deep merging with it. config-defaults/server.ymlreplaces embeddedserver.yml.config-cache/server.ymlreplaces local config during runtime loads.- Embedded
values.ymlis overridden by localvalues.yml. - Remote
values.ymloverrides embedded and filesystem values. - Missing required config reports all searched layers.
- Source provenance is recorded for resolved module configs.
--print-default-configand--dump-default-configsexpose embedded templates without changing normal startup behavior.- Controller service-info output includes resolved values from embedded defaults plus local, cached, remote, and environment overrides.
Add application-level smoke tests for:
light-gatewaystartup with no filesystemserver.yml, using embedded templates plus localvalues.yml.light-agentprovider config loading from embedded templates after bootstrap.portal-service/apps/portal-servicepre-runtimeportal-service.ymlload from embedded templates.portal-service/apps/config-serverstandaloneserver.ymlload from embedded templates.
Migration Plan
- Add embedded source support to
config-loaderandlight-runtime. - Add shared build-time template embedding for
light-gateway. - Wire
light-gatewayto pass embedded templates toLightRuntimeBuilder. - Keep Docker
config-defaultscopies unchanged and verify container parity. - Add native startup tests that run without a copied template directory.
- Roll the same pattern to
light-agentandlight-deployer. - Add the shared embedded-source helper to
portal-serviceand migrateportal-service,config-server, andlight-oauth. - Add print and dump commands for embedded templates.
- After several releases, deprecate Docker
config-defaultscopies and rely on embedded defaults plus explicit dump commands for inspectability.
Risks And Mitigations
| Risk | Mitigation |
|---|---|
| Embedded templates drift from source templates | Embed the committed config/ directory directly with include_dir, or generate a manifest from that directory at build time |
| Operators cannot inspect templates in native deployment | Keep source templates in repo and add print/dump commands for embedded templates |
| Docker behavior changes unexpectedly | Keep config-defaults above embedded defaults during migration |
| Config-server remote values stop overriding defaults | Preserve remote values as the highest non-env value layer |
| Apps duplicate merge logic | Move embedded-source merging into shared loader/runtime helpers |
| Secrets accidentally embedded | Embed only committed template files and keep secrets in values, env, or external files |
| Structured config becomes hard to reason about | Use source-level override for config files and reserve key-level merging for values.yml |
Resolved Decisions
- Native operators should get
--print-default-config <name>and--dump-default-configs <directory>commands. - Module registry should expose resolved config first, with source provenance as metadata when available.
- Docker images should keep
/app/config-defaultsduring migration, then deprecate it once embedded templates and dump commands are stable. - Rust deployments should standardize on embedded templates plus remote
values.yml. Config-server should not normally deliver full template files for Rust services.
Decision Summary
Embed app config/*.yml templates into the binary as the lowest-priority
default configuration source. The initial implementation uses a shared
build-time manifest generator, with include_dir remaining a possible future
implementation detail. Keep the existing source config directories for
documentation and build input. Use source-level override for structured config
files and key-level overlay for values.yml. Preserve current filesystem and
remote value layers so container deployments keep working, while native
deployments can run with only the binary and a small deployment-specific config
directory.
Handler Chain
Status: Phases 1, 2, 3, 4, 5, 6, 7, and 8 implemented; further transport phases proposed
Purpose
Light Fabric needs a light-pingora handler chain for the Rust
light-gateway product.
The first implementation should focus on light-pingora, not a generic
cross-framework abstraction. A Pingora-first design is simpler and matches the
gateway family of use cases: gateway, sidecar, proxy server, proxy client, load
balancer, and BFF.
The deployment model should use one light-gateway binary. Different runtime
behaviors should come from product-specific configuration managed in
light-portal and delivered by config-server. A BFF deployment, a sidecar
deployment, and a load-balancer deployment can therefore run the same binary
with different handler.yml, traffic/resource config, and handler-specific
config files.
The design should preserve the useful part of light-4j handler.yml: ordered
configuration of cross-cutting request and response concerns. It should not copy
the Java reflection model, mutable next handler pattern, or class-name-based
configuration.
Goals
- Add middleware handler-chain support to
frameworks/light-pingora. - Use one
apps/light-gatewaybinary for the Pingora gateway family. - Keep
handler.ymlas the chain and ordering configuration. - Let
light-portalmanage product-specific configuration and config-server deliver it at startup. - Support virtual hosts selected from the HTTP
Hostheader. - Serve static SPA content directly from Pingora.
- Proxy API, BFF, sidecar, and balancer routes to upstream services.
- Use stable handler IDs instead of Rust type names.
- Use explicit handler registration. Do not require
inventory. - Integrate loaded handler and traffic/resource config with
ModuleRegistry. - Keep the design compatible with runtime config reload.
Non-Goals
- Do not build a transport-neutral
light-handlercrate in the first phase. - Do not add an Axum/Tower adapter in the first phase.
- Do not create separate binaries for gateway, sidecar, proxy server, proxy client, load balancer, and BFF in the first phase.
- Do not dynamically load handler crates from
handler.yml. - Do not use Java-style reflection or string-to-type construction.
- Do not make Rust type names part of the public config contract.
- Do not support multi-certificate TLS SNI selection in the first phase.
- Do not implement streaming static-file delivery in the first phase unless it is needed for a concrete SPA asset size problem.
Current Shape
light-pingora already adapts a Pingora proxy into the shared runtime:
#![allow(unused)]
fn main() {
pub trait PingoraApp: Send + Sync + 'static {
type Proxy: ProxyHttp + Send + Sync + 'static;
fn proxy(&self, config: &RuntimeConfig) -> Result<Self::Proxy, RuntimeError>;
}
}
PingoraTransport calls app.proxy(config) and passes the result to
pingora::proxy::http_proxy_service(...).
Pingora’s ProxyHttp lifecycle already has the hooks needed for the gateway
family:
request_filter: validate, authenticate, rate limit, or directly write a local response such as a static fileupstream_peer: select the upstream for proxy routesupstream_request_filter: mutate the request sent to upstreamupstream_response_filter: mutate the upstream response before cachingresponse_filter: mutate the response sent to the browser
The current light-gateway already writes /health directly from
request_filter. Static SPA serving can use the same pattern.
Product Model
The Rust light-gateway binary should link all built-in Pingora gateway
capabilities:
- virtual host routing
- static SPA serving
- reverse proxy routing
- outbound proxy behavior
- upstream load balancing
- sidecar token/header behavior
- shared middleware handlers
The active behavior is selected by configuration, not by compiling a different binary. The six product personas are configuration profiles:
gatewaysidecarproxy-serverproxy-clientbalancerbff
These profiles can be represented in light-portal as product-specific config
sets. At runtime, light-gateway only sees the resolved files returned by
config-server. The binary should not need to know whether the files came from a
portal product template, an environment override, or a local fallback.
This keeps deployment simple:
- one binary
- one container image
- one
light-pingoraframework - different behavior by remote config
The tradeoff is that config validation must be strong. A product config should not silently start in a different mode if a static root, virtual host, upstream, or chain is wrong.
High-Level Flow
The Pingora gateway request flow should be:
request
-> match handler.yml paths by path and method
-> fall back to handler.yml defaultHandlers when no path matches
-> run request handlers
-> proxy fixed upstream, route by service_id/service_url, serve static file,
or return error
-> run response handlers
-> response
For static handlers such as virtual-host or path-resource, request_filter
writes the response and returns Ok(true) so Pingora does not proxy the
request.
For proxy or router handlers, request_filter stores the selected upstream
decision in the per-request context and returns Ok(false). upstream_peer
and upstream_request_filter then use that context to connect to the right
upstream and set headers.
Crate Layout
Keep the first implementation inside frameworks/light-pingora.
Suggested modules:
frameworks/light-pingora/src/
lib.rs
handler.rs
correlation.rs
cors.rs
metrics.rs
proxy.rs
resource.rs
router.rs
service.rs
token.rs
Responsibilities:
- parse and validate
handler.yml - parse
handler.yamlas a compatibility fallback - parse and validate
proxy.yml,router.yml,path-resource.yml, andvirtual-host.yml - build explicit handler registry
- resolve handler chains
- match handler paths and fallback handlers
- capture Java-style
{name}path-template variables - load active handler-specific config files
- serve static SPA content
- select fixed proxy upstreams from
proxy.yml - select dynamic sidecar/router upstreams from
router.yml - resolve sidecar
service_idvalues frompathPrefixService.yml - retrieve and cache OAuth client-credentials tokens from
client.yml - expose module-registry entries for active handler and traffic/resource config
This keeps the first implementation close to the Pingora lifecycle and avoids premature abstractions for Axum.
If Axum later needs the same handler semantics, extract the framework-neutral parts after the Pingora implementation has stabilized.
Configuration Split
Use handler.yml for the Java-compatible handler middleware contract:
handler declarations, reusable chains, path-to-chain mappings, and fallback
handlers.
Use Java-compatible product-specific config files for traffic and static resource behavior:
proxy.yml: fixed inbound reverse proxy targets for gateway, proxy server, balancer, and simple BFF API forwarding.router.yml: dynamic outbound routing byservice_idorservice_url, mainly for sidecar-style deployments.path-resource.ymlorpath-resource.yaml: a single static resource mount.virtual-host.ymlorvirtual-host.yaml: host-based static resource mounts for BFF/SPA deployments.
The product profile selected in light-portal decides which of these files are
included and which handlers are active in handler.yml. The Rust binary should
not require a separate gateway.yml to duplicate these existing contracts.
Handler-specific files such as correlation.yml, cors.yml, metrics.yml,
header.yml, security.yml, apikey.yml, basic-auth.yml,
unified-security.yml, and limit.yml stay separate. They are loaded only
when the corresponding handler is active in the resolved path/default
execution model. Phase 3 implements this active loading for correlation.yml,
cors.yml, and metrics.yml. Phase 4 extends the same active-loading and
reload model to header.yml, security.yml, apikey.yml,
basic-auth.yml, unified-security.yml, and limit.yml.
Remote Config Source
light-gateway starts with enough local bootstrap configuration to contact
config-server. The existing Light Fabric runtime then resolves local and remote
configuration before light-pingora builds the runtime handler/resource/proxy
model.
Startup flow:
- load local bootstrap files from the configured config directory
- contact config-server using the configured service identity, environment, and authorization
- download remote product configuration managed by
light-portal - merge remote config with local fallback config
- load
handler.yml, applicable traffic/resource config files, and active handler-specific config files - validate the complete route and handler model
- bind Pingora listeners
- register the runtime instance with the controller
The remote product config should include:
handler.ymlproxy.ymlfor fixed inbound proxy profilesrouter.ymlfor sidecar/router profilespath-resource.ymlorvirtual-host.ymlfor static/BFF profiles- active handler config files
- TLS, trust, or client files required by the runtime
- optional product-specific static file references or mount paths
handler.yml decides which linked handlers are active. A handler that is
registered in the binary but not referenced by any configured paths entry or
defaultHandlers chain should not be instantiated, should not load its config
file, and should never run.
Handler Config
Example handler.yml:
enabled: ${handler.enabled:true}
reportHandlerDuration: ${handler.reportHandlerDuration:false}
handlerMetricsLogLevel: ${handler.handlerMetricsLogLevel:DEBUG}
basePath: ${handler.basePath:/}
handlers: ${handler.handlers:[]}
chains: ${handler.chains:{}}
paths: ${handler.paths:[]}
defaultHandlers: ${handler.defaultHandlers:[]}
The config-server values managed by light-portal provide the concrete arrays
and maps:
handler.handlers:
- correlation
- headers
- metrics
- cors
- jwt
- rate-limit
handler.chains:
spa:
exec:
- correlation
- headers
- metrics
- cors
api:
exec:
- correlation
- headers
- metrics
- cors
- jwt
- rate-limit
public:
exec:
- correlation
- headers
- metrics
handler.paths:
- path: /api/
method: GET
exec:
- api
handler.defaultHandlers:
- public
This keeps the same top-level handler.yml contract as the Java framework:
enabled, reportHandlerDuration, handlerMetricsLogLevel, basePath,
handlers, chains, paths, and defaultHandlers.
The Rust implementation also accepts the Java extension fields
additionalHandlers, additionalChains, and additionalPaths. They are
merged into the effective handler model before validation.
Unlike Java, the Rust handlers list uses stable short handler IDs. It does
not use fully qualified class names, and it does not need @alias because the
IDs are already short and stable.
handler.yml is the preferred Rust file name. handler.yaml is accepted as a
compatibility fallback because some Java modules and templates use that suffix.
Fixed Proxy Config
proxy.yml should keep the Java inbound reverse-proxy contract. It is used
when the deployment has a known set of target upstream URIs.
enabled: ${proxy.enabled:true}
http2Enabled: ${proxy.http2Enabled:false}
hosts: ${proxy.hosts:http://localhost:8080}
connectionsPerThread: ${proxy.connectionsPerThread:20}
maxRequestTime: ${proxy.maxRequestTime:1000}
rewriteHostHeader: ${proxy.rewriteHostHeader:true}
reuseXForwarded: ${proxy.reuseXForwarded:false}
maxConnectionRetries: ${proxy.maxConnectionRetries:3}
maxQueueSize: ${proxy.maxQueueSize:0}
forwardJwtClaims: ${proxy.forwardJwtClaims:false}
metricsInjection: ${proxy.metricsInjection:false}
metricsName: ${proxy.metricsName:proxy-response}
The Rust implementation should parse proxy.hosts as one or more comma
separated http:// or https:// targets and select a target with round-robin
load balancing. It should preserve rewriteHostHeader, reuseXForwarded,
request timeout, retry, and queue settings where Pingora exposes equivalent
behavior.
Router Config
router.yml should keep the Java outbound router contract. This is primarily
for the sidecar pattern, where earlier handlers resolve service_id,
service_url, tokens, and discovery context before the router connects to the
downstream service.
http2Enabled: ${router.http2Enabled:true}
httpsEnabled: ${router.httpsEnabled:true}
maxRequestTime: ${router.maxRequestTime:1000}
pathPrefixMaxRequestTime: ${router.pathPrefixMaxRequestTime:{}}
connectionsPerThread: ${router.connectionsPerThread:10}
softMaxConnectionsPerThread: ${router.softMaxConnectionsPerThread:5}
maxQueueSize: ${router.maxQueueSize:0}
rewriteHostHeader: ${router.rewriteHostHeader:true}
reuseXForwarded: ${router.reuseXForwarded:false}
maxConnectionRetries: ${router.maxConnectionRetries:3}
preResolveFQDN2IP: ${router.preResolveFQDN2IP:false}
hostWhitelist: ${router.hostWhitelist:[]}
serviceIdQueryParameter: ${router.serviceIdQueryParameter:false}
urlRewriteRules: ${router.urlRewriteRules:[]}
methodRewriteRules: ${router.methodRewriteRules:[]}
queryParamRewriteRules: ${router.queryParamRewriteRules:{}}
headerRewriteRules: ${router.headerRewriteRules:{}}
metricsInjection: ${router.metricsInjection:false}
metricsName: ${router.metricsName:router-response}
The Java router chooses the target from service_url first, guarded by
hostWhitelist, or from service_id plus optional env_tag through service
discovery.
Phase 5 implements the Pingora router execution path and keeps the Java
configuration shape. The active router handler loads and registers
router.yml, selects direct service_url targets after hostWhitelist
validation, supports serviceIdQueryParameter, and removes router selection
headers before forwarding upstream. It also applies Java-style URL, method,
query-parameter, and header rewrite rules.
Phase 6 adds the sidecar path-prefix and token flow. Phase 7 adds
controller-backed service_id discovery while keeping the same request
contract. For local/static deployments, use direct-registry.yml as the
fallback service map.
Sidecar Path Prefix And Token Config
pathPrefixService.yml maps request path prefixes to downstream service IDs.
The handler writes service_id only when the request does not already provide
one.
enabled: ${pathPrefixService.enabled:true}
mapping: ${pathPrefixService.mapping:{}}
Rust intentionally selects the longest path-boundary prefix. This avoids map
iteration ambiguity when prefixes overlap and prevents /v1/address from
matching /v1/address2.
token.yml gates when the token handler should run:
enabled: ${token.enabled:false}
appliedPathPrefixes: ${token.appliedPathPrefixes:}
The token handler reads the Java-compatible client credentials section from
client.yml:
tls:
verifyHostname: ${client.verifyHostname:true}
oauth:
multipleAuthServers: ${client.multipleAuthServers:false}
token:
cache:
capacity: ${client.tokenCacheCapacity:200}
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
server_url: ${client.tokenServerUrl:}
serviceId: ${client.tokenServiceId:com.networknt.oauth2-token-1.0.0}
proxyHost: ${client.tokenProxyHost:}
proxyPort: ${client.tokenProxyPort:}
enableHttp2: ${client.tokenEnableHttp2:true}
client_credentials:
uri: ${client.tokenCcUri:/oauth2/token}
client_id: ${client.tokenCcClientId:}
client_secret: ${client.tokenCcClientSecret:}
scope: ${client.tokenCcScope:}
serviceIdAuthServers: ${client.tokenCcServiceIdAuthServers:}
pathPrefixServices: ${client.pathPrefixServices:}
request:
connectTimeout: ${client.connectTimeout:2000}
timeout: ${client.timeout:3000}
enableHttp2: ${client.enableHttp2:true}
In single-auth-server mode, the handler uses the configured token server and
client credentials for all matched paths. In multipleAuthServers mode, it
uses service_id or pathPrefixServices to select
client_credentials.serviceIdAuthServers[service_id].
The token request follows the Java request shape:
POSTtoserver_url + uriContent-Type: application/x-www-form-urlencodedAccept: application/json- HTTP Basic authentication with
client_id:client_secret - form fields
grant_type=client_credentialsand optional space-joinedscope
The injected header follows the Java gateway rule:
- if the inbound request has no
Authorization, injectAuthorization: Bearer <token> - if the inbound request already has
Authorization, injectX-Scope-Token: Bearer <token>
The Rust cache is local to the gateway process and is registered as
light-pingora/token-cache when a runtime cache registry is available. Cache
summaries expose key and expiry metadata but never expose bearer token values.
Tokens are refreshed synchronously inside the configured renew-before-expiry
window. Async background renewal can be added later if blocking refresh latency
becomes visible.
When server_url is not configured, phase 7 discovers the token service from
serviceId through the runtime portal-registry client. This requires
server.enableRegistry and a live controller registration. A disconnected
registry client returns a clear configuration/runtime error instead of silently
falling back to an unknown token endpoint.
Static Resource Config
For a single static site, keep path-resource.yml:
path: ${path-resource.path:/public}
base: ${path-resource.base:/opt/light-4j/public}
prefix: ${path-resource.prefix:true}
transferMinSize: ${path-resource.transferMinSize:1024}
directoryListingEnabled: ${path-resource.directoryListingEnabled:false}
For host-based BFF/static sites, keep virtual-host.yml:
hosts: ${virtual-host.hosts:[]}
Example config-server values:
virtual-host.hosts:
- domain: local.localhost
path: /
base: /lightapi/dist
transferMinSize: 10245760
directoryListingEnabled: false
- domain: signin.localhost
path: /
base: /signin/dist
transferMinSize: 10245760
directoryListingEnabled: false
Rust should preserve the Java domain, path, base, transferMinSize, and
directoryListingEnabled fields. It should also add the Rust improvement for
SPA fallback: when a static virtual host cannot find a requested browser route
and the path does not look like an asset, it should serve index.html from the
matched static root.
BFF Wiring Example
The Java BFF config in portal-config-loc/all-in-lt/light-gateway uses
handler.paths to send API routes through the default chain, which includes
path-prefix service resolution, token handling, and the router. It then uses:
handler.defaultHandlers:
- cors
- virtual
That means unmatched browser routes fall through to CORS plus virtual-host
static serving. Rust should keep this pattern: handler.yml decides whether a
request goes to proxy/router/static handling, based on paths and fallback
handlers.
Other product personas use different config file combinations. A BFF commonly
uses handler.yml, router.yml, path-prefix/token configs, and
virtual-host.yml. A simple proxy or balancer can use handler.yml and
proxy.yml. A sidecar uses handler.yml, router.yml, token/cache config,
registry/discovery config, and usually no static resource config.
Phase 3 Handler Config
Phase 3 implements the first three Java-compatible cross-cutting handlers.
correlation.yml:
enabled: ${correlation.enabled:true}
autogenCorrelationID: ${correlation.autogenCorrelationID:true}
correlationMdcField: ${correlation.correlationMdcField:cId}
traceabilityMdcField: ${correlation.traceabilityMdcField:tId}
The Rust handler reads X-Correlation-Id and X-Traceability-Id, generates a
Java-compatible URL-safe UUID value when correlation is missing, passes the
correlation ID to the upstream request, and echoes X-Traceability-Id on the
response. It stores the values in the Pingora request context instead of MDC.
cors.yml:
enabled: ${cors.enabled:true}
allowedOrigins: ${cors.allowedOrigins:}
allowedMethods: ${cors.allowedMethods:}
pathPrefixAllowed: ${cors.pathPrefixAllowed:}
The Rust handler accepts the same list/string forms as Java, supports
pathPrefixAllowed, short-circuits preflight OPTIONS, rejects disallowed
origins with 403, and adds the CORS response headers before static or proxied
responses are sent. Rust intentionally uses longest-prefix selection for
pathPrefixAllowed so overlapping prefixes are deterministic.
metrics.yml:
enabled: ${metrics.enabled:true}
enableJVMMonitor: ${metrics.enableJVMMonitor:false}
serverProtocol: ${metrics.serverProtocol:http}
serverHost: ${metrics.serverHost:localhost}
serverPath: ${metrics.serverPath:/apm/metricFeed}
serverPort: ${metrics.serverPort:8086}
serverName: ${metrics.serverName:metrics}
serverUser: ${metrics.serverUser:admin}
serverPass: ${metrics.serverPass:admin}
reportInMinutes: ${metrics.reportInMinutes:1}
productName: ${metrics.productName:http-sidecar}
sendScopeClientId: ${metrics.sendScopeClientId:false}
sendCallerId: ${metrics.sendCallerId:false}
sendIssuer: ${metrics.sendIssuer:false}
issuerRegex: ${metrics.issuerRegex:}
Phase 3 parses and registers this config with serverPass masked, records
request counts and status classes in memory, and logs request metrics with the
matched endpoint and correlation ID. enableJVMMonitor is parsed for config
compatibility but is not applicable to Rust. External Influx/APM reporters are
deferred until the metrics sink decision is made.
Phase 4 Handler Config
Phase 4 implements the security-oriented Java-compatible handlers that fit the Pingora request metadata model.
header.yml:
enabled: ${header.enabled:false}
request:
remove: ${header.request.remove:}
update: ${header.request.update:}
response:
remove: ${header.response.remove:}
update: ${header.response.update:}
pathPrefixHeader: ${header.pathPrefixHeader:}
The Rust handler applies request header remove/update rules before proxying and
response header remove/update rules before static or proxied responses are
sent. Rust intentionally uses longest-prefix selection for pathPrefixHeader
so overlapping prefixes are deterministic.
apikey.yml:
enabled: ${apikey.enabled:true}
hashEnabled: ${apikey.hashEnabled:false}
pathPrefixAuths: ${apikey.pathPrefixAuths:[]}
The Rust handler follows the Java rule that no matching path prefix means the
handler passes the request. A matching rule validates the configured header
against either a plain API key or the Java iterations:saltHex:hashHex
PBKDF2-HMAC-SHA1 hash format.
basic-auth.yml:
enabled: ${basic.enabled:false}
enableAD: ${basic.enableAD:true}
allowAnonymous: ${basic.allowAnonymous:false}
allowBearerToken: ${basic.allowBearerToken:false}
users: ${basic.users:[]}
The Rust handler supports configured local users, anonymous path users, and the Java-compatible bearer pass-through mode. LDAP/AD authentication is parsed for configuration compatibility but is not implemented in phase 4.
security.yml:
enableVerifyJwt: ${security.enableVerifyJwt:true}
ignoreJwtExpiry: ${security.ignoreJwtExpiry:false}
enableH2c: ${security.enableH2c:false}
enableMockJwt: ${security.enableMockJwt:false}
jwt:
clockSkewInSeconds: ${security.jwt.clockSkewInSeconds:60}
skipPathPrefixes: ${security.skipPathPrefixes:[]}
passThroughClaims: ${security.passThroughClaims:{}}
The Rust handler verifies Bearer JWTs with configured JWKs, honors
kid when present, supports RSA and EC algorithms handled by the Rust JWT
library, applies clock skew and optional expiry bypass, caches decoded claims,
and forwards configured pass-through claims as request headers. Dynamic JWK key
service bootstrap and SWT/SJWT verification are deferred until the runtime has
the discovery and key-service client surface needed by those flows.
unified-security.yml:
enabled: ${unified-security.enabled:true}
anonymousPrefixes: ${unified-security.anonymousPrefixes:[]}
pathPrefixAuths: ${unified-security.pathPrefixAuths:[]}
The Rust handler supports Java-style path-prefix selection across Basic, JWT, and API-key authentication. Anonymous prefixes bypass authentication. SWT/SJWT rules return a clear not-implemented response until the discovery-backed key flow is added.
limit.yml:
enabled: ${limit.enabled:false}
concurrentRequest: ${limit.concurrentRequest:0}
queueSize: ${limit.queueSize:0}
errorCode: ${limit.errorCode:429}
rateLimit: ${limit.rateLimit:}
headersAlwaysSet: ${limit.headersAlwaysSet:false}
key: ${limit.key:server}
server: ${limit.server:{}}
address: ${limit.address:{}}
client: ${limit.client:{}}
user: ${limit.user:{}}
The Rust handler implements in-memory request rate limiting by server, client
address, JWT client ID, or JWT user ID. It emits X-RateLimit-Limit,
X-RateLimit-Remaining, X-RateLimit-Reset, and Retry-After when a request
is rejected, and it can always emit the rate-limit headers when
headersAlwaysSet is enabled. Cluster-wide distributed counters are deferred
until there is a concrete gateway clustering requirement.
Handler Registry
Use explicit registration.
#![allow(unused)]
fn main() {
let handlers = PingoraHandlerRegistry::new()
.register(correlation::descriptor())
.register(headers::descriptor())
.register(metrics::descriptor())
.register(cors::descriptor())
.register(jwt::descriptor())
.register(rate_limit::descriptor());
}
No inventory is needed for the first version. Explicit registration is
deterministic, testable, and makes the compiled-in handler set clear from the
service binary.
The light-gateway binary can register every built-in handler it supports.
Registration only makes a handler available. Activation is controlled by
handler.yml.
Build the active handler set lazily:
- parse
handler.yml - resolve
pathsanddefaultHandlers - expand any referenced chains
- compute the set of referenced handler IDs
- instantiate only referenced handlers
- load config only for referenced handlers
This allows one binary to support gateway, sidecar, proxy, balancer, and BFF profiles without requiring unused handler config files.
The registry maps stable config IDs to factories:
#![allow(unused)]
fn main() {
pub struct PingoraHandlerDescriptor {
pub id: &'static str,
pub kind: PingoraHandlerKind,
pub factory: PingoraHandlerFactory,
}
}
Suggested first handler IDs:
correlationheadersmetricscorsjwtapi-keybasic-authrate-limitrequest-size-limit
Trace headers should be handled by correlation; there should not be a
separate traceability handler.
Handler API
Use Pingora phases directly. Avoid a generic exchange abstraction until another framework needs it.
The current implementation keeps PingoraHandler as a descriptor/factory
surface and executes the built-in phase 3 handlers from light-gateway’s
Pingora lifecycle. This keeps the first implementation straightforward:
request_filterresolves the configured chain and runs request-stage handlers in order.- A request-stage handler can continue, short-circuit with a local response, or select a terminal action such as proxy/static/health.
upstream_request_filterapplies upstream request mutations such as generated correlation IDs.response_filterapplies response-stage headers and records proxied response metrics.- Static responses call the same response decoration and metrics code before writing the local response.
Once security/rate-limit handlers are added, this can be lifted into a richer trait with request/upstream/response hooks if the duplication becomes real. It is intentionally not generalized before the Pingora behavior stabilizes.
Response handlers should run before both static and proxied responses are sent.
For proxied responses, this maps to Pingora response_filter. For static
responses, the static-file renderer calls the same response handler chain
before writing the local response.
Request Context
The per-request context should carry route decisions across Pingora phases.
#![allow(unused)]
fn main() {
pub struct GatewayRequestContext {
pub upstream: Option<ProxyTarget>,
pub endpoint: String,
pub method: String,
pub path_params: BTreeMap<String, String>,
pub correlation: CorrelationState,
pub cors: Option<CorsResponseHeaders>,
pub metrics_enabled: bool,
}
}
The context is created by ProxyHttp::new_ctx() and populated in
request_filter.
upstream_peer should only select an upstream after a proxy or router handler
has selected one. If no upstream is selected for a proxied request, the
implementation should return a clear configuration error rather than silently
falling back.
Virtual Hosts
Virtual-host static serving should use the HTTP Host header.
Host normalization rules:
- lowercase the host
- strip the port when present
- reject empty or invalid hosts unless a default virtual host is configured
- exact host match first
- wildcard match such as
*.example.comafter exact hosts, with the longest matching suffix winning
HTTP host routing is enough for the first implementation.
TLS certificate selection by SNI is separate. The current light-pingora
transport uses one Rustls TLS setting for the listener, so the first production
options are:
- terminate TLS at ingress or a load balancer
- use a wildcard certificate
- use one certificate with all required SANs
Phase 8 evaluated dynamic multi-cert SNI selection. The current
light-pingora build uses Pingora’s Rustls listener, and Pingora 0.8 Rustls
TLS settings do not support certificate callbacks. For now the production
options remain terminating TLS before light-gateway, using a wildcard
certificate, or using one certificate with all required SANs. Native multi-cert
SNI can be added only after moving to a Pingora TLS backend/version that
supports server certificate callbacks or certificate resolution through Rustls.
Static SPA Rendering
Static SPA rendering should be part of the Pingora resource engine, not a
generic middleware handler. It is enabled by path-resource.yml or
virtual-host.yml, typically for BFF profiles.
Rules for the first implementation:
- support
GETandHEAD - return
405for unsupported methods on static routes - canonicalize requested paths under the configured static root
- reject path traversal
- do not serve files outside the static root
- deny dotfiles by default
- do not list directories
- serve
index.htmlfor the root path - support SPA fallback to
index.htmlfor non-asset routes - infer
Content-Typefrom file extension - set
Cache-Control: no-cacheforindex.html - set long immutable cache headers for hashed assets
- allow static route prefixes to be bypassed by API routes such as
/api/,/oauth/,/mcp/, or/ws/
Recommended cache behavior:
index.html Cache-Control: no-cache
*.js, *.css with hash Cache-Control: public, max-age=31536000, immutable
images/fonts with hash Cache-Control: public, max-age=31536000, immutable
other assets Cache-Control: public, max-age=3600
Phase 8 keeps small static files on the simple read-then-write path and streams
files whose size is greater than or equal to the configured
transferMinSize. Static responses include ETag and Last-Modified, honor
If-None-Match and If-Modified-Since, and return 304 without a response
body when the browser cache is current.
Proxy And Router Behavior
proxy.yml selects from configured upstream URIs. This is the simpler inbound
reverse-proxy case and should be implemented before dynamic sidecar routing.
Fixed proxy target behavior:
- parse comma-separated
proxy.hosts - support
http://andhttps:// - duplicate a single host internally if retry/load-balancer behavior needs at least two entries
- select upstream with round-robin
- apply timeout, retry, queue, and host-forwarding settings where Pingora supports them
router.yml selects from request metadata. Phase 5 implements direct
service_url targets, host whitelist enforcement, and rewrite behavior. Phase
7 adds controller-backed service_id lookup through the runtime
portal-registry client with direct-registry.yml as the static fallback.
Router target behavior:
- prefer
service_urlwhen present and allowed byrouter.hostWhitelist - otherwise use
service_idplus optionalenv_tag - optionally allow
service_idfrom the query string whenserviceIdQueryParameteris true - resolve
service_idfrom controller discovery when the portal-registry client is connected - fall back to
direct-registry.directUrlsfor local/static deployments or controller lookup failures - support URL, method, query-parameter, and header rewrite rules
- remove
service_urlandservice_idheaders before forwarding
upstream_peer creates the HttpPeer from the selected upstream:
- address
- TLS enabled
- SNI
- optional host header
upstream_request_filter should set or override upstream headers such as:
HostX-Forwarded-ForX-Forwarded-ProtoX-Forwarded-Host
It must also remove client-supplied internal trust markers such as
X-Light-Gateway. The generic gateway does not synthesize a Portal-specific
identity marker; deployments that require authenticated upstream identity must
use a mutually authenticated transport or equivalent infrastructure control.
Handler-specific upstream mutations should also run from this phase.
Chain Resolution
Startup should validate handler and selected traffic/resource configuration before binding listeners.
Validation rules:
- every handler ID in
handler.ymlmust exist in the explicit registry - every chain item must resolve to a registered handler or another chain
- recursive chain references are invalid
- every
handler.pathsentry must reference existing chains or handlers - every
handler.defaultHandlersentry must reference existing chains or handlers proxy.ymlhosts must be validhttp://orhttps://URIs when the proxy handler is activerouter.ymlrewrite rules must be parseable when the router handler is active- every static virtual host must have a static root
- static roots must be absolute or resolved relative to a configured base
- duplicate exact virtual hosts are invalid
- duplicate handler IDs in the registry are invalid
The resolved model should be immutable and cheap to read:
#![allow(unused)]
fn main() {
pub struct GatewayRuntimeModel {
pub virtual_hosts: BTreeMap<String, Arc<VirtualHost>>,
pub default_host: Option<Arc<VirtualHost>>,
pub chains: BTreeMap<String, Arc<ResolvedHandlerChain>>,
pub proxy_targets: Vec<Arc<ProxyTarget>>,
}
}
Config reload should continue to swap loaded models atomically. In-flight requests should keep using the handler/resource/proxy/router model they already selected.
Runtime Integration
light-runtime remains responsible for bootstrap, config loading, lifecycle,
controller registration, and module registry. light-pingora should load its
Pingora-specific handler, traffic, and resource config through the existing
runtime config loader.
Module IDs:
light-pingora/handlerlight-pingora/proxylight-pingora/routerlight-pingora/path-prefix-servicelight-pingora/tokenlight-client/clientlight-pingora/path-resourcelight-pingora/virtual-hostlight-pingora/correlationlight-pingora/corslight-pingora/metricslight-pingora/headerlight-pingora/securitylight-pingora/apikeylight-pingora/basic-authlight-pingora/unified-securitylight-pingora/limit
The module registry should expose:
- handler config snapshot, masked
- proxy, router, path-resource, and virtual-host config snapshots, masked
- active handler IDs
- active chains
- active virtual hosts
- active proxy/router/static capabilities
- reloadable status
The implemented phases use the existing ReloadableModule pattern for active
handler, proxy, router, resource, virtual-host, path-prefix service, token, and
handler-specific config files. Phase 7 exposes a capabilities summary from
get_service_info, including active modules, traffic capabilities, active
handlers, chain names, path mappings, default handlers, virtual hosts, and
path-resource config.
Suitable First Handlers
Start with handlers that map cleanly to Pingora request and response metadata:
- correlation ID and trace headers
- response headers
- metrics
- CORS
- JWT verification
- API key verification
- basic auth
- request size limit from headers
- simple rate limiting by principal, IP, host, or route
Defer handlers that require deeper body handling:
- request decompression
- response compression policy beyond Pingora modules
- request body sanitizer
- generic body parser
- WebSocket message handlers
Error Model
Handlers and proxy/resource selection should return structured errors that render consistently.
#![allow(unused)]
fn main() {
pub struct HandlerError {
pub status: u16,
pub code: Cow<'static, str>,
pub message: Cow<'static, str>,
pub metadata: serde_json::Value,
}
}
Security handlers should avoid returning sensitive validation details to the browser. Detailed diagnostics should go to logs with correlation IDs.
Common gateway errors:
- unknown host:
404 - no matching handler path or static resource:
404 - unsupported method for static route:
405 - static file outside root:
403 - missing upstream: startup validation error
- auth failure:
401or403 - rate limit:
429
Testing Strategy
Unit tests in light-pingora:
- build active handler set from referenced
pathsanddefaultHandlers - ignore registered but unreferenced handlers
- do not require config files for unreferenced handlers
- parse valid
handler.yml - reject unknown handler IDs
- reject recursive chains
- resolve path/default handler chains in order
- parse
handler.yamlfallback - merge
additionalHandlers,additionalChains, andadditionalPaths - capture path-template variables
- parse CORS list/string and path-prefix config
- classify metrics status codes
- normalize host names and strip ports
- reject duplicate virtual hosts
- match exact virtual hosts
- parse and validate
proxy.ymlhosts - parse and validate
router.ymlrewrite-rule config - select router targets from direct
service_url - reject direct router targets that do not match
hostWhitelist - select router targets from controller discovery and
direct-registry.yml - apply router URL, method, query-parameter, and header rewrites
- parse
pathPrefixService.ymland avoid partial-segment path matches - parse
token.ymland the client credentials subset ofclient.yml - support single and multiple auth-server token configuration
- discover token service endpoints from
client.ymltokenserviceId - mask token cache summaries and never expose bearer token values
- expose gateway capabilities in
get_service_info - prevent static path traversal
- deny dotfiles by default
- serve
index.htmlfor/ - serve SPA fallback for non-asset paths
- avoid SPA fallback for
/api/proxy routes - select cache headers for
index.htmland hashed assets - stop handler execution on early response
- run response handlers before static response write
Integration tests:
- same binary starts with BFF profile config
- same binary starts with proxy or balancer profile config
- BFF profile can route API paths through configured handlers and serve SPA
fallback through
defaultHandlers - static SPA route returns
index.html - static asset route returns correct content type and cache header
- virtual host A and virtual host B serve different roots
- API route is proxied to the configured
proxy.ymlupstream - auth handler blocks protected API routes
- public static route does not require auth unless configured
Rollout Plan
Phase 1: Product config and active handler model (implemented)
- keep a single
apps/light-gatewaybinary - register all built-in handler descriptors explicitly
- resolve active handler IDs from
handler.yml - instantiate only active handlers
- load config only for active handlers
- document product profiles managed by
light-portal
Phase 2: BFF and fixed proxy engine (implemented)
-
load and register
proxy.yml,path-resource.yml, andvirtual-host.yml -
match
handler.ymlpaths and fallback handlers in Java-compatible order -
select fixed proxy upstreams from
proxy.yml -
match virtual hosts by
Host -
serve single-site and virtual-host static content
-
implement safe static path resolution
-
serve static files from
request_filter -
add Rust SPA fallback improvement
-
add content type and cache headers
-
add traversal, dotfile, fallback, proxy-host, and virtual-host tests
Phase 3: Handler chain execution (implemented)
- run request and response handlers around static and proxied responses
- implement correlation, CORS, and basic metrics
- parse
correlation.yml,cors.yml, andmetrics.yml - pass generated correlation IDs upstream
- apply response headers to both static and proxied responses
- log handler duration when
reportHandlerDurationis enabled - defer generic response headers to a handler-specific follow-up
Phase 4: Security and request/response policy handlers (implemented)
- implement JWT, API key, basic auth, and rate-limit handlers
- implement the generic header handler for request and response mutation
- implement unified-security path-prefix selection for Basic, JWT, and API key
- parse Java-compatible
security.yml,apikey.yml,basic-auth.yml,unified-security.yml,header.yml, andlimit.yml - add JWT pass-through claim request header mutation
- add path-level chain selection for public SPA and protected API routes
Phase 5: Sidecar router (implemented)
- load and register
router.yml - implement dynamic target selection by
service_urlorservice_id - enforce
hostWhitelist - support router URL, method, query-parameter, and header rewrites
- apply router request mutation in
upstream_request_filter - remove router selection headers before forwarding
- include router config in the active reload model
- add sidecar-focused tests
Phase 6: Sidecar path-prefix and token flow (implemented)
- load and register
pathPrefixService.yml - resolve
service_idby longest path-boundary prefix - load and register
token.yml - load and register the token-related view of
client.yml - support single-auth-server and
multipleAuthServersclient credentials - cache tokens locally and expose masked cache summaries through the runtime cache registry
- inject
AuthorizationorX-Scope-Tokenaccording to inbound request state - extend reload coverage to
pathPrefixService.yml,token.yml, and token-relatedclient.yml - add sidecar token/path-prefix tests
Phase 7: Discovery and control plane (implemented)
- expose the runtime portal-registry client to framework transports
- add
discovery/lookupsupport to the portal-registry client - resolve router
service_idtargets through controller discovery - fall back to
direct-registry.ymlfor local/static profiles - discover token service endpoints from
client.ymltokenserviceId - expose active capabilities, hosts, paths, handlers, and chains through
get_service_info - atomically replace resolved handler/resource/proxy models on reload
Phase 8: Advanced transport features (implemented)
- add streaming static-file delivery for files at or above
transferMinSize - add conditional static requests with
ETagandLast-Modified - add wildcard virtual hosts with exact-host precedence
- evaluate multi-cert TLS SNI support and document the Rustls limitation
Phase 2 Decisions
- Static roots can be absolute, matching the Java deployment model, or relative to the runtime config directory for local Rust development.
- SPA fallback applies only to browser routes. Paths that look like assets,
such as
/app.jsor/favicon.ico, return 404 when the file is missing. - Handler path matching supports exact paths and Java/OpenAPI-style
{name}path-template segments.
Open Questions
- Should static content support
ETagin the first implementation if portal deployments depend on browser cache validation?
MCP Router
Status
Phases 1, 2, 3, and 4 are implemented in light-pingora and
light-gateway. The configurable tokenization client remains deferred until
light-tokenization is migrated to portal-service/apps/portal-service and
the protocol is selected. Stateful backend MCP session mapping is implemented
for the single-process gateway session store and documented below.
This page describes the implemented legacy stateful profile. The design for
serving that profile concurrently with the sessionless and stateless
2026-07-28 profile is documented in
MCP 2026-07-28 Dual-Profile Gateway Design.
Purpose
The Java mcp-router module exposes a configured Model Context Protocol
endpoint, /mcp by default, and turns configured gateway targets into MCP
tools. AI agents can call initialize, tools/list, and tools/call; the
router then forwards the tool call to an HTTP service or another MCP server.
In light-fabric this should be a light-pingora handler that is activated by
light-gateway through handler.yml. The same gateway binary can contain the
MCP router implementation, but each product decides whether it runs by including
the mcp handler and the mcp-router.yml configuration from the config server.
This feature is separate from the existing runtime MCP control plane in
light-runtime. Runtime MCP is an internal management surface exposed through
the portal registry connection. The MCP router is an HTTP-facing gateway
feature and is subject to the normal inbound handler chain.
The implemented legacy transport baseline is MCP Streamable HTTP as defined by
the 2025-06-18 transport specification:
https://modelcontextprotocol.io/specification/2025-06-18/basic/transports.
It must not be interpreted as the current MCP protocol revision; later legacy
and stateless revision support is governed by the dual-profile design linked
above.
Goals
- Keep the Java configuration model recognizable:
enabled,path, andtools. - Allow
mcp-router.toolsto be injected by the config server the same wayhandler.handlers,handler.chains,handler.paths, andhandler.defaultHandlersare injected. - Activate the router with the existing
mcphandler id inhandler.yml. - Expose one MCP endpoint with Streamable HTTP semantics, so
/mcpis the only public MCP path for both POST messages and optional GET streams. - Support MCP JSON-RPC methods needed by the Java module:
initialize,notifications/initialized,tools/list, andtools/call. - Route tools to direct
targetHostendpoints, discoveredserviceIdtargets, and backend MCP servers. - Reuse existing cross-cutting handlers such as correlation, security, CORS, rate limit, header, metrics, and proxy routing where the chain order allows.
- Register the router configuration with the module registry so it can be inspected and reloaded consistently with other light-fabric modules.
Non-Goals
- Do not use Rust dynamic plugins or
inventoryfor runtime tool registration. The active tools are product configuration, not compile-time discovery. - Do not merge the public MCP router and the internal runtime MCP control plane into one handler.
- Do not implement a full MCP server framework in the first pass. The gateway only needs the methods used by agents to discover and call configured tools.
- Do not copy Java’s legacy HTTP+SSE endpoint split as the target transport. Streamable HTTP is the Rust target; legacy SSE can be considered only as a compatibility mode if an older client requires it.
- Do not hardcode tokenization or masking service URLs. Java currently has a hardcoded tokenization endpoint in this path; the Rust port should make that configurable when masking/tokenization is added.
Java Behavior To Map
The Java module has three main pieces:
McpConfigloadsmcp-router.ymlwithenabled,path, andtools.McpHandlerowns the HTTP MCP endpoint and JSON-RPC protocol handling.McpToolRegistrystores configured tool implementations by name.
Java configuration:
enabled: ${mcp-router.enabled:true}
path: ${mcp-router.path:/mcp}
maxSessions: ${mcp-router.maxSessions:10000}
maxSessionsPerClient: ${mcp-router.maxSessionsPerClient:100}
tools: ${mcp-router.tools:}
Each tool supports these fields:
- name: weather
description: Get weather information
protocol: http
serviceId: com.networknt.weather-1.0.0
envTag: dev
targetHost: http://localhost:7081
path: /weather
method: GET
endpoint: /weather@get
apiType: http
inputSchema:
type: object
properties:
city:
type: string
toolMetadata: {}
The Java handler currently supports:
GET /mcpas an SSE compatibility endpoint. It creates a session id and emits anendpointevent pointing to/mcp?sessionId=....POST /mcpfor JSON-RPC messages.initialize, returning protocol version, tool capabilities, and server info.notifications/initialized, returning no response.tools/list, optionally filtered byparams.queryorparams.intent.tools/call, forwarding arguments to the configured tool.
The Java tool execution supports two target types:
- HTTP tools call a configured HTTP endpoint.
GETmaps arguments to query parameters. Other methods send the arguments as a JSON body. - MCP proxy tools call a backend MCP server by sending a JSON-RPC
tools/callrequest to the configured backend path.
Java also includes rule-based access checks, response filtering, masking, and tokenization around tool calls. The Rust version now implements access checks, response filtering, and schema-driven request masking without hardcoded service endpoints. Tokenization is intentionally deferred.
The Rust implementation should map this behavior to MCP Streamable HTTP rather
than keeping Java’s legacy HTTP+SSE transport as the default. Streamable HTTP
uses one MCP endpoint path. Clients send JSON-RPC messages with POST /mcp;
the server can return either a single application/json response or
text/event-stream from that same POST when streaming is needed. Clients may
also issue GET /mcp to open an optional server-to-client SSE stream on the
same endpoint.
Resolved Decisions
- Use Streamable HTTP so only one public MCP endpoint, normally
/mcp, is exposed. - Defer the tokenization client design until
light-tokenizationis migrated intoportal-service/apps/portal-serviceand its protocol is selected. - Reuse the light-4j
access-control.ymlcompatibility contract for MCP, REST, and JSON-RPC authorization. - Do not add configured per-tool outbound headers. Backend tool calls should pass through the headers received from the agent, subject only to headers that the HTTP client must regenerate for a new outbound request and MCP session headers that the gateway must map or regenerate.
Rust Architecture
Add the MCP router to light-pingora because it is a request/response gateway
handler. light-gateway should wire it into the existing handler descriptor
table and runtime state.
Proposed modules:
frameworks/light-pingora/src/access_control.rs
frameworks/light-pingora/src/mcp.rs
Primary types:
#![allow(unused)]
fn main() {
pub struct McpRouterConfig {
pub enabled: bool,
pub path: String,
pub tools: Vec<McpToolConfig>,
}
pub struct McpToolConfig {
pub name: String,
pub description: String,
pub protocol: Option<String>,
pub service_id: Option<String>,
pub env_tag: Option<String>,
pub target_host: Option<String>,
pub path: String,
pub method: HttpMethod,
pub endpoint: Option<String>,
pub api_type: McpToolType,
pub input_schema: serde_json::Value,
pub tool_metadata: serde_json::Value,
}
pub struct McpRouterRuntime {
pub config: ArcSwap<McpRouterConfig>,
pub client: reqwest::Client,
pub registry_client: Option<Arc<PortalRegistryClient>>,
}
}
The exact field names should follow the existing light-fabric serde naming style while accepting the Java config names through aliases:
serviceIdenvTagtargetHostapiTypeinputSchematoolMetadata
mcp-router.yml should be the primary Rust file name, but the loader should
also accept mcp-router.yaml for Java compatibility.
Tool Registration
The router does not need global static registration. Build an immutable tool map
when mcp-router.yml is loaded:
McpRouterConfig -> BTreeMap<String, McpToolConfig> -> Arc<McpRouterState>
On reload, build a new state and atomically swap the Arc. In-flight requests
continue with the old state.
This is simpler than Java’s static McpToolRegistry and avoids Rust plugin
complexity. It also matches the light-fabric product model: all handlers can be
linked into one binary, while the config server decides which handlers and tools
are active for a product.
Request Flow
The mcp handler should participate in the normal handler chain:
request
-> correlation
-> metrics
-> cors
-> security or unified security
-> limit
-> mcp
-> proxy or route handler, only if mcp did not consume the request
response
-> header
-> metrics
-> access log
When the request path matches mcp-router.path:
POSTparses a JSON-RPC message. Requests return eitherapplication/jsonfor a single response ortext/event-streamfor a streamed response on the same endpoint. Notifications and JSON-RPC responses sent by the client return202 Acceptedwith no body when accepted.GETwithAccept: text/event-streammay open a server-to-client SSE stream on the same endpoint. If the gateway has no server-initiated messages to stream, it should return405 Method Not Allowed.DELETEshould terminate the gateway session and any mapped backend MCP sessions. Until session termination is implemented, it can return405 Method Not Allowed.- Other methods return
405 Method Not Allowed.
When the path does not match, the handler continues to the next handler in the configured chain.
The handler must be safe to include in shared chains. If mcp-router.enabled is
false, or the mcp handler is not in handler.yml, no MCP route is exposed.
JSON-RPC Handling
Supported methods:
initialize
notifications/initialized
tools/list
tools/call
initialize response:
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"protocolVersion": "2024-11-05",
"capabilities": {
"tools": {
"listChanged": true
}
},
"serverInfo": {
"name": "light-gateway-mcp",
"version": "1.0.0"
}
}
}
tools/list returns configured tools with name, description, and
inputSchema. It should preserve Java’s simple filtering:
params.querymatches tool name or description.params.intentmatches tool name or description.
tools/call validates params.name, finds the tool, validates or forwards
params.arguments, and returns either:
{
"content": [
{
"type": "text",
"text": "..."
}
]
}
or the structured result returned by the backend MCP server.
JSON-RPC errors should use the same codes as Java where practical:
-32700 parse error
-32601 method or tool not found
-32602 invalid params
-32000 tool execution failed
-32001 access denied
Rust improvement: malformed transport payloads should return a clear HTTP 400
with a JSON-RPC error body instead of a generic HTTP 500.
For Streamable HTTP:
- Clients must send each JSON-RPC message as a separate
POSTto the MCP endpoint. - Clients should send
Accept: application/json, text/event-stream. - The router should negotiate and honor
MCP-Protocol-Version. - The router terminates the client-facing MCP session.
initializeresponses should include a gateway-ownedMcp-Session-Id, and later client requests should be validated against that gateway session.
MCP Session Management
The MCP router should use a facade model. To the agent, light-gateway is the
MCP server. To upstream MCP targets, light-gateway is an MCP client. This
keeps gateway security, access-control policy, masking, response filtering, and
tool aggregation in one place while still respecting upstream MCP session
state.
There are two distinct session scopes:
- Frontend session: the session between the MCP client and
light-gateway. - Backend session: one upstream MCP server session owned by the gateway for a specific frontend session and backend target.
The frontend session is created during client initialize:
- The client sends
initializetomcp-router.path. - The gateway returns the MCP capabilities it exposes and a gateway-generated
Mcp-Session-Id. - The gateway stores session state keyed by that id. The state should include the negotiated protocol version, client info, security principal or relevant auth context, and any backend MCP sessions created for this client session.
- Later client requests must include the gateway session id. Unknown or expired session ids should fail before tool execution.
- A client
DELETErequest, explicit expiry, or gateway shutdown should close all backend sessions associated with the frontend session.
Session Duration And Limits
Frontend MCP sessions last until one of these events happens:
- The session is idle for 30 minutes.
- The client sends
DELETEto the MCP endpoint with the gatewayMcp-Session-Id. - The gateway process exits.
The 30-minute value is an idle timeout, not a fixed session lifetime. Each
valid session-bound request refreshes last_accessed, so an active client can
keep using the same frontend MCP session for longer than 30 minutes. Once the
session has been idle for 30 minutes, the next validation or lazy purge removes
it and later requests with that session id fail as an unknown session.
The idle timeout is currently compiled into light-pingora as
MCP_SESSION_IDLE_TIMEOUT and is not configurable from mcp-router.yml. The
lazy purge throttle is also compiled in as MCP_SESSION_PURGE_INTERVAL with a
60-second interval. Expired sessions may therefore remain in memory briefly
until another MCP request triggers validation or purge, but they are rejected
when used after the idle timeout.
The configurable session settings are capacity limits:
enabled: ${mcp-router.enabled:true}
path: ${mcp-router.path:/mcp}
maxSessions: ${mcp-router.maxSessions:10000}
maxSessionsPerClient: ${mcp-router.maxSessionsPerClient:100}
tools: ${mcp-router.tools:[]}
maxSessionslimits the total number of frontend MCP sessions held by one gateway process. The default is10000.maxSessionsPerClientlimits sessions for one client key. The default is100.
Both capacity values must be greater than zero. When a new initialize request
would exceed either limit, the router first forces an expired-session purge. If
the limit is still reached, the request fails without issuing another
Mcp-Session-Id: total store exhaustion returns 503, and per-client
exhaustion returns 429.
Expired sessions are purged lazily during later MCP requests, and any mapped backend MCP sessions are closed during that purge. If a frontend session is deleted or expires, the gateway also terminates every backend MCP session mapped to that frontend session.
The per-client key is derived from the authenticated principal when available,
preferring client_id, then user_id, email, and host. If no security
principal is available, the key falls back to MCP clientInfo.name and
clientInfo.version from the initialize request.
For a single gateway process, the session store can start in memory. In a
multi-pod deployment, the store should be external, such as Redis, or ingress
must provide sticky routing for all requests that carry the same
Mcp-Session-Id.
Backend handling depends on the tool type.
For apiType: http, the backend is a normal stateless API:
- No backend MCP session is created.
- The gateway translates
tools/callarguments into a normal HTTP request. GETtools serialize arguments into the query string; body-capable methods send JSON.- The gateway wraps the HTTP response into an MCP
tools/callresult. - User-specific auth, tenant, correlation, and trace headers come from the frontend session or inbound request and are applied to the outbound HTTP call as normal gateway headers.
For apiType: mcp, the backend is a stateful MCP server:
- The gateway lazily initializes the backend session the first time a frontend
session calls a tool for that backend target. If future dynamic tool
discovery depends on the backend, this initialization can happen before
tools/listinstead. - The gateway sends
initializeto the backend MCP endpoint as an MCP client. It should use the client-requested protocol version when supported and pass only the capabilities it needs upstream. - If the backend returns
Mcp-Session-Id, the gateway stores it in a mapping keyed by the gateway session id and backend target identity. - The gateway sends
notifications/initializedto the backend when the backend session is established. - For later backend calls, the gateway sends the backend session id to that backend. It must not forward the frontend gateway session id as if it were a backend session id.
- The gateway still performs access checks before calling the backend and response filtering after the backend response.
- When the frontend session ends, the gateway should terminate each mapped backend MCP session to avoid leaking backend resources.
The backend target identity used in the session map should be stable across
requests. It should include the resolved route information that distinguishes
one backend MCP endpoint from another, such as targetHost or serviceId,
envTag, protocol, and tool path.
When the router aggregates tools from both MCP servers and normal APIs, the
client still sees one gateway MCP session and one tools/list response. The
gateway registry decides how each tools/call is executed:
| Feature | MCP server backend | Normal API backend |
|---|---|---|
| Config type | apiType: mcp | apiType: http or omitted |
| Backend session | Yes, mapped from gateway session to backend target | No |
| Initialization | Gateway initializes backend as an MCP client | No upstream initialization |
| Message handling | JSON-RPC tools/call through backend MCP session | Translate JSON-RPC arguments to HTTP |
| Backend session header | Send backend Mcp-Session-Id only to that backend | Do not send MCP session state |
| Tear-down | Close backend session on client session end | Nothing backend-specific |
The configured tools/list remains the gateway’s public contract. A future
dynamic-discovery mode may call backend MCP tools/list and merge those tools
with configured HTTP tools, but that must still preserve the gateway’s policy
surface and avoid exposing backend tools that are not authorized for the
product.
HTTP Tool Execution
For apiType: http or missing apiType:
- Resolve the target base URL.
- Build the target URL from base URL plus tool
path. - For
GET, serialize arguments withurl::form_urlencoded. - For
POST,PUT, andPATCH, send arguments as JSON. - Pass through the inbound agent headers to the backend tool call so caller identity, authorization, correlation, tenant, locale, and tracing context are preserved.
- Let the HTTP client regenerate transport-specific headers for the new
outbound request, such as
Host,Content-Length,Transfer-Encoding, and connection management headers. - Treat 2xx as success.
- Parse JSON responses as structured MCP results.
- Wrap non-JSON responses as MCP text content.
- Return an empty 2xx response as
{ "result": "success" }.
Target resolution:
- Prefer
targetHostfor direct calls. - Otherwise use
serviceId,protocol, andenvTagthrough the existing portal registry discovery client. - If neither is available, return a tool execution error.
MCP Proxy Tool Execution
For apiType: mcp:
- Resolve the target base URL the same way as HTTP tools.
- Ensure a backend MCP session exists for the current gateway session and
backend target. If none exists, initialize the backend MCP endpoint and store
the returned backend
Mcp-Session-Id. - POST to the configured backend
path. - Pass through the inbound agent headers to the backend MCP server, with
transport-specific headers regenerated for the new outbound request.
Replace any frontend gateway
Mcp-Session-Idwith the mapped backend session id for this backend target. - Send a backend JSON-RPC request:
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "tool-name",
"arguments": {}
}
}
- If the backend returns
error, map it to-32000. - If the backend returns
result, return it to the caller. - On frontend session termination or expiry, close the backend MCP session.
This preserves the Java McpProxyTool behavior while using Rust’s typed
JSON-RPC models where possible and adds the MCP session mapping required by
stateful backend MCP servers.
Configuration Loading
The router should be loaded as a normal light-fabric module:
config-server product values
-> mcp-router.yml placeholders
-> light-gateway startup
-> light-pingora mcp router state
Example product values:
mcp-router.enabled: true
mcp-router.path: /mcp
mcp-router.tools:
- name: get_pet
description: Get a pet by id.
targetHost: http://petstore:8080
path: /v1/pets
method: GET
inputSchema:
type: object
properties:
id:
type: string
Example handler.yml path wiring:
handlers:
- correlation
- metrics
- cors
- jwt
- mcp
- proxy
chains:
default:
- correlation
- metrics
- cors
- jwt
- proxy
mcp:
- correlation
- metrics
- cors
- jwt
- mcp
paths:
- path: /mcp
method: POST
exec:
- mcp
- path: /mcp
method: GET
exec:
- mcp
defaultHandlers:
- proxy
The exact chain names are product choices. The important point is that /mcp
can have a narrow chain while normal API proxy traffic keeps the normal proxy
chain.
Module Registry
The MCP router should register its configuration with the module registry:
- module name:
mcp-router - config files:
mcp-router.yml, withmcp-router.yamlas compatibility fallback - enabled status
- configured path
- tool count
- tool names
The module registry should mask any future secret fields in toolMetadata,
headers, or credential configuration.
Reload behavior:
- Reload
mcp-router.yml. - Validate duplicate tool names, missing paths, unsupported methods, and target resolution fields.
- Build a new immutable router state.
- Swap the runtime state atomically.
- Report the updated module registry status.
Security And Policy
The first layer of protection should be the handler chain. Products can place
JWT, API key, basic auth, unified security, CORS, rate limit, and header
handlers before or after mcp as needed.
Because MCP Streamable HTTP is browser-reachable, the mcp handler must also
validate the Origin header according to the configured CORS or security
policy. Invalid origins should fail before tool execution.
Fine-grained tool authorization should be added after the base router:
- Reuse the existing light-4j
access-control.ymlmodel as the compatibility contract.access-control.ymlcontrolsenabled,accessRuleLogic,defaultDeny,defaultInclude, andskipPathPrefixes;rule.ymlprovidesruleBodiesandendpointRules. - Make the access policy endpoint stable. Java uses the tool
endpointfield, such as/weather@get; when omitted, Rust derives{path}@{method}. - Include correlation id, caller claims, request headers, tool name, endpoint, and arguments in the policy input.
- Support default deny when access control is enabled and no
req-accrule matches. - Support fail-closed row filtering when response filtering is enabled and no caller claim matches a configured row-filter entry.
- Provide built-in Rust actions compatible with the Java class names used by
current config:
RoleBasedAccessControlAction,ResponseColumnFilterAction, andResponseRowFilterAction.
Response filtering should be implemented as a second policy stage:
- Apply policy after backend execution and before JSON-RPC response emission.
- Support both
structuredContentand single text content responses, matching Java’s behavior. - Match endpoint rules exactly first, then Java-style path templates and
parent path entries such as
/v1/accounts@getfor/v1/accounts/123@get.
req-acc And res-fil Rule Design
The MCP router should treat endpoint rules as two separate policy stages:
req-acc: request access rules. These run before backend tool execution. They decide whether the caller can invoke the tool at all.res-fil: response filter rules. These run after backend tool execution and before the JSON-RPC result is sent back to the caller. They can remove rows, remove columns, or otherwise reduce the returned payload.
Both stages use Light-Rule with CEL rule conditions. Light-Fabric only supports
CEL conditions for new rule execution. The old native condition-row format is a
legacy Java yaml-rule format and is not the Light-Fabric runtime contract. CEL
keeps the rule predicate explicit while still letting the Java-compatible action
classes perform the stable role, row, and column filter behavior. In this model,
CEL decides whether a rule is eligible to run, and permission carries the
endpoint-specific role, row, and column policy values.
The endpoint key must be stable and should not include the query string. For example, the demo request:
curl -s "http://127.0.0.1:8086/offers?segment=premium&state=ON&category=travel"
maps to the endpoint rule key /offers@get. The query parameters are part of
the tool arguments and can be inspected by CEL through toolArguments.
For the offer demo, the backend should return enough rows to prove filtering, for example:
[
{
"offerId": "OFFER-TRAVEL-01",
"title": "Premium travel credit",
"segment": "premium",
"state": "ON",
"category": "travel",
"priority": 1,
"active": true
},
{
"offerId": "OFFER-TRAVEL-50",
"title": "Premium lounge bundle",
"segment": "premium",
"state": "ON",
"category": "travel",
"priority": 50,
"active": true
},
{
"offerId": "OFFER-TRAVEL-OLD",
"title": "Retired companion fare",
"segment": "premium",
"state": "ON",
"category": "travel",
"priority": 10,
"active": false
}
]
Two roles are enough to demonstrate the behavior:
offer-viewer: can invoke the offer tool, but can only see rows wherepriority < 50andactive == true. Theactivecolumn must not be returned.offer-admin: can invoke the offer tool and can see every row and every column, including inactive offers and all priorities.
Example rule mapping:
ruleBodies:
allowOfferSearch:
common: Y
ruleId: allowOfferSearch
ruleName: Allow offer search
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
filterOfferRows:
common: Y
ruleId: filterOfferRows
ruleName: Filter offer rows
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200
&& responseBody != ""
&& auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.ResponseRowFilterAction
filterOfferColumns:
common: Y
ruleId: filterOfferColumns
ruleName: Filter offer columns
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200
&& responseBody != ""
&& auditInfo.subject_claims.ClaimsMap.role != null
actions:
- actionClassName: com.networknt.rule.ResponseColumnFilterAction
endpointRules:
/offers@get:
req-acc:
- allowOfferSearch
res-fil:
- filterOfferRows
- filterOfferColumns
permission:
roles: offer-viewer offer-admin
row:
role:
offer-viewer:
- colName: priority
operator: "<"
colValue: 50
- colName: active
operator: "="
colValue: true
col:
role:
offer-viewer: offerId,title,segment,state,category,priority
res-fil order matters. Row filtering must run before column filtering when a
row predicate depends on a column that may be hidden from the final response.
In the example above, active is needed to select rows but is then removed for
offer-viewer.
res-fil always executes as a sequential pipeline. accessRuleLogic only
controls how multiple req-acc rules are combined; it does not apply to
response filters. The pipeline should parse the MCP result JSON once, pass the
same mutable JSON value through each res-fil action, and serialize it back
into the MCP result once after all filters complete.
The current compatibility actions support the existing permission model:
rolesis used byRoleBasedAccessControlAction.row.role,row.group,row.position,row.attribute, androw.userprovide row filters for the matching caller claim. Additional dimensions are supported when their claim names are declared inclaimMappings.- If
permission.rowexists but no row-filter entry matches the caller’s claims,access-control.defaultIncludedecides the miss behavior.defaultInclude: falsereturns no rows and is the secure default.defaultInclude: truekeeps the legacy include-all behavior. - Row filters support
=,!=,<,>,<=,>=,in, andnot in. col.role,col.group,col.position,col.attribute, andcol.userprovide the returned field list for the matching caller claim. Mapped custom dimensions are also supported. A field list prefixed with!is a remove list; otherwise it is a keep list.- Column filtering must apply to top-level JSON objects as well as top-level
arrays and object payloads containing
items. Row filtering treats a top-level JSON object as a single candidate row and returns an MCP tool error withisError: truewhen it is denied.
This should remain the default design for Java and Rust parity. The Java row
action must apply the same single-object behavior instead of returning maps
unchanged. If a policy needs arbitrary per-row CEL predicates, use the explicit
ResponseCelRowFilterAction rather than changing the declarative permission
format of ResponseRowFilterAction.
CEL must not directly manipulate the MCP result JSON. The rule-level CEL
expression decides whether a res-fil rule applies; the action performs the
mutation. This keeps result extraction, structuredContent handling,
text-content handling, single-pass JSON parsing, failure behavior, and audit
logging inside tested Rust pipeline and action code.
For MCP, defaultInclude is evaluated inside the same response-filter pipeline
as HTTP access-control. The router must apply it before the final JSON-RPC
result is emitted:
- Execute the backend tool call.
- Extract the response payload from
structuredContentor supported text content. - Run
res-filactions in order. - If
ResponseRowFilterActionsees configured row filters but no matching caller claim, retain no rows whendefaultInclude: false. - Serialize the filtered result back into the MCP response.
This makes direct HTTP endpoints and MCP-routed tools share the same row-filter security behavior.
A CEL-aware row action can be added when the permission row-filter format is not expressive enough:
ruleBodies:
filterOfferRowsWithCel:
ruleId: filterOfferRowsWithCel
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
statusCode == 200 && responseBody != ""
actions:
- actionClassName: com.networknt.rule.ResponseCelRowFilterAction
actionValues:
rowExpression: >
auditInfo.subject_claims.ClaimsMap.role == "offer-admin"
|| (row.priority < 50 && row.active == true)
Even in that case, CEL only returns a boolean for each row. The action owns row iteration and mutation of the parsed response value; the response-filter pipeline owns final serialization and updating the MCP response.
The row action should not deep-clone the full rule context for every response
row. It should use a child CEL context that shadows row, or reuse one mutable
context and update only the row binding for each evaluation. If a row-level CEL
evaluation fails because the row is missing a referenced field, the action should
drop that row and continue. Invalid action configuration, such as an expression
that fails to compile, should fail the whole filter action closed.
Masking and tokenization handling:
- Preserve Java schema extensions:
x-mask,x-mask-pattern, andx-tokenize. - Parse these extensions from
inputSchemaasserde_json::Value. - Apply schema-driven
x-maskrequest masking before backend tool execution. - Keep
x-tokenizeas a future extension point. Do not call a tokenization service until the portal-service tokenization protocol is finalized. - Do not hardcode a tokenization service URL. The tokenization client should be
designed after
light-tokenizationis migrated intoportal-service/apps/portal-service, whether the final protocol is JSON-RPC, MCP, or gRPC.
Per-tool outbound headers would mean headers that the MCP router adds from tool
configuration when it calls a specific backend target, for example a configured
Authorization, X-API-Key, tenant routing header, or vendor-specific version
header. We do not need that feature. The required behavior is header
pass-through: backend tool calls receive the headers that came from the agent,
while the HTTP client regenerates only the transport-specific headers required
for a valid outbound request. MCP session headers are not normal pass-through
headers. The gateway owns the frontend Mcp-Session-Id and maps it to backend
session ids when an upstream MCP server is involved.
Relationship To Existing Runtime MCP
light-runtime already has RuntimeMcpHandler for runtime management tools.
That should remain internal and registry-facing.
The gateway MCP router should not automatically expose runtime management tools. If a product needs that bridge later, add an explicit configured tool type, for example:
apiType: runtime
That keeps public agent-facing tools separate from management tools and avoids accidentally exposing cache, module, or service operations through a public gateway route.
Phased Implementation
Phase 1: Core Router
- Add
mcp-router.ymlconfig parsing inlight-pingora. - Accept
toolsas either a YAML array or a JSON string to match Java config server injection behavior. - Add immutable tool map validation.
- Implement the base Streamable HTTP single endpoint: unary
POST /mcp,Acceptvalidation forapplication/jsonandtext/event-stream,202 Acceptedfor accepted notifications, and405for unsupported methods. - Implement JSON-RPC
initialize,notifications/initialized,tools/list, andtools/call. - Implement direct
targetHostHTTP tools. - Pass through agent request headers to direct HTTP and backend MCP tool calls, except MCP session headers that the gateway must map separately.
- Wire the existing
mcphandler id inlight-gateway. - Register module status and config with the module registry.
- Add parser and handler tests.
Status: implemented.
Phase 2: Discovery And MCP Proxy
- Resolve
serviceId,protocol, andenvTagthrough the existing portal registry discovery client. - Implement
apiType: mcpbackend proxy tools. - Add reload support with atomic state swap.
- Add tests with fake discovery and backend MCP responses.
Status: implemented.
Phase 3: Streamable HTTP Streaming
- Add streamed
text/event-streamresponses fromPOST /mcpfor long-running tool calls or server-to-client messages related to the originating request. - Add optional
GET /mcpserver-to-client streams on the same endpoint. - Track frontend sessions when
Mcp-Session-Idis issued. Return405for standalone GET streams until server-initiated messages are implemented. - Add tests for content negotiation,
202 Acceptednotifications, streamed POST responses, and optional GET behavior.
Status: implemented.
Phase 4: Policy, Filtering, Masking
- Add tool-level authorization using the
access-control.ymlcompatibility contract. - Add response filtering for structured and text MCP results.
- Add schema-driven request masking.
- Add MCP tool-call log fields for tool name, endpoint, duration, status, and policy outcome.
Status: implemented for access control, response filtering, and request masking. Tokenization is deferred until the portal-service tokenization client is designed.
Phase 5: Stateful MCP Backend Sessions
- Add a gateway session store keyed by frontend
Mcp-Session-Id. - Validate later client requests against the gateway session.
- For
apiType: mcp, maintain backend session mappings keyed by gateway session id and backend target identity. - Lazily initialize backend MCP sessions by sending backend
initialize, capturing backendMcp-Session-Id, and sendingnotifications/initialized. - Replace the frontend session id with the mapped backend session id on upstream MCP calls.
- Terminate mapped backend MCP sessions when the frontend session is deleted, expires, or the gateway shuts down.
- Add tests for frontend session validation, backend session creation, backend session reuse, and backend session termination.
Status: implemented for the in-memory frontend session store, configurable
global and per-client session caps, 30-minute lazy idle expiry, lazy backend
initialization, backend Mcp-Session-Id mapping, backend session reuse, and
explicit DELETE teardown. Shutdown cleanup, external session storage, and
multi-backend isolation tests remain future hardening for multi-pod
deployments.
Testing Strategy
- Config tests:
- empty config
- disabled config
- duplicate tool names
toolsas YAML arraytoolsas JSON stringinputSchemaas object and string
- JSON-RPC tests:
initializenotifications/initialized- notification returns
202 Accepted tools/listtools/listwithqueryandintent- missing method
- invalid params
- malformed JSON
- Streamable HTTP tests:
- single
/mcpendpoint handles POST - POST validates
Accept - unsupported methods return
405 - optional GET stream returns
405until enabled
- single
- Tool execution tests:
- direct
GETwith encoded arguments - direct
POSTwith JSON arguments - non-JSON backend response
- empty 2xx backend response
- non-2xx backend response
- agent headers are forwarded to backend tool calls
- discovered service target
- backend MCP proxy success and error
- direct
- Handler chain tests:
/mcpconsumed bymcp- non-MCP path continues to the next handler
- disabled router does not expose
/mcp
- Reload tests:
- tool added
- tool removed
- invalid reload keeps the prior good state
Remaining Decisions
- Confirm whether Phase 1 includes only unary Streamable HTTP POST or also streamed POST responses.
- Decide the tokenization client protocol after
light-tokenizationis migrated intoportal-service/apps/portal-service. - Map the Java
access-control.ymlschema to Rust policy execution and define how it will be shared by REST, JSON-RPC, and MCP handlers.
MCP 2026-07-28 Dual-Profile Gateway Design
Status
The final 2026-07-28 specification is published. The gateway implements the
dual-profile architecture, but final-contract corrections and release
qualification remain open under issue #379.
Stateless support is built in, without an enable/disable switch (issue #380).
The stateless revision list must be non-empty and contain only implemented
revisions; it is not an alternate disable switch. The removed enabled field
is rejected as unknown. Development configurations must remove it; binary
rollback requires the matching configuration, not mixed-version compatibility. Final specification artifacts are pinned in the implementation
Phase 7 manifest; conformance and operational evidence remain pending.
The implementation plan in implementation/light-gateway/2026-07-16-Mcp20260728DualProfileGatewayImplementationPlan.md
Section 25 is the current R0-R6 execution sequence. Historical July closures
do not establish final compliance. 2025-11-25 is the compatibility target;
2024-11-05 retirement is effective immediately for development (R5, 2026-09-09).
The design-to-implementation phase crosswalk is: design 0 → implementation 0–1; design 1 → 2–3; design 2 → 4–5; design 3 → 6; design 4 → 9; design 5 → 8; design 6 → 7. Phase 9 remains explicitly closed without an automatic backend probe or stateless-to-legacy bridge.
R1 final-contract corrections are implemented: Region generates
Mcp-Param-Region; header annotations require direct properties paths.
Legacy non-object output is exposed under an object value member with a
matching, reference-rebased output schema. March 2025 uses text output and
bounded session-authenticated batches (maximum 64 messages, no initialize in
a batch). Modern IDs/errors and subscription metadata are validated.
Primary specification sources:
- 2026-07-28 release candidate
- Final protocol changelog
- SEP-2567: Sessionless MCP via Explicit State Handles
- SEP-2575: Make MCP Stateless
- SEP-2243: HTTP Header Standardization
- Final base protocol
- Final Streamable HTTP transport
- Final discovery contract
- Final tools contract
- Final authorization contract
- SEP-1303: Input Validation Errors as Tool Execution Errors
- SEP-1319: Decouple Request Payload from RPC Method Definitions
Specification Interpretation
The consolidated specification defines the wire contract. Individual SEPs explain why a change was made, its security implications, rejected alternatives, and migration guidance, but their proposal text can predate later integration edits.
When sources differ, use this order:
- The final revision’s TypeScript schema, which the specification identifies as the protocol message source of truth.
- The final normative specification pages and generated JSON Schema.
- The final revision changelog and error-code registry.
- Final SEP text for rationale and requirements not changed during integration.
- Release-candidate blog examples and non-normative guidance.
The published final schema occupies the first position. Every difference from
the historical release candidate must be reviewed before production enablement.
For example, SEP-2575 proposal text and the consolidated draft differ in the
normative strength of clientInfo and in error-code allocations. The gateway
must implement the consolidated schema rather than preserve superseded SEP
examples.
The release coverage matrix later in this document records every consolidated
changelog area that affects light-gateway. A row marked deferred still needs
an explicit capability or compatibility boundary; deferred does not mean that
arbitrary messages may pass through unchecked.
Executive Decision
light-gateway will support the legacy stateful and 2026-07-28 stateless MCP
profiles at the same configured endpoint, normally /mcp.
The gateway selects a profile from the request’s protocol contract. It must
not select the stateless profile merely because Mcp-Session-Id is absent.
An absent session id can also mean that a legacy request is malformed, and
treating it as stateless would bypass the legacy session boundary.
Both profiles terminate at one transport-neutral application core for tool visibility, authorization, request masking, execution, response filtering, auditing, and metrics. Protocol adapters own only the lifecycle, wire fields, response envelope, and transport behavior specific to their revision.
The first production milestone will support server/discover, tools/list,
and tools/call for stateless clients. Long-lived subscriptions/listen,
multi-round-trip requests, and an optional stateless-to-legacy backend bridge
are separate gates and must not be advertised before they are implemented.
Context
The existing Rust MCP router implements a stateful Streamable HTTP facade:
- A client calls
initialize. light-gatewaycreates a gateway-owned frontend session and returns anMcp-Session-Id.- Every later request is validated against that session.
- For an
apiType: mcptool, the gateway lazily initializes a backend MCP session and maps it to the frontend session and backend target. - Deleting or expiring the frontend session terminates its backend sessions.
The current implementation keeps frontend and backend sessions in the
McpRouterRuntime. It preserves them across configuration reloads within the
same process. It also maintains an authorization-aware tools-list cache.
The remediation implementation retains 2025-11-25, 2025-06-18 and
2025-03-26 Streamable HTTP, with November 2025 as the legacy default and
revision-specific result/schema handling. March supports session-bound batches.
2024-11-05 is rejected by frontend, backend and configuration policy. The product owner confirmed no consumers and authorized immediate retirement
on 2026-09-09. The local all-in-lt gateway has been upgraded and its rejection
verified; no consumer migration is required. Original two-endpoint HTTP+SSE is unsupported.
The 2026-07-28 profile changes that lifecycle:
initializeandnotifications/initializedare removed.Mcp-Session-Idis removed.- Protocol version and client capabilities travel with every request.
server/discoverreports supported versions, server capabilities, and server identity.- Streamable HTTP requests expose routing fields through
Mcp-Methodand, when applicable,Mcp-Name. - List results include freshness and cache-scope information.
- Ordinary results carry
resultType: "complete". - Long-lived server notifications use a POST response stream created with
subscriptions/listenrather than the HTTP GET endpoint.
Supporting both generations therefore requires two versioned protocol adapters, not an optional-session branch inside one wire contract.
Goals
- Serve legacy and
2026-07-28clients concurrently on one MCP path. - Preserve existing legacy initialization, session validation, backend session mapping, teardown, authorization, and filtering behavior.
- Make every stateless request independently understandable, authenticated, authorized, bounded, and routable to any gateway replica.
- Reuse one application core for tool listing and tool execution so the two profiles do not drift in policy behavior.
- Support HTTP tools, legacy MCP backends, and stateless MCP backends through explicit compatibility rules.
- Return truthful capabilities and explicit incompatibility errors rather than silently downgrading or emulating unsupported semantics.
- Serve configured stateless protocol versions without a separate enable switch.
Non-Goals
- Do not add an application setting that changes the meaning of one protocol version between stateful and stateless operation.
- Do not infer protocol profile from user agent, HTTP version, connection reuse, missing headers, or backend topology.
- Do not remove the legacy profile while supported clients still depend on it.
- Do not pool hidden legacy backend sessions by user identity for stateless clients.
- Do not interpret or authorize arbitrary application state handles in the gateway. Handles are normal tool arguments and results.
- Do not implement prompts, resources, sampling, roots, logging, MCP Apps, Tasks, or multi-round-trip requests merely because the protocol schema can represent them.
- Do not advertise
subscriptions/listenuntil the Pingora response path can keep an SSE response open and cancel it safely.
Terminology
| Term | Meaning |
|---|---|
| Legacy profile | MCP 2025-11-25 and supported earlier revisions using initialization and protocol-level sessions |
| Stateless profile | MCP 2026-07-28, with per-request version/capabilities and no protocol-level session |
| Frontend | The MCP client-to-light-gateway side |
| Backend | An HTTP API or MCP server invoked by a configured gateway tool |
| Frontend session | A gateway-owned legacy client session identified by Mcp-Session-Id |
| Backend session | A legacy upstream MCP session owned by the gateway |
| Explicit state handle | An opaque application identifier returned by a tool and supplied to later tool calls; it is not an MCP protocol primitive |
| Principal fingerprint | A stable, non-secret digest of the authenticated subject, issuer, tenant, and other identity fields required to bind state or caches |
Protocol Profiles
| Concern | Legacy stateful profile | 2026-07-28 stateless profile |
|---|---|---|
| Lifecycle | initialize, then notifications/initialized | No initialization handshake |
| Version | Negotiated and retained in the session | Sent in the HTTP header and request params._meta |
| Client capabilities | Retained from initialization | Supplied for every request |
| Client identity | Supplied during initialization | clientInfo SHOULD be supplied in request params._meta |
| Gateway session | Mcp-Session-Id | Prohibited |
| Discovery | Initialization result | server/discover |
| Tool lists | May be session-sensitive | Must not vary by connection; may vary by authenticated principal or deployment state |
| Application state | May exist behind a legacy session | Explicit tool arguments and server-minted handles |
| Server notifications | Legacy Streamable HTTP behavior | subscriptions/listen POST response stream |
| Horizontal routing | Requires affinity or shared session routing | Any replica can handle any ordinary request |
| Stream resumption | Legacy-version behavior | No SSE event replay or Last-Event-ID resumption |
The version adapter must apply the complete contract for its selected profile. It must not mix a legacy lifecycle with stateless result shapes or accept a stateless lifecycle under a legacy version.
HTTP Method Matrix
| HTTP method | Legacy stateful profile | 2026-07-28 stateless profile |
|---|---|---|
POST | Initialize, notifications, requests, and responses permitted by the negotiated legacy revision | Single self-contained JSON-RPC message; the only request entry point |
DELETE | Terminates the identified frontend session and its backend sessions | Not a protocol operation; reject without touching state |
GET | Keep the currently implemented behavior; this gateway returns 405 unless a separately supported legacy transport requires it | No server-notification endpoint; use subscriptions/listen through POST |
| Other methods | 405 Method Not Allowed | 405 Method Not Allowed |
The deprecated HTTP+SSE transport is not added as part of dual-profile support. Compatibility with that older two-endpoint transport requires a separate, explicit design and route so it cannot be confused with Streamable HTTP.
Request Classification
Classification Rules
The gateway classifies a POST only after enforcing request-body limits and parsing a single JSON-RPC message. Batch requests remain unsupported unless a future design explicitly adds them.
Use this ordered decision table:
| Condition | Result |
|---|---|
Method is initialize, no session id, and requested version is an enabled legacy version | Legacy initialization path |
Mcp-Session-Id is present and the request does not claim 2026-07-28 | Legacy session path |
Version header and request params._meta both select enabled 2026-07-28, agree exactly, and required routing headers are valid | Stateless path |
Version header and body select enabled 2026-07-28, with a stale Mcp-Session-Id also present | Select stateless, ignore the legacy header, and never read, mint, echo, or delete session state |
| Legacy non-initialize request has no session id | Reject as missing legacy session id |
Stateless version is present only in the header or only in params._meta | Reject as a header mismatch |
| Version is missing or unsupported and no valid legacy initialization can negotiate it | Reject; do not guess |
The classifier must be a pure, unit-tested component. Tool execution must not start and no frontend or backend state may be mutated until classification, version validation, authentication, header/body validation, and authorization have succeeded.
Stateless HTTP Header Validation
For a 2026-07-28 POST:
Content-Typemust identify JSON and the body must be one UTF-8 JSON-RPC request or notification permitted by the protocol. Client-sent JSON-RPC responses are prohibited because this revision has no server-initiated requests.Acceptmust list bothapplication/jsonandtext/event-stream.MCP-Protocol-Versionis required and must equalparams._meta["io.modelcontextprotocol/protocolVersion"].Mcp-Methodis required for every JSON-RPC request and must equal the JSON-RPCmethod. This revision does not define routing-header requirements for notification POSTs; the gateway must not invent them.Mcp-Nameis required where SEP-2243 defines a named operation, includingtools/call, and must equal the correspondingparams.nameorparams.urivalue.- Header names are compared case-insensitively; method and name values are case-sensitive after decoding. Ambiguous duplicate semantic headers are rejected.
Mcp-NameandMcp-Param-*use the specified visible-ASCII representation. A non-ASCII, control-containing, leading/trailing-whitespace, or literal sentinel-shaped value uses the exact=?base64?{base64-utf8}?=encoding. Decode before comparing with the body; compare integer parameters numerically rather than requiring one decimal spelling.- This sentinel is MCP-specific and is not an RFC 2047 MIME encoded-word.
Implementations must not substitute
=?utf-8?B?...?=or apply MIME header decoding. The lowercase=?base64?prefix and?=suffix are literal, case-sensitive protocol markers. - A missing, malformed, or mismatched required header returns HTTP
400withHeaderMismatchcode-32020. - An unsupported version returns HTTP
400withUnsupportedProtocolVersioncode-32022and includes the requested and supported versions. - A required capability that the client did not declare returns
MissingRequiredClientCapabilitycode-32021.
The core 2026-07-28 protocol defines no client-to-server notification over
Streamable HTTP. Because the first milestone also advertises no extension that
defines one, it rejects notification POSTs as unsupported. If a supported
extension notification is added later, acceptance returns HTTP 202 with no
body; rejection uses an HTTP error and may include an id-less JSON-RPC error.
A JSON-RPC request returns either one application/json object or an SSE
response stream. The adapter never returns a JSON-RPC response to a notification
and always rejects a client-sent JSON-RPC response.
Origin validation happens before JSON parsing or state mutation. When Origin
is present, compare the complete normalized origin against an exact allowlist;
suffix, substring, wildcard-host, and reflected-origin matching are prohibited.
An invalid origin returns HTTP 403. An empty or absent allowlist rejects
requests that carry Origin while still allowing non-browser clients that do
not send the header. This is an MCP transport security boundary, not merely a
response CORS-header concern. The same rule applies to both enabled Streamable
HTTP profiles; legacy compatibility does not weaken it.
Trusted reverse proxies, WAFs, ingresses, and load balancers must preserve the
browser’s Origin header exactly. Stripping or rewriting it is deployment
nonconformance because the gateway cannot reliably distinguish that browser
from a real non-browser client. User-Agent and other spoofable headers are not
acceptable recovery signals. Deployment conformance tests must exercise Origin
preservation through the complete external path.
If a tool schema uses x-mcp-header, the gateway terminates and parses the
request, so it must validate each applicable Mcp-Param-* header against the
tool argument before policy evaluation or execution. A gateway must never
authorize or route on a header value and then execute a different body value.
The final July 28 schema and error registry override release-candidate details if they change before publication.
Legacy Validation Hardening
Legacy behavior remains version-specific, but simultaneous support must not leave the legacy session id as an authorization bearer token.
When creating a frontend session on a protected route, store a principal fingerprint derived from the independently verified request identity. Every session-bound POST or DELETE must recompute and compare the fingerprint before touching the session or a backend. Missing identity, a different identity, or a changed tenant binding fails closed.
An anonymous MCP route must be an explicit product decision rather than the result of missing or failed authentication. An anonymous legacy session uses a separate gateway-derived anonymous client binding for capacity and abuse controls; its cryptographically random session id necessarily remains a bearer capability. A protected and anonymous session must never share the same binding namespace.
The stored fingerprint must contain no bearer token, cookie, CSRF value, or other reusable credential. The gateway may retain the existing client key for capacity accounting, but capacity identity and security identity must be separate concepts.
Common Application Core
Both adapters normalize accepted messages into an effective request context:
#![allow(unused)]
fn main() {
enum FrontendProtocol {
Legacy {
session_id: String,
negotiated_version: String,
},
Stateless,
}
struct EffectiveMcpRequestContext {
protocol: FrontendProtocol,
protocol_version: String,
client_info: Option<ClientInfo>,
client_capabilities: ClientCapabilities,
requested_log_level: Option<LoggingLevel>,
auth: Option<AuthPrincipal>,
correlation_id: Option<String>,
delegation: Option<DelegationClaims>,
}
}
The common core owns:
- Method allowlisting.
- Delegated-authority validation.
- Tool-list visibility and deterministic ordering.
- Request access control.
- Input-schema validation and request masking.
- Backend target resolution and SSRF protection.
- Tool execution and bounded retry policy.
- Response filtering and output-schema handling.
- JSON-RPC application error mapping.
- Audit events, metrics, and safe diagnostics.
The common core returns a protocol-neutral result. The selected adapter adds
the correct result envelope, _meta, cache fields, protocol headers, or
legacy session headers.
Stateless Methods
server/discover
The stateless adapter must implement server/discover and return:
- enabled protocol versions in deterministic preference order;
- only capabilities implemented and enabled by this gateway instance;
- optional instructions that describe the configured tool facade;
resultType: "complete";- gateway server identity in
_meta["io.modelcontextprotocol/serverInfo"]with the normative strength from the final schema; ttlMsandcacheScopebecause discovery is cacheable.
Discovery is independently authenticated and authorization-aware. Its cache key
uses the same principal, protocol, policy, and configuration revisions as the
capability result. The first milestone reports gateway capabilities derived
from enabled handlers and configured tools; it does not depend on call-time
portal-registry discovery. cacheScope defaults to private; public is valid
only when the complete discovery result is identical for every caller. The
advertised TTL must not exceed the internal entry lifetime.
A configuration or policy swap advances its revision and makes old entries
unreachable before a response is served from the new runtime. TTL is an expiry
backstop, not the primary reload-coherence mechanism. If a future dynamically
discovered catalog makes discovery/list results depend on portal-registry
state, that work must first add a monotonic registry generation or canonical
snapshot revision plus a push/watch invalidation signal. The current
request/response DiscoverySnapshot has neither. Until that contract exists,
the gateway must not claim immediate backend-discovery invalidation or
advertise listChanged; a documented short TTL is the only available
staleness bound.
The initial stateless release advertises tools only. It must not advertise prompts, resources, notifications, multi-round-trip requests, Tasks, Apps, or extensions that are not wired through the application core.
Legacy clients continue to use initialize. A dual-version client may probe
server/discover; failure may cause legacy fallback only under the downgrade
rules defined later in this document.
tools/list
Stateless tools/list uses the same authorization-aware visibility logic as
legacy tools/list, with these additional rules:
- The returned order is deterministic.
- The list must not vary because of a connection or prior tool call.
- It may vary by current authenticated principal, scopes, tenant, active policy,
or gateway configuration. The first milestone does not hide or add configured
tools based on call-time backend discovery; backend availability is checked
by
tools/call. - The result contains
resultType: "complete",ttlMs, andcacheScope. cacheScopedefaults toprivatebecause the visible catalog can vary by authenticated principal and access-control policy.
The first milestone returns the complete bounded visible catalog, omits
nextCursor, and rejects a non-empty cursor that it did not issue. It does not
pretend to paginate. If the computed visible catalog exceeds
maxToolsListItems, or its encoded response would exceed
maxResponseBodyBytes, the gateway fails the whole request with the bounded
implementation-defined -32000 resource-limit error locked in Phase 0. The
message states that the visible catalog exceeds a gateway limit and that
pagination is not supported. It must not truncate the catalog, emit a cursor,
or cache a partial result. Cursor generation, integrity, principal binding,
expiry, and reload invalidation require a separate pagination design before
larger visible catalogs are accepted.
The internal cache key must include at least:
- protocol profile and version;
- normalized query or intent parameters;
- authenticated-principal fingerprint;
- relevant forwarded-header fingerprint;
- MCP router configuration revision;
- access-control policy revision;
The initial configured catalog has no backend-discovery component in its cache key. A future discovery-dependent catalog must add the registry generation or snapshot revision described above; a TTL alone is insufficient to claim immediate invalidation.
Cache entries must expire no later than the advertised ttlMs. A policy or
router reload invalidates affected entries before new responses are served.
tools/call
Stateless tools/call is authorized and executed independently. It must not
read, create, touch, or delete a frontend session. A successful ordinary
result includes resultType: "complete" and the server identity fields
required or recommended by the final protocol schema.
Mutating calls are not automatically replayed after ambiguous transport failure. Existing retry metadata remains authoritative, but retries must be limited to operations explicitly declared safe or idempotent.
Stateless calls use a typed outbound-header allowlist. The gateway regenerates
profile routing, correlation, trace, tenant, locale, and backend credential
headers from trusted request context and target configuration. It does not
copy raw frontend X-Forwarded-*, cookies, authorization, backend-specific
credentials, unknown Mcp-*, or arbitrary extension headers to a backend.
Legacy header forwarding remains a separately versioned compatibility contract
and must not be reused as the modern default.
Unsupported Methods
Methods not implemented by the gateway return a normal JSON-RPC
method-not-found response. For the stateless Streamable HTTP profile this is
HTTP 404 Not Found with JSON-RPC code -32601; the body distinguishes a
modern unknown method from a missing legacy HTTP+SSE endpoint. Capabilities
must not imply that those methods are available. This rule is especially
important for deprecated roots, sampling, and logging features and for
extensions that are not part of the first milestone.
Tool and JSON Schema Contract
The current router stores and advertises configured inputSchema values and
uses schema annotations for request masking. That is not equivalent to full
JSON Schema validation. Supporting the 2026-07-28 tools contract requires a
dedicated, bounded schema compilation and validation path.
Schema Loading
At configuration load, before a runtime swap:
inputSchemamust be a valid JSON Schema object and must describe an object at the root. A no-argument tool should use{ "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": {}, "additionalProperties": false }.outputSchema, when present, may describe any JSON value, including arrays, primitives, ornull.- A schema without
$schemauses JSON Schema 2020-12. - JSON Schema 2020-12 is mandatory. Any additional dialect is an explicit, documented configuration choice; an unsupported dialect is rejected rather than interpreted as 2020-12 or treated permissively.
- Local
$refand$defsresolution is supported within the schema document. - Network dereferencing of external
$refURIs is disabled. An unresolved external reference rejects the tool. A future opt-in resolver needs a separate SSRF review, exact host allowlist, byte/depth/time limits, and cache policy. - Schema compilation is bounded by document bytes, nesting depth, subschema count, reference expansions, regular-expression complexity where supported, and a time budget.
Compilation produces immutable validators stored with the tool configuration. Invalid schemas fail the candidate configuration before it becomes active. One invalid configured tool must not silently turn into an unconstrained tool.
Composition does not replace the MCP object-root declaration. allOf,
anyOf, oneOf, conditionals, and local references are supported as siblings
of root type: object. If an older Portal-generated tool contains only
allOf or another composition keyword, add type: object at the same level
as an immediate repair, then regenerate the tool. Remove or reset any stale
selected-tool schema override so it cannot replace the regenerated schema.
An empty string is not an empty schema; no-argument tools use the explicit
closed object above.
Tool Names and Aggregation
Names exposed through the stateless profile are case-sensitive, unique within the gateway catalog, 1 to 128 characters, and limited to ASCII letters, digits, underscore, hyphen, and dot. Enabling the stateless profile fails validation if an exposed configured name violates this contract. Legacy-only deployments may retain their existing names until migrated.
If future dynamic backend discovery introduces collisions, configuration must
provide a stable explicit alias or prefix. Backend serverInfo.name is not a
unique identifier and must not be used as the automatic disambiguation key.
x-mcp-header Validation
At schema load, every x-mcp-header annotation must:
- be non-empty and match the HTTP field-name token syntax;
- contain no control, carriage-return, or line-feed characters;
- be case-insensitively unique within the tool schema;
- annotate only a statically reachable primitive string, integer, or boolean;
- keep integers within the protocol’s safe IEEE-754 integer range;
- not name a pseudo-header, hop-by-hop or proxy-authentication header, HTTP framing header, credential header, or gateway-owned MCP routing/session header;
- not annotate the same property as sensitive or schema-masked data.
Invalid annotations fail a configured tool. If a future backend-discovery
client receives an invalid backend tool definition, it excludes that tool and
emits a bounded warning without exposing schema values. Sensitive parameters,
tokens, secrets, and PII must never be mirrored into Mcp-Param-* headers.
At call time, extract the annotated value according to the final transport
encoding rules and require the regenerated or received header to match the
validated argument exactly. Missing argument and JSON null mean that the
header is absent.
Input and Output Validation
For tools/call, the common core applies this order:
- Apply transport byte/depth limits and authenticate or accept the explicitly anonymous route.
- Parse, classify, validate routing headers, resolve the configured tool, and apply principal/tool-level visibility and coarse authorization.
- Validate the original arguments against
inputSchemawithin the validation budget. - Evaluate delegation and argument-dependent request access policy against the validated, unmasked arguments.
- Apply request masking and invoke the backend.
- Parse the backend result, apply response policy and filtering, and then
validate the final
structuredContentagainstoutputSchemawhen present. - Construct the version-specific result envelope.
The error boundary is intentionally split:
- an unknown tool or malformed
tools/callenvelope that does not satisfy the protocol’sCallToolRequestschema returns a JSON-RPC protocol error; - arguments that fail the selected tool’s
inputSchemareturn a completed tool result withresultType: "complete",isError: true, and bounded, model-actionable content, without backend traffic.
The validation result must identify the failing field and constraint when safe, but must not echo the complete input, secrets, schema-masked values, or unbounded validator diagnostics. Coarse authorization runs first so validation details cannot be used to probe a tool the principal cannot invoke. This SEP-1303 distinction lets a model correct tool arguments while preserving protocol errors for malformed envelopes and unknown tool names.
A configured output schema makes final output conformance a gateway responsibility because the gateway terminates and may filter the backend result. A backend or response filter that produces non-conforming structured content results in a tool error; the gateway must not emit data that contradicts its advertised schema.
structuredContent may be any JSON value. Masking and response filtering must
therefore handle objects, arrays, strings, numbers, booleans, and null without
assuming an object root. When structured content is returned, the gateway
should also provide its serialized JSON in a text content block for backward
compatibility unless the backend result already supplies an equivalent block.
Result Types, MRTR, and Extensions
Result Types
Every successful 2026-07-28 result contains a recognized resultType.
Ordinary results use "complete". A stateless backend response that omits the
field is invalid for that backend version. When consuming a response from an
earlier negotiated backend version, the gateway treats an absent field as
"complete" for backward compatibility.
An unknown result type is invalid unless it belongs to an extension explicitly
supported and negotiated by both sides. The gateway must not relabel an unknown
or input_required result as complete merely to fit a legacy frontend.
Multi Round-Trip Requests
The first stateless milestone does not implement MRTR and must not advertise
the associated client or server capabilities. If a backend returns
resultType: "input_required" before the gateway implements MRTR, the gateway
replaces that backend result with a terminal gateway-generated tool error. For
a modern frontend it has resultType: "complete", isError: true, and the
bounded message light-gateway does not support MCP multi round-trip bridging for this backend. This is not a schema-validation failure and does not relabel
the backend’s input request as a successful complete result. The gateway does
not expose backend requestState, inputRequests, or other opaque retry state
to the frontend.
A later MRTR design must specify:
- the exact supported input-request methods;
- required client-capability validation;
- authorization and integrity protection for opaque
requestState; - bounds on state bytes, input-request count, nesting, and retry count;
- a new JSON-RPC id for every retry while preserving correlation and audit lineage;
- translation behavior for every frontend/backend profile combination;
- how policy is reevaluated on every retry.
Request-scoped notifications or server input requests belong to the response
stream of the initiating request, not to subscriptions/listen.
The per-request io.modelcontextprotocol/logLevel value is never retained as
gateway session state. The gateway must not emit or forward
notifications/message for a stateless request that omitted it. If request-
scoped logging is later implemented, the requested threshold applies only to
that response stream and is independently bounded and filtered; the deprecated
Logging capability and logging/setLevel remain absent.
Extensions
The core ClientCapabilities and ServerCapabilities contain an extensions
map. Extension support is optional, independently versioned, and disabled by
default.
For the first milestone:
server/discoveromits or returns an empty extension map;- unknown client extensions do not change core behavior;
- an operation that requires an unsupported extension fails explicitly;
- extension metadata and result types are not forwarded through the gateway unless a versioned gateway adapter validates and translates that extension;
- MCP Apps and the Tasks extension remain separate designs and are not advertised;
- tool task support is omitted or normalized to
forbiddenunless the Tasks extension is implemented and negotiated.
The gateway is not a transparent byte proxy, so backend extension support does not automatically make the same extension available to frontend clients. Each supported extension needs an owner, version allowlist, capability intersection, resource limits, authorization review, and conformance tests.
Deprecated and Removed Core Features
The stateless profile does not implement removed initialize,
notifications/initialized, ping, or logging/setLevel methods. It does not
add the removed HTTP GET notification endpoint or SSE resumption.
Roots, sampling, and logging are deprecated in this release. Because the
gateway does not currently implement them, it leaves their capabilities absent
rather than introducing new deprecated functionality. The deprecated
HTTP+SSE transport, deprecated includeContext values, and deprecated Dynamic
Client Registration are likewise not added by the MCP router.
Frontend and Backend Compatibility
The gateway is both an MCP server to the frontend and, for apiType: mcp, an
MCP client to the backend. These protocol profiles are independent.
| Frontend profile | HTTP backend | Legacy MCP backend | Stateless MCP backend |
|---|---|---|---|
| Legacy stateful | Existing direct translation | Existing mapped backend session | Translate stored legacy client metadata into each stateless backend request |
| Stateless | Direct translation | Reject by default; optional per-request bridge | Direct stateless proxy |
Controller WebSocket Control Plane Is Separate
The browser control-plane route /ctrl/mcp is not the Streamable HTTP /mcp
endpoint described here. It is routed by websocket-router, remains
payload-opaque at light-gateway, and continues to use the separately frozen
controller JSON/WebSocket contract. This design does not route it through
mcp-router, translate it to the stateless profile, or require
statelessToLegacyBridge.
Only a future tool explicitly configured with apiType: mcp and a controller
Streamable HTTP target would enter the compatibility matrix above. Such a tool
must not be marked sessionIndependent: true merely because controller
operations appear request/response-shaped. The preferred choices are to keep
that frontend/backend path legacy or upgrade the target to stateless; a bridge
still requires the proof and opt-in defined below.
Backend Profile Configuration
Each MCP tool target has an explicit backend profile:
backendMcpProtocol: legacy # legacy, stateless, or auto
sessionIndependent: false
legacy is the compatibility default for existing apiType: mcp tools.
The fields are ignored for apiType: http.
All tools resolving to the same normalized MCP backend target must declare a compatible backend profile. Configuration loading fails if one target is simultaneously declared legacy and stateless or has conflicting bridge properties.
Legacy Frontend to Stateless Backend
This direction is supported. The gateway retains the legacy frontend’s negotiated client information and capabilities, converts them into stateless per-request metadata, generates the required backend headers, and does not create a backend session.
The gateway returns a legacy-shaped result to the frontend. Fields introduced
only in 2026-07-28 are consumed or translated deliberately; they must not be
copied blindly into an older result schema.
Stateless Frontend to Stateless Backend
This is the preferred proxy path. The gateway:
- Re-authorizes the configured gateway tool.
- Resolves and validates the backend target.
- constructs a new backend request using the backend’s supported stateless version;
- regenerates
MCP-Protocol-Version,Mcp-Method,Mcp-Name, and applicableMcp-Param-*headers; - propagates the effective client capabilities and safe trace context;
- filters the backend response before returning it.
Ingress routing headers are never forwarded without regeneration and body/header consistency validation.
Stateless Frontend to Legacy Backend
This direction is rejected by default. A stateless frontend has no lifecycle scope that can safely own, route, or terminate a legacy backend session. Pooling backend sessions by authenticated principal would mix independent agents, conversations, browser tabs, or subagents and recreate hidden application state.
An optional perRequest bridge may be implemented later only when all of these
conditions hold:
- the administrator explicitly enables the bridge;
- the tool is declared
sessionIndependent: true; - one request can be completed without state from a prior backend call;
- initialize, call, and delete are bounded by independent deadlines;
- backend session creation has separate global and per-principal limits;
- every success and failure path attempts backend teardown;
- metrics expose incomplete teardown without logging session ids.
The bridge performs initialize, one operation, and delete. It must never be
selected silently by auto discovery.
Backend auto Discovery
auto is optional and must be conservative:
- Probe
server/discoverusing the preferred enabled stateless version. - Cache the result by normalized target, principal fingerprint, configuration revision, and a bounded TTL.
- Select stateless only after a valid discovery response.
- Fall back to legacy only for an explicit unsupported-version or method-not-found response indicating an older server.
Do not fall back after authentication or authorization rejection, TLS failure, timeout, malformed response, header mismatch, DNS/SSRF rejection, or HTTP 5xx. Those failures are terminal for that attempt because fallback could become a downgrade path.
Explicit Application State Handles
The stateless protocol does not prohibit stateful applications. A backend may
return an opaque handle such as basket_id or browser_id, and later tools may
accept that handle as an ordinary argument.
The gateway does not introduce a generic handle type or handle registry. It continues to authorize the tool call, validate the input schema, apply masking, and filter the result. The backend that owns the handle must:
- validate
(handle, auth_context)on every call; - avoid treating possession as authorization when authentication exists;
- document lifetime and recovery behavior in the tool description;
- return a useful expired-handle error;
- provide bounded expiry and optional cleanup tools;
- use at least 128 bits of cryptographic entropy and a bounded lifetime when an unauthenticated handle necessarily acts as a bearer capability.
Handles, session ids, request ids, and bearer credentials must not be metric labels or appear in normal logs.
Authentication and Authorization
The MCP handler remains inside the normal light-gateway handler chain. The
security or unified-security handler must establish McpRequestContext.auth
before the MCP handler runs on a protected route.
For both profiles:
- complete authentication, or an explicit anonymous-route decision, before protocol state mutation or backend traffic;
- apply delegation binding before the requested operation;
- evaluate tools-list visibility against the current request identity;
- evaluate request access control before tool execution;
- apply response filtering after backend execution;
- fail closed when policy is unavailable under a default-deny deployment;
- forward only explicitly allowed identity/delegation material to a backend;
- strip and regenerate hop-by-hop, MCP routing, protocol, and session headers.
For the stateless profile, every protected request carries fresh authorization input and must be independently authenticated and authorized. An explicitly anonymous request is still independently rate-limited and evaluated against the route’s anonymous policy. No previous request, connection, discovery response, or subscription grants authority to a later request.
For the legacy profile on a protected route, the current request must both authenticate successfully and match the session’s stored principal fingerprint. Session validation does not replace current-token expiry, revocation, audience, issuer, or scope validation.
Frontend Resource-Server Boundary
For a protected MCP route, light-gateway is the OAuth resource server even
though the MCP router delegates token parsing to security or
unified-security. The product deployment must provide:
- OAuth Protected Resource Metadata for the canonical MCP resource URI;
- exact audience/resource validation for every bearer token;
Authorization: Beareron every protected HTTP request, never a query-string access token;- HTTP
401with an appropriateWWW-Authenticatechallenge for a missing, invalid, or expired token; - HTTP
403and aninsufficient_scopechallenge with the required scope set when the authenticated principal lacks permission; - a
resource_metadatalink and scope guidance consistent with the canonical MCP resource; - bounded step-up behavior on clients, without repeatedly replaying an ambiguous mutation.
JSON-RPC authorization errors may accompany the HTTP response where permitted, but they do not replace the required HTTP status and challenge headers.
Backend Client and Token Audience Boundary
For an apiType: mcp target, light-gateway is also an MCP client. A bearer
token accepted for the frontend gateway resource must not be copied to a
different backend MCP resource merely because it arrived in an agent header.
That would violate audience binding and create a confused-deputy path.
Each backend target must select one credential strategy:
- forward a caller token only when independent validation proves that the backend is an intended audience/resource for that exact token;
- exchange or mint a bounded delegated token for the backend resource;
- use a configured service credential when the call is intentionally performed as the gateway rather than the end user;
- use no credential only for an explicitly anonymous backend.
The strategy is part of the normalized backend identity and cache key. Tokens,
refresh tokens, client secrets, PKCE verifiers, and registered client metadata
are owned by the security/client runtime, not stored in McpGatewaySession,
backend discovery caches, or tool configuration.
For caller forwarding, the gateway security boundary verifies aud against
the normalized configured backendResource before opening backend traffic;
backend validation is defense in depth, not the gateway’s authorization
decision. String and array claims use exact audience membership. Missing or
mismatched audience evidence fails closed. An opaque token cannot use caller
mode unless trusted introspection returns the required audience evidence.
Release Authorization Dependencies
The MCP router consumes an authenticated principal, but the complete release
also changes OAuth behavior. Before claiming 2026-07-28 compliance, the
relevant light-fabric security and client modules must verify or explicitly
defer:
- authorization-response
issvalidation against previously validated issuer metadata; - binding persisted client credentials to the issuer that created them;
- correct OpenID Connect
application_typewhen deprecated Dynamic Client Registration is used for compatibility; - Client ID Metadata Documents as the preferred dynamic registration model;
.well-knownprotected-resource and authorization-server discovery rules;- confidential refresh-token storage, rotation requirements for public clients,
and correct optional
offline_accessbehavior; - bounded scope accumulation and step-up retries.
These are cross-cutting security dependencies rather than duplicate MCP-router implementations. Their release-gate evidence must nevertheless be linked from the coverage matrix.
Response Model and Streaming
The existing response model stores a complete Vec<u8> and a streamed
boolean. The Pingora writer sends that body once with end = true. That is
sufficient for a single JSON response or one buffered SSE frame, but it cannot
implement subscriptions/listen.
Before adding subscriptions, replace it with an explicit response body:
#![allow(unused)]
fn main() {
enum McpResponseBody {
Empty,
Buffered(Bytes),
Stream(McpResponseStream),
}
struct McpResponseStream {
receiver: BoundedReceiver<Bytes>,
cancellation: CancellationToken,
}
}
The gateway writer must keep a streaming response open, apply backpressure, detect disconnect, cancel producers, and finish exactly once. Buffered SSE and long-lived SSE must not share a misleading boolean flag.
For a stateless SSE response, closing the HTTP stream cancels that request;
notifications/cancelled is not expected on Streamable HTTP. The writer stops
work as soon as practical and emits nothing after cancellation. It also sends
X-Accel-Buffering: no. A long-lived subscription may emit bounded SSE comment
keep-alives to survive intermediary idle timeouts; comments carry no JSON-RPC
meaning and consume the subscription byte/rate budget.
subscriptions/listen
When implemented, a client sends subscriptions/listen through POST and
explicitly requests supported notification types. The response is a long-lived
SSE stream.
The gateway must:
- Authorize the listen request independently.
- Enforce global and per-principal subscription limits.
- Open a bounded event channel.
- Send
notifications/subscriptions/acknowledgedas the first JSON-RPC message. - Identify the subscription with the original request id.
- Tag each emitted notification with
io.modelcontextprotocol/subscriptionId. - Emit only notification types requested by the client and supported by the gateway.
- Cancel the producer when the HTTP stream closes, expires, reloads incompatibly, or encounters a slow consumer.
On deliberate server teardown, the gateway sends the empty
subscriptions/listen result (with the original request id and the modern
complete discriminator) before closing the SSE response. A saturated slow
consumer may instead observe a remote close when the bounded channel cannot
admit that terminal result; the gateway never grows or replays the queue to
make graceful close succeed.
The first supported notification should be toolsListChanged, produced after
a successful MCP router or relevant policy reload. Prompts and resources remain
unadvertised until the gateway owns equivalent event sources.
A disconnected stream is not resumable. The client must create a new request
with a new JSON-RPC id, re-fetch authoritative state when necessary, and
re-subscribe. The gateway ignores Last-Event-ID and keeps no replay buffer for
this profile.
An access token does not gain an indefinite lifetime because its response stream remains open. A protected subscription closes at token expiry, policy revocation convergence deadline, configured maximum duration, or gateway shutdown, whichever occurs first. Reconnection performs fresh authentication, authorization, discovery/list rehydration when needed, and subscription creation.
Configuration
The following is the locked configuration contract. All new fields have Serde defaults so existing configuration continues to load unchanged.
enabled: ${mcp-router.enabled:true}
path: ${mcp-router.path:/mcp}
maxSessions: ${mcp-router.maxSessions:10000}
maxSessionsPerClient: ${mcp-router.maxSessionsPerClient:100}
maxRequestBodyBytes: ${mcp-router.maxRequestBodyBytes:1048576}
maxResponseBodyBytes: ${mcp-router.maxResponseBodyBytes:4194304}
maxJsonDepth: ${mcp-router.maxJsonDepth:128}
originAllowlist: ${mcp-router.originAllowlist:[]}
schema:
defaultDialect: ${mcp-router.schema.defaultDialect:https://json-schema.org/draft/2020-12/schema}
allowExternalRefs: ${mcp-router.schema.allowExternalRefs:false}
maxSchemaBytes: ${mcp-router.schema.maxSchemaBytes:1048576}
maxDepth: ${mcp-router.schema.maxDepth:64}
maxSubschemas: ${mcp-router.schema.maxSubschemas:4096}
maxConcurrentValidations: ${mcp-router.schema.maxConcurrentValidations:32}
validationWatchdogMs: ${mcp-router.schema.validationWatchdogMs:50}
protocols:
legacy:
enabled: ${mcp-router.protocols.legacy.enabled:true}
versions: ${mcp-router.protocols.legacy.versions:["2025-11-25", "2025-06-18", "2025-03-26"]}
stateless:
versions: ${mcp-router.protocols.stateless.versions:["2026-07-28"]}
discoverTtlMs: ${mcp-router.protocols.stateless.discoverTtlMs:30000}
discoverCacheScope: ${mcp-router.protocols.stateless.discoverCacheScope:private}
maxDiscoverCacheEntries: ${mcp-router.protocols.stateless.maxDiscoverCacheEntries:1024}
toolsListTtlMs: ${mcp-router.protocols.stateless.toolsListTtlMs:30000}
toolsListCacheScope: ${mcp-router.protocols.stateless.toolsListCacheScope:private}
maxToolsListCacheEntries: ${mcp-router.protocols.stateless.maxToolsListCacheEntries:4096}
maxToolsListItems: ${mcp-router.protocols.stateless.maxToolsListItems:1024}
maxConcurrentRequests: ${mcp-router.protocols.stateless.maxConcurrentRequests:1024}
maxConcurrentRequestsPerPrincipal: ${mcp-router.protocols.stateless.maxConcurrentRequestsPerPrincipal:32}
maxConcurrentBackendCallsPerTarget: ${mcp-router.protocols.stateless.maxConcurrentBackendCallsPerTarget:32}
maxSubscriptions: ${mcp-router.protocols.stateless.maxSubscriptions:10000}
maxSubscriptionsPerPrincipal: ${mcp-router.protocols.stateless.maxSubscriptionsPerPrincipal:4}
maxSubscriptionDurationMs: ${mcp-router.protocols.stateless.maxSubscriptionDurationMs:900000}
statelessToLegacyBridge: ${mcp-router.protocols.stateless.statelessToLegacyBridge:reject}
tools: ${mcp-router.tools:[]}
The legacy version list above is the active compatibility baseline for this release.
Example backend declaration:
tools:
- name: weather
description: Get weather information
apiType: mcp
targetHost: https://weather.internal
path: /mcp
method: call
backendMcpProtocol: stateless
sessionIndependent: true
backendCredentialMode: service
backendResource: https://weather.internal/mcp
inputSchema:
type: object
properties:
city:
type: string
Configuration validation must reject:
- no enabled protocol versions;
2026-07-28listed under the legacy adapter;- a legacy version listed under the stateless adapter;
- an invalid or non-normalized origin or a wildcard/suffix origin rule;
- unsupported
cacheScopeor bridge values; allowExternalRefs: truewithout a separately approved resolver policy;- invalid, unsupported-dialect, or over-limit tool schemas;
- invalid tool names when the stateless profile exposes those tools;
- invalid, duplicate, sensitive, or unreachable
x-mcp-headerannotations; - zero or internally inconsistent resource limits;
sessionIndependenton a normal HTTP tool when treated as an MCP bridge control;- a
caller/exchangetarget withoutbackendResource; - the deprecated
caller-compatcredential mode on a new stateless target; - conflicting backend profiles for one normalized backend target;
- a stateless MCP target without explicit
backendCredentialMode; autowhen backend discovery is disabled by product policy.
discoverCacheScope: public and toolsListCacheScope: public are invalid
whenever authentication, delegation, or access-control policy can change the
corresponding result. An empty origin allowlist means browser-originated MCP
requests are rejected; it does not reject non-browser requests without an
Origin header.
These field names are the locked configuration contract. Stateless MCP tools
must declare backendCredentialMode. Portal publication defaults it to
workflow only for the com.networknt.workflow-1.0.0 service when service
discovery is used without a direct targetHost. Gateway enforces the same
target restriction before sending the caller access token in Authorization
and its shared service token in X-Scope-Token. Other stateless backends must
select a credential mode explicitly. The legacy profile retains its
compatibility default. The secure defaults are normative: legacy enabled,
stateless available by default, and stateless-to-legacy bridging rejected.
The runtime models the fixed-value cache-scope and bridge fields explicitly.
Only private and reject, respectively, are accepted in this release;
unknown stateless protocol fields fail configuration loading instead of being
silently ignored.
For an explicit targetHost that does not set
toolMetadata.runtime.allowPrivateTargetHost: true, SSRF protection is split
across two checks. Literal IP addresses are rejected during static URL
validation when they are loopback, private, link-local, or metadata addresses.
Hostnames are also checked by the HTTP client’s connection-time DNS resolver;
the resolver rejects the entire lookup when any returned address is non-public,
closing the validation-to-connect DNS rebinding window.
Targets resolved from the privileged service registry control plane are
already approved internal targets. This includes both
direct-registry.directUrls and portal-discovered nodes. The resolved target
carries that trust decision to the separate private-target client without
requiring duplicate per-tool metadata. The public client is never silently
downgraded for an explicit target. The public-target resolver cannot be combined
with an HTTP proxy because the proxy would resolve the origin outside the
gateway’s connection-time policy; such a client configuration fails closed.
Resource Limits
Legacy session capacity and stateless request capacity are separate budgets.
One must not consume or release permits from the other.
maxResponseBodyBytes must be at least 2048 so the gateway can always return a
bounded protocol error; error messages are truncated safely and backend error
bodies are never copied into client-visible errors.
The stateless profile requires bounded:
- request body bytes and JSON nesting depth;
- tools returned per list result, in addition to the buffered response-byte limit;
- schema document bytes, schema depth, subschema count, reference expansion, and validation time;
- concurrent requests globally and per principal;
- concurrent backend calls per target;
- response bytes for buffered responses;
- tools-list cache entries and TTL;
- discovery cache entries and TTL;
- extension metadata and MRTR state, even when the initial limit is zero because those features are unsupported;
- open subscriptions globally and per principal;
- events and bytes queued per subscription;
- subscription lifetime, additionally capped by current credential expiry on a protected route;
- initialization and teardown work for an enabled compatibility bridge.
Schema validation runs on a dedicated fixed-size worker pool, not Tokio’s core
workers or shared blocking pool. The pool is bounded by available parallelism
and maxConcurrentValidations; admission and queue capacity are bounded by the
same configured ceiling, and overload fails before enqueue. The duration
watchdog is observational because a running validator cannot be cancelled.
Structural bounds, linear-time regex, the adversarial corpus, and isolation
bound the gateway-wide impact without claiming a hard per-validation timeout.
Overload responses must be explicit and observable. A slow subscription is closed rather than allowed to grow an unbounded queue. Limits should use RAII permits so completion, timeout, cancellation, panic unwinding, and disconnect all release capacity.
Reload, Scaling, and Failure Semantics
Configuration Reload
Legacy frontend sessions and their mapped backend sessions continue to survive a compatible in-process reload. If a reload disables a version still used by an established legacy session, the rollout policy must decide whether to retain that version until session expiry or terminate affected sessions explicitly; it must not reinterpret the session under another version.
Stateless ordinary requests retain no protocol state across calls. New requests immediately use the new router and policy revisions.
An active subscription owns a live response stream, not a protocol session.
The subscription hub should be shared across compatible runtime swaps so a
successful catalog reload can publish toolsListChanged. An incompatible
reload closes affected streams; clients reconnect and re-subscribe.
Multiple Replicas
Ordinary stateless requests must work through round-robin routing without affinity or a shared protocol session store. Caches may remain per replica if their keys and invalidation rules are safe.
Legacy requests still require affinity to the gateway process that owns the frontend session, or a separately designed shared session/routing layer. Adding the stateless profile does not make legacy sessions horizontally portable.
Subscriptions remain attached to the replica holding their HTTP stream. They do not require subsequent ordinary requests to return to that replica.
Retry and Downgrade
The gateway and clients must distinguish compatibility failures from security and availability failures.
Permitted compatibility fallback:
server/discoverreturns method not found from a server believed to predate the stateless profile;- the server returns a well-formed unsupported-version error listing a mutually supported legacy version.
Terminal failures with no automatic downgrade:
- authentication or authorization rejection;
- missing or mismatched headers;
- TLS or certificate failure;
- DNS or SSRF rejection;
- malformed JSON-RPC or discovery response;
- timeout, connection failure, or HTTP 5xx;
- missing required client capability;
- policy or internal gateway failure.
Mutations with ambiguous outcomes are not replayed unless their tool metadata explicitly allows a safe retry. Read-only or declared-idempotent calls may use the existing bounded retry policy.
Observability
Record bounded labels for:
- frontend profile and protocol version;
- JSON-RPC method;
- configured tool name;
- backend type and backend profile;
- status and normalized error class;
- compatibility fallback decision;
- cache hit or miss;
- active request and subscription counts;
- subscription termination reason;
- legacy session and bridge capacity utilization.
Structured diagnostics may include correlation id, method, configured tool name, status, elapsed time, and byte counts. They must omit request arguments, result bodies, cookies, authorization headers, CSRF values, session ids, explicit state handles, and raw principal identifiers.
Valid W3C trace context received in the protocol-defined _meta fields may be
propagated through the gateway’s tracing model. Conflicting or malformed trace
context must not replace independently generated correlation identifiers.
2026-07-28 Release Coverage Matrix
This matrix is the traceability contract between the consolidated release,
this design, the responsible light-fabric boundary, and the delivery gate.
Required means the feature is needed for release qualification of the stateless profile.
Deferred means the gateway must omit the capability and reject or translate
the feature explicitly. Dependency means another light-fabric module owns the
behavior but must provide release evidence.
| Release area | Primary sources | Gateway decision and owner | Status and gate |
|---|---|---|---|
Remove protocol sessions and Mcp-Session-Id | SEP-2567 | Stateless frontend adapter never touches the session store; legacy adapter remains isolated | Required, Phases 1-2 |
| Remove initialize handshake; per-request version and capabilities | SEP-2575 | Classifier and versioned frontend adapters normalize into EffectiveMcpRequestContext | Required, Phases 1-2 |
server/discover | SEP-2575, discovery spec | Implement auth-aware cacheable discovery with truthful capabilities, server identity, TTL, and cache scope | Required, Phase 2 |
| Standard HTTP routing and parameter headers | SEP-2243 | Validate and regenerate Mcp-Method, applicable Mcp-Name, and valid Mcp-Param-* headers | Required, Phases 2-3 |
| Streamable HTTP request rules and Origin protection | Transport spec | Enforce POST body/content negotiation, exact Origin allowlist, profile-specific GET/DELETE rules, initial notification rejection, and 202 only for a future supported notification | Required, Phase 2 |
| Session-independent list results | SEP-2567 | Catalog may vary by principal/config/policy, never by connection or prior call | Required, Phase 2 |
| Cache TTL and scope | SEP-2549 | server/discover and tools/list return bounded ttlMs and normally private cacheScope | Required, Phase 2 |
| Deterministic tool ordering | Tools spec | Preserve BTreeMap ordering after authorization filtering and test byte-stable results | Required, Phase 2 |
| Tool name format and collision handling | SEP-986, tools spec | Validate stateless-exposed names; require explicit aliases for aggregate collisions | Required, Phases 1-2 |
| JSON Schema 2020-12 and schema dialects | SEP-1613, SEP-2106, base spec | Add bounded compile/validation; 2020-12 mandatory; external network references disabled | Required, Phases 1-2 |
| Tool input validation error semantics | SEP-1303, tools spec | Return schema failures as bounded isError: true tool results; reserve protocol errors for malformed envelopes and unknown tools | Required, Phases 1-2 |
Arbitrary structuredContent and output schemas | SEP-2106, tools spec | Support all JSON roots and validate final filtered output against configured outputSchema | Required, Phase 2 |
| Standalone request/result payload definitions | SEP-1319 | No wire-format change; pin generated schemas and keep protocol-neutral payload models separate from JSON-RPC adapters | Required architecture boundary, Phases 0-1 |
Required resultType | SEP-2322, base spec | Emit complete; accept absent only from earlier negotiated backends; reject unknown values | Required, Phases 2-3 |
| Multi Round-Trip Requests | SEP-2322, SEP-2260 | Do not advertise initially; reject input_required from a backend without down-conversion | Deferred, separate design |
| Elicitation and MRTR migration changes | SEP-1034, SEP-1036, SEP-1330, changelog | Not reachable while MRTR and elicitation are unadvertised; a later design must cover defaults, URL/enum schemas, and removal of notifications/elicitation/complete and elicitationId | Deferred with MRTR |
subscriptions/listen | SEP-2575 | Replace buffered streaming abstraction, then add bounded tools-list-change streams | Deferred to Phase 5 |
Remove SSE replay and Last-Event-ID | SEP-2575 | No replay buffer; clients retry with a new request id and rehydrate state | Required, Phase 2/5 |
Remove ping, logging/setLevel, roots-list-changed, old subscribe methods | SEP-2575 | Return method not found and keep capabilities absent | Required boundary, Phase 2 |
| Per-request log level and message notification rule | SEP-2575, changelog | Do not emit or forward notifications/message unless that request supplied io.modelcontextprotocol/logLevel; retain no logging state | Required boundary, Phase 2 |
Trace context in _meta | SEP-414 | Validate and bridge W3C trace context without overriding trusted gateway correlation state | Required, Phase 2 |
| Core extensions map and independent extension versions | SEP-2133 | Empty/absent initially; only validated adapters may advertise or forward an extension | Required boundary, Phase 2 |
| Tasks extension | SEP-2663 | Do not advertise; task support is forbidden/absent until a separate extension design | Deferred |
| MCP Apps extension | SEP-1865 | Do not advertise or forward UI metadata without a separate sandbox/consent design | Deferred |
Deprecate roots, sampling, logging, and sampling includeContext values | SEP-2577, SEP-2596 | Do not add new deprecated features or the deprecated thisServer/allServers values to this tool-only gateway profile | No new implementation |
| Feature lifecycle and HTTP+SSE deprecation | SEP-2596 | Keep legacy Streamable HTTP only; no implicit two-endpoint HTTP+SSE fallback | Required boundary, Phase 0 |
| MCP error-code allocation | Base spec/changelog | Preserve legacy implementation codes by version; reserve -32020 through -32022 for their assigned stateless meanings; lock implementation-defined -32000 to the catalog resource-limit error | Required, Phases 0-2 |
Resource-not-found error becomes -32602 | SEP-2164 | No direct tool-only behavior; any future resource adapter must be version-aware | Deferred resource design |
| Authorization issuer validation and mix-up defense | SEP-2468 | Security/client runtime validates recorded issuer and returned iss | Dependency, Phase 0 gate |
| Client registration type and issuer-bound credentials | SEP-837, SEP-2352 | Client runtime owns registration metadata and keys credentials by issuer | Dependency, Phase 0 gate |
| Protected-resource discovery, refresh tokens, and scope step-up | Authorization spec, SEP-2207, SEP-2350, SEP-2351 | Security/client runtime owns metadata, token, challenge, and bounded step-up behavior | Dependency, Phase 0 gate |
| Dynamic Client Registration deprecation | Changelog | Do not add DCR to the MCP router; prefer Client ID Metadata Documents in owning client code | Dependency/boundary |
| Authorization extensions | SEP-2133, authorization extensions | Disabled and unadvertised unless separately configured, negotiated, and tested | Deferred |
| Conformance scenarios required for standards | SEP-2484 | Pin official schema/scenarios and map every supported feature to CI evidence | Required, Phase 6 |
| Schema generator numeric correction | Changelog | Pin final generated schema; do not maintain hand-copied numeric field types | Required, Phase 0 |
| Governance and SEP process changes | SEP-1850 and governance entries | No runtime behavior; retain source links and final-spec refresh procedure | No runtime impact |
The matrix must be updated in the same change whenever the implementation advertises another core capability or extension. A capability without an owner, resource bounds, authorization behavior, and conformance evidence is invalid.
Delivery Sequence
Phase 0: Contract and Legacy Baseline
- Add the current stable
2025-11-25revision to the legacy compatibility suite before introducing2026-07-28. - Freeze legacy initialize, notification, list, call, SSE, DELETE, access-control, reload, and backend-session fixtures.
- Vendor or pin the RC TypeScript and generated JSON schemas plus conformance scenario revision, then replace them with the final July 28 artifacts.
- Record schema/checksum provenance so generated field types are not copied by hand.
- Lock the
-32000catalog resource-limit error and over-limit no-truncation fixtures before implementing statelesstools/list. - Add legacy principal-to-session binding.
- Close or assign every authorization dependency in the release coverage matrix, including backend token-audience strategy.
Phase 1: Protocol-Neutral Core
- Extract the classifier and
EffectiveMcpRequestContext. - Separate protocol validation/envelopes from common authorization and tool execution.
- Replace version constants with configured profile registries.
- Add configuration fields with backward-compatible defaults.
- Add bounded JSON Schema 2020-12 compilation and input/output validators.
- Validate stateless tool names and
x-mcp-headerannotations at load time.
Phase 2: Stateless Frontend Vertical Slice
- Implement
server/discover,tools/list, andtools/call. - Enforce Origin, content negotiation, per-request metadata, method rules, and SEP-2243 headers.
- Add stateless result, cache, and error envelopes.
- Add the empty extension boundary and reject unsupported result types/MRTR.
- Keep subscriptions and backend bridging disabled and unadvertised.
Phase 3: Stateless Backend Adapter
- Add explicit backend profiles.
- Implement legacy-to-stateless and stateless-to-stateless translation.
- Validate backend schemas, result types, extension capabilities, and output conformance according to the negotiated backend version.
- Add an audience-correct credential strategy per backend target.
- Add conservative backend discovery and downgrade tests.
Phase 4: Optional Compatibility Bridge
- Keep rejection as the default.
- If a real compatibility requirement exists, implement the bounded per-request bridge only for explicitly session-independent tools.
Phase 5: Streaming and Subscriptions
- Replace the buffered response abstraction.
- Implement cancellation-safe long-lived POST response streams.
- Add
toolsListChangedsubscriptions and truthful discovery capability.
Phase 6: Conformance and Canary
- Run the final official conformance suite where available.
- Verify every
RequiredandDependencycoverage-matrix row has linked test or operational evidence. - Exercise the full frontend/backend compatibility matrix.
- Test reload, multi-replica routing, downgrade resistance, limits, disconnect, and credential leakage.
- Enable stateless support for a canary client and target before changing the product default.
Verification Matrix
At minimum, automated tests must cover:
| Area | Required cases |
|---|---|
| Classification | Legacy initialize, legacy session request, stateless request, missing session, stale session header ignored by a fully identified stateless request, header/meta mismatch, unsupported version |
| HTTP transport | JSON content type, Accept requires JSON and SSE, exact Origin allowlist, empty allowlist with browser/non-browser clients, POST/GET/DELETE matrix, client-response rejection, unsupported notification rejection, future accepted-extension notification 202, unknown-method HTTP 404 plus JSON-RPC -32601 |
| Legacy regression | Existing JSON and SSE responses, DELETE, expiry, reload preservation, backend session reuse and teardown |
| Stateless discovery | Deterministic versions/capabilities, no false capabilities, server identity in _meta, TTL/scope, principal-aware cache, unsupported version details |
| Stateless list | Per-principal visibility, deterministic order, bounded complete catalog without nextCursor, rejection of unissued cursors, whole-request -32000 failure without truncation/caching when item or response limits are exceeded, private cache, TTL, policy/config invalidation |
| Tool schemas | Default and explicit dialects, invalid schema, local and external $ref, composition/depth/time bounds, arbitrary output roots, post-filter output validation |
| Tool error semantics | Malformed envelope and unknown tool produce protocol errors; input-schema failures produce bounded resultType: complete, isError: true results before backend traffic |
| Tool headers and names | Valid/invalid names, collision handling, all x-mcp-header constraints, encoding, missing/null values, mismatch, sensitive/header conflict |
| Stateless call | HTTP backend, stateless MCP backend, authorization denial, masking, filtering, output conformance, safe retries, no session-store mutation |
| Results and MRTR | Required complete, missing field by backend version, unknown extension result, unsupported input_required mapped to the exact gateway-generated tool error without opaque state, new request id on a future retry |
| Request-scoped notifications | No notifications/message without per-request log level; no retained log-level state; initiating response stream used instead of subscription stream |
| Extensions | Empty capability map, unknown optional extension, required unsupported extension, no backend extension smuggling, Tasks and Apps absent |
| Authorization | Protected-resource metadata, token on every request, 401/403 challenges, issuer/audience binding, frontend token not forwarded to wrong backend, bounded scope step-up |
| Backend compatibility | All six frontend/backend matrix cells with explicit success or incompatibility outcome |
| Downgrade resistance | No fallback after 401, 403, TLS, timeout, malformed response, mismatch, SSRF rejection, or 5xx |
| Resource safety | Body/depth/concurrency/cache/subscription bounds, cancellation, permit release, slow consumer |
| Streaming | First acknowledgment, subscription id tagging, disconnect cleanup, reload, no replay or resumption |
| Scaling | Stateless calls across alternating replicas; legacy behavior requires documented affinity |
| Secrets | No token, cookie, CSRF, session id, handle, key, or private payload in logs and gate output |
The release gate must prove that enabling the stateless adapter does not alter legacy fixtures when the same legacy configuration is loaded.
Risks and Mitigations
| Risk | Mitigation |
|---|---|
| Missing session id is misclassified as stateless | Require the complete 2026-07-28 version and metadata contract before selecting stateless |
| Legacy session id crosses principals | Bind the session to a verified principal fingerprint and revalidate current authentication on every request |
| Wire behavior drifts between profiles | Share application logic and isolate only versioned adapters and envelopes |
| Gateway silently downgrades after a security failure | Permit fallback only for explicit method/version compatibility responses |
| Stateless client is proxied through shared hidden legacy state | Reject by default; allow only bounded per-request bridging for declared session-independent tools |
| Tool catalog leaks across users | Use private cache scope and auth/policy/config-aware cache keys |
| Capability advertisement exceeds implementation | Build capabilities from enabled handlers and tested features, not schema availability |
| Buffered SSE is mistaken for subscription support | Replace the response type and add disconnect/cancellation tests before advertising subscriptions |
| Config reload leaves stale list results | Include revisions in cache keys, invalidate on reload, and publish list change only after a successful swap |
| A future discovery-dependent catalog is stale until TTL | Do not enable it until the registry provides a monotonic revision plus push/watch invalidation; treat TTL only as a documented backstop |
| Complex schemas exhaust CPU or trigger SSRF | Compile with depth/subschema/time limits and disable external network $ref resolution |
| Filtering produces output that violates the advertised schema | Validate final structuredContent after filtering and return a tool error on mismatch |
| Unknown extensions or result types pass through the facade | Advertise and forward only explicitly adapted, versioned extensions; reject unknown result types |
| Frontend bearer token is replayed to another resource | Require an audience-correct backend credential strategy and never transit arbitrary tokens |
| Raw frontend or unrecognized headers reach a stateless backend | Use a typed outbound allowlist and regenerate admitted context from trusted gateway state |
| Browser Origin is treated as ordinary CORS metadata | Enforce exact request Origin validation before JSON parsing; empty allowlist rejects browser origins |
| Release-candidate schema changes | Pin the final schema and error registry and run conformance tests before release |
Acceptance Criteria
This design is implemented when:
- One
/mcpendpoint concurrently accepts enabled legacy and2026-07-28clients without heuristic profile selection. - Existing legacy contract tests remain unchanged and pass.
- A stateless list or call touches no frontend session state and can be routed to alternating gateway replicas.
- Both profiles use the same access-control, masking, execution, filtering, audit, and metrics core.
- Backend protocol compatibility follows the explicit matrix and never silently pools hidden state.
- Discovery and capabilities describe only implemented behavior.
- Required stateless headers, metadata, result fields, cache fields, and error
codes conform to the final
2026-07-28specification. - Tool names,
inputSchema, optionaloutputSchema, arbitrary structured content, andx-mcp-headerbehavior conform to the final tools and schema contracts under bounded validation. - Origin validation, frontend OAuth challenges, issuer/audience binding, and backend credential selection pass their release gates.
- Unknown or deferred core features and extensions remain unadvertised and cannot pass transparently through the gateway.
- Every
RequiredandDependencyrelease-coverage row links to test or operational evidence. - All request, response, schema, concurrency, cache, and optional streaming resources are bounded and cancellation-safe.
- Stateless support is built in; release qualification remains a separate process.
Resolved Conditional Profile Decision
The Phase 9 dependency assessment found no configured backend that requires
auto discovery and no named legacy backend with an approved
session-independence proof. The release therefore keeps backend profiles
explicit, keeps stateless-to-legacy behavior at reject, and adds no dormant
discovery cache or hidden backend-session machinery. This decision may be
reopened only for a named dependency with the compatibility fixtures, security
review, administrator opt-in, lifecycle bounds, and teardown evidence required
by the applicable design section.
The browser controller route /ctrl/mcp does not qualify: it remains a
separate JSON/WebSocket control plane routed by websocket-router, not an
mcp-router Streamable HTTP backend.
Open Decisions Before Implementation Lock
- Confirm the final July 28 schema’s exact
clientInfoandserverInforequirements; the SEP text and consolidated RC use different normative strength. - Decide whether MCP Origin policy is stored directly in
mcp-router.ymlor supplied by a shared exact-origin security module. One component must be the authoritative validator; duplicated allowlists are not acceptable. - Select additional JSON Schema dialects, if any. Supporting only mandatory 2020-12 is the safest initial profile.
- Select each backend target’s audience-correct credential strategy and define which component performs token exchange or service-token acquisition.
- Select final default request, cache, and subscription limits from baseline measurements rather than treating the illustrative values as tuned limits.
- Decide whether subscription state should survive an in-process MCP router configuration swap or close and force re-subscription. The implementation must make either outcome deterministic and tested.
WebSocket Router
Status
Phases 1, 2, and 3 are implemented. Phase 1 added configuration parsing,
Java-compatible pathPrefixService normalization, route resolution, and
upstream URI cleanup in light-pingora. Phase 2 wired the websocket handler
into light-gateway with WebSocket upgrade detection, discovery-based upstream
selection, request context storage, and upstream header/query cleanup. Phase 3
added a real gateway-to-backend WebSocket integration test for text, binary,
close, subprotocol, and header behavior.
Purpose
The Java light-websocket-4j websocket-router module routes WebSocket
traffic through a gateway or sidecar. A client connects to the gateway, the
router resolves the downstream service from headers, query parameters, or path
prefix configuration, and the gateway connects to the target WebSocket service.
In light-fabric this should be a light-pingora traffic handler activated by
light-gateway through handler.yml. The same light-gateway binary can link
the WebSocket router implementation, while each product decides whether it runs
by including the websocket handler and websocket-router.yml configuration
from config-server.
The Rust implementation should preserve the Java routing semantics and most of
the Java configuration shape, but it should not copy Java’s enabled flag or
frame-bridging architecture. Pingora already supports HTTP/1 upgrade proxying,
so the first implementation should resolve the target and let Pingora tunnel
the upgraded connection.
Goals
- Add a Java-compatible WebSocket router to
frameworks/light-pingora. - Activate the router with the existing
websockethandler id inapps/light-gateway. - Keep the Java
websocket-routerrouting configuration recognizable:defaultProtocol,defaultEnvTag, andpathPrefixService. - Allow
websocket-router.pathPrefixServiceto be injected by config-server at startup the same way other handler-specific config is injected. - Resolve downstream services from header, query parameter, or longest path prefix.
- Reuse the existing light-gateway discovery and upstream selection model.
- Preserve WebSocket handshake headers and pass normal agent/browser headers through to the downstream service.
- Register the router configuration with the module registry and support the same reload model as other light-pingora handler configs.
- Keep the design suitable for gateway, sidecar, and BFF deployments.
Non-Goals
- Do not implement a separate WebSocket server framework in light-fabric.
- Do not terminate and re-create WebSocket frames in the first phase.
- Do not multiplex multiple client WebSocket sessions over one downstream connection.
- Do not support HTTP/2 extended CONNECT for WebSocket in the first phase.
- Do not use Rust dynamic plugins or
inventoryfor WebSocket route registration. - Do not create a separate gateway binary for WebSocket routing.
- Do not use
enabledinwebsocket-router.yml. The handler is active whenhandler.ymlincludeswebsocketin the matched execution chain.
Resolved Decisions
- Activation is controlled only by
handler.yml. If a matched chain includeswebsocket, the router is enabled for that request. websocket-router.ymlshould not containenabled.- WebSocket-specific controls should cover both request/upgrade rate and active upgraded connection count.
- The first implementation should use Pingora HTTP/1 upgrade passthrough, not a frame-aware WebSocket bridge.
- Invalid
websocket-router.ymlconfiguration should fail startup. Invalid reloads should be rejected while the last valid runtime state keeps serving existing traffic.
Java Behavior To Map
Java configuration includes enabled, but the Rust target config removes it:
# Light websocket router configuration
defaultProtocol: ${websocket-router.defaultProtocol:http}
defaultEnvTag: ${websocket-router.defaultEnvTag:}
pathPrefixService: ${websocket-router.pathPrefixService:}
preserveRoutingHeaders: ${websocket-router.preserveRoutingHeaders:false}
idleTimeoutMs: ${websocket-router.idleTimeoutMs:3600000}
maxConnectionDurationMs: ${websocket-router.maxConnectionDurationMs:}
maxActiveConnections: ${websocket-router.maxActiveConnections:}
maxUpgradeRequestsPerSecond: ${websocket-router.maxUpgradeRequestsPerSecond:}
The Java enabled field is intentionally not carried forward. In Rust, the
handler chain is the activation contract. Removing websocket from a path or
default chain disables WebSocket routing for that path.
Production controls are optional. idleTimeoutMs defaults to one hour; blank
or zero values disable the matching control. preserveRoutingHeaders defaults
to false, so routing-only
Service-Id, service_id, and serviceId headers are stripped before the
upstream handshake unless a backend explicitly needs them.
pathPrefixService accepts three forms:
pathPrefixService:
/chat:
serviceId: com.networknt.llmchat-1.0.0
protocol: http
envTag: dev
pathPrefixService:
/chat: com.networknt.llmchat-1.0.0
pathPrefixService: {"/chat":{"serviceId":"com.networknt.llmchat-1.0.0","protocol":"http","envTag":"dev"}}
The Java handler resolves the downstream service in this order:
- Header: first non-blank value from
Service-Id,service_id, orserviceId. - Query parameter: first non-blank value from
service_idorserviceId. - Path prefix:
pathPrefixServicematch against the request path.
If a target is found, query parameters can override the target protocol and environment tag:
protocolenv_tagenvTag
The Java handler removes router-only query parameters before connecting to the downstream service:
protocolservice_idserviceIdenv_tagenvTag
The Java implementation accepts client WebSocket subprotocols, opens a new JDK
WebSocket client connection to the downstream service, forwards Authorization,
forwards the selected subprotocols, and then bridges text and binary frames in
both directions.
Rust Architecture
Add the WebSocket router to light-pingora because it is a Pingora gateway
traffic handler.
Proposed module:
frameworks/light-pingora/src/websocket.rs
Primary types:
#![allow(unused)]
fn main() {
pub struct WebSocketRouterConfig {
pub default_protocol: String,
pub default_env_tag: Option<String>,
pub path_prefix_service: BTreeMap<String, WebSocketServiceTarget>,
}
pub struct WebSocketServiceTarget {
pub service_id: String,
pub protocol: String,
pub env_tag: Option<String>,
}
pub struct WebSocketRouteDecision {
pub service_id: String,
pub protocol: String,
pub env_tag: Option<String>,
pub upstream_path_and_query: String,
}
}
The serde layer should accept Java field names through aliases:
defaultProtocoldefaultEnvTagpathPrefixServiceserviceIdenvTag
Use websocket-router.yml as the preferred Rust file name. Accept
websocket-router.yaml as a compatibility fallback.
Config Normalization
Normalize pathPrefixService at load time:
raw config
-> validate defaultProtocol/defaultEnvTag
-> parse pathPrefixService YAML map, JSON string map, or legacy key/value string
-> apply defaults to entries missing protocol or envTag
-> sort prefixes by length for longest-prefix matching
-> build Arc<WebSocketRouterState>
An invalid entry should fail config loading instead of being ignored silently. This is stricter than Java and is safer for remote config delivered by config-server.
Handler Registration
apps/light-gateway already reserves the websocket handler id as a traffic
handler. The implementation should attach that id to the WebSocket router
runtime:
handlers:
- correlation
- metrics
- jwt
- limit
- websocket
paths:
- path: /chat
method: GET
exec:
- correlation
- metrics
- jwt
- limit
- websocket
The router should only run for chains that include websocket. This lets a BFF
serve static SPA assets, REST APIs, MCP, JSON-RPC, and WebSocket endpoints from
the same gateway binary with path-specific handler chains.
Request Flow
The target flow should be:
client request
-> handler.yml path/chain match
-> cross-cutting request handlers
-> websocket handler
-> verify WebSocket upgrade
-> resolve service target
-> strip router-only query parameters
-> store WebSocketRouteDecision in request context
-> Pingora upstream_peer selects discovered target
-> Pingora upstream_request_filter preserves WebSocket handshake headers
-> Pingora proxies the HTTP/1 upgraded stream
-> response/metrics handlers observe completion
The router should not read the request body and should not buffer WebSocket messages. Once the request is upgraded, Pingora owns the tunnel.
Upgrade Detection
The handler should require the normal WebSocket handshake:
- method
GET ConnectioncontainsupgradeUpgradeequalswebsocketSec-WebSocket-Keyexists- HTTP version is compatible with HTTP/1 upgrade
If the websocket handler is selected by handler.yml but the request is not
a WebSocket upgrade, return 426 Upgrade Required.
HTTP/2 extended CONNECT can be considered later, but should not block the first implementation.
Target Resolution
Target resolution should match Java precedence:
1. service id header
2. service id query parameter
3. pathPrefixService longest-prefix match
Header names:
Service-Id
service_id
serviceId
Query names:
service_id
serviceId
protocol
env_tag
envTag
For path-prefix matches, use the request path without the query string. When multiple prefixes match, choose the longest prefix.
The resolved protocol should be http or https. Conceptually this maps to
ws or wss, but Pingora should still connect to the upstream as HTTP or
HTTPS and then perform the WebSocket upgrade.
Header And Query Policy
Because the Rust implementation should use Pingora upgrade passthrough, it should preserve the original handshake headers:
UpgradeConnectionSec-WebSocket-KeySec-WebSocket-VersionSec-WebSocket-ProtocolSec-WebSocket-ExtensionsAuthorization- cookies
- normal agent/browser headers
The router should strip only router-control query parameters from the upstream URI:
protocolservice_idserviceIdenv_tagenvTag
The service-id routing headers should be removed before the upstream request by default:
Service-Idservice_idserviceId
This keeps gateway routing controls separate from backend application headers. If a backend later needs these headers, add an explicit config option rather than leaking them by default.
Discovery And Upstream Selection
The WebSocket router should reuse the same discovery/runtime model as
router.yml and the existing Pingora proxy flow.
Resolved target:
protocol + serviceId + envTag
Discovery returns an upstream HTTP or HTTPS endpoint. upstream_peer creates
the Pingora peer:
http: non-TLS upstreamhttps: TLS upstream with normal SNI/hostname handling
For the first implementation, require HTTP/1.1 to the backend for WebSocket upgrade. HTTP/2 WebSocket tunneling can be a later feature.
Error Handling
Errors should be returned before the connection is upgraded:
| Condition | Response |
|---|---|
| Handler selected but request is not WebSocket upgrade | 426 Upgrade Required |
| No service id and no path-prefix match | 403 Forbidden |
| Invalid protocol override | 400 Bad Request |
| Discovery has no usable endpoint | 502 Bad Gateway |
| Upstream connect/upgrade failure | 502 Bad Gateway |
Returning HTTP errors before upgrade is clearer than Java’s close-frame behavior because the Rust implementation does not accept the WebSocket until the target is known.
Module Registry And Reload
Register the loaded configuration with the module registry:
module id: light-pingora/websocket-router
config name: websocket-router
config file: websocket-router.yml or websocket-router.yaml
On reload:
- Load and validate the new config.
- Build a new immutable route state.
- Atomically swap the state.
- Let in-flight upgraded connections continue with the old decision.
Existing WebSocket tunnels should not be interrupted by a config reload unless the gateway process is restarted.
Observability
The handler should integrate with existing correlation and metrics handlers:
- include correlation id in pre-upgrade logs
- record target resolution result
- record route source:
header,query, orpathPrefixService - count upgrade attempts, successful upgrades, rejected upgrades, and upstream connection failures
- optionally record tunnel duration once Pingora exposes completion
Do not log full query strings by default because they may contain application data.
Test Plan
Parser and resolver tests:
- YAML object
pathPrefixService - string service id entries
- JSON string map entries
- legacy key/value string entries
- default protocol and env tag application
- invalid entries fail load
- header beats query and path prefix
- query beats path prefix
- longest prefix wins
- query protocol/envTag override
- router query params are stripped
Gateway tests:
- non-upgrade request to a WebSocket chain returns
426 - missing target returns
403 - unknown discovery target returns
502 - upgrade request preserves
Sec-WebSocket-Protocol Authorizationand normal browser/agent headers pass through- service-id routing headers are stripped before upstream
Integration tests:
- connect through light-gateway to a local WebSocket echo backend
- text message round trip
- binary message round trip
- close frame behavior
- subprotocol negotiation
- TLS upstream smoke test when a local test certificate is available
Implementation Phases
Phase 1: Config And Resolver
Status: implemented.
- Add
frameworks/light-pingora/src/websocket.rs. - Parse
websocket-router.ymlandwebsocket-router.yaml. - Normalize all Java-compatible
pathPrefixServiceforms. - Implement target resolution and upstream URI cleanup.
- Add unit tests.
Phase 2: Gateway Handler Wiring
Status: implemented.
- Connect the existing
websockethandler id to the router runtime. - Detect WebSocket upgrade requests in the Pingora request flow.
- Store
WebSocketRouteDecisionin the request context. - Select the discovered upstream in
upstream_peer. - Strip router query params and service-id headers in
upstream_request_filter.
Phase 3: WebSocket Integration Tests
Status: implemented.
- Add a local test WebSocket echo service.
- Verify text, binary, close, subprotocol, and header behavior through light-gateway.
- Verify HTTP and HTTPS upstream paths if practical in CI.
Phase 4: Production Controls
Status: implemented.
- Add optional idle timeout and max connection duration.
- Add WebSocket-specific limit controls for both upgrade/request rate and active upgraded connection count.
- Add explicit config for preserving routing headers if a backend requires them.
- Add access-control integration once the same access-control model is shared across REST, JSON-RPC, MCP, and WebSocket routes.
Implementation notes:
maxUpgradeRequestsPerSecondgates accepted upgrade attempts before discovery lookup.maxActiveConnectionstracks proxied upgraded sessions with a permit that is released when Pingora finishes the request context. The active counter is preserved across router and policy reloads.idleTimeoutMsis applied to downstream and upstream tunnel IO. Pingora’s body-filter hooks also check idle age when either side sends tunneled data.maxConnectionDurationMsis checked by the tunnel body filters and is also used as an IO timeout when it is the only timeout configured. A connection that continuously exchanges frames is closed on the next tunneled body chunk after the duration is exceeded.- WebSocket access-control uses the shared
access-control.ymlandrule.ymlmodel. The rule context uses tool namewebsocket, endpoint fromhandler.yml, and tool arguments containingserviceId,protocol,envTag,upstreamPathAndQuery, and routesource.
Open Questions
None.
Stateless Auth Handler
Status
Initial Rust implementation is complete in light-pingora and
light-gateway. It includes the shared SPA session runtime, authorization-code
entrypoint, logout, cookie handling, CSRF validation, refresh-token renewal,
Google/Facebook/GitHub callback entrypoints, handler wiring, config stubs, and
runtime-load tests.
Purpose
The Java light-spa-4j stateless-auth module is the BFF login bridge for
SPA deployments that use OAuth 2.0 authorization code flow in the cloud. The
browser completes the provider redirect, calls the gateway callback path with
the authorization code, and the gateway exchanges that code for light-oauth
tokens. The gateway then stores the internal access token, refresh token, user
metadata, and CSRF value in browser cookies.
In light-fabric this should be a light-pingora security handler used by
light-gateway. The handler should be activated by handler.yml, loaded from
config-server with the same product-level configuration model as the rest of
the gateway, and implemented with the same shared SPA session runtime used by
the MSAL exchange handler.
Goals
- Preserve the Java BFF behavior for authorization code login, logout, CSRF
validation, refresh-token renewal, and downstream
Authorizationinjection. - Keep the Java
statelessAuth.ymlfield names recognizable so light-portal can injectstatelessAuth.*values into config-server output. - Use
handler.ymlas the primary activation and ordering contract. - Keep the existing
statelesshandler id as the public handler-chain name. - Share cookie, CSRF, JWT parsing, refresh-token single-flight, and
Authorizationinjection code with the MSAL exchange handler. - Use the existing
client.ymlOAuth token configuration for authorization code and refresh-token calls. - Register the loaded config in
ModuleRegistryand reject invalid config at startup. - Support BFF chains that also use static SPA serving, proxy/router, WebSocket routing, and MCP routing.
- Support Google, Facebook, and GitHub login entrypoints in addition to the generic authorization-code callback.
Non-Goals
- Do not use Rust dynamic plugins or
inventory. - Do not create a separate BFF binary.
- Do not store server-side browser sessions in the first implementation.
- Do not require the Rust social-login implementation to copy Java’s provider-specific classes. Rust should preserve the external behavior and config contract, but it can use established OAuth/OIDC crates for provider protocol handling.
- Do not redirect the browser from the gateway by default. Java returns a JSON
body containing
redirectUri,denyUri, andscopes; Rust should preserve that behavior.
Resolved Decisions
- Google, Facebook, and GitHub login handlers are in scope. The existing
google,facebook, andgithubhandler ids should remain as public handler-chain names. - Rust should prefer provider-appropriate crates instead of hand-rolling every
provider flow.
openidconnectis a good fit for OpenID Connect providers such as Google, andoauth2is a good fit for plain OAuth 2.0 providers or provider-specific extensions. cookieTimeoutUrishould be used by Rust to return a structured session-expired response when a browser session cannot be renewed.
Java Behavior To Map
Java config file:
enabled: ${statelessAuth.enabled:true}
redirectUri: ${statelessAuth.redirectUri:https://localhost:3000/#/app/dashboard}
denyUri: ${statelessAuth.denyUri:https://localhost:3000/#/app/dashboard}
enableHttp2: ${statelessAuth.enableHttp2:false}
authPath: ${statelessAuth.authPath:/authorization}
logoutPath: ${statelessAuth.logoutPath:/logout}
logoutCsrfEnforced: ${statelessAuth.logoutCsrfEnforced:false}
cookieDomain: ${statelessAuth.cookieDomain:localhost}
cookiePath: ${statelessAuth.cookiePath:/}
cookieTimeoutUri: ${statelessAuth.cookieTimeoutUri:/}
cookieSecure: ${statelessAuth.cookieSecure:true}
sessionTimeout: ${statelessAuth.sessionTimeout:3600}
rememberMeTimeout: ${statelessAuth.rememberMeTimeout:604800}
bootstrapToken: ${statelessAuth.bootstrapToken:token}
googlePath: ${statelessAuth.googlePath:/google}
googleClientId: ${statelessAuth.googleClientId:google_client_id}
googleClientSecret: ${statelessAuth.googleClientSecret:secret}
googleRedirectUri: ${statelessAuth.googleRedirectUri:https://localhost:3000}
facebookPath: ${statelessAuth.facebookPath:/facebook}
facebookClientId: ${statelessAuth.facebookClientId:facebook_client_id}
facebookClientSecret: ${statelessAuth.facebookClientSecret:secret}
githubPath: ${statelessAuth.githubPath:/github}
githubClientId: ${statelessAuth.githubClientId:github_client_id}
githubClientSecret: ${statelessAuth.githubClientSecret:secret}
Java request behavior:
GET authPath, normally/authorization, expects query parametercodeand optionalstate.- Missing
codereturnsERR10035. - The handler generates a CSRF value and sends an authorization-code token
request through
http-clientusingclient.ymloauth.token.authorization_code. - On success, it sets browser cookies and returns JSON containing
scopes,redirectUri, anddenyUri. POST logoutPath, normally/logout, validates the readable CSRF cookie/header pair when enforcement is enabled, clears BFF cookies, and returns204 No Content.- Other requests are treated as downstream BFF requests. The handler reads the
accessTokencookie, verifies/parses it, validates CSRF, refreshes the token if it expires within 90 seconds, and injectsAuthorization: Bearer <access-token>before the proxy/router handler runs. - If no access token exists but a refresh token exists, the handler attempts refresh and then injects the new access token.
- If neither cookie exists, Java allows the request to continue. The downstream service can still decide whether the endpoint is anonymous or protected.
Java error codes to preserve:
| Code | Meaning |
|---|---|
ERR10035 | Authorization code is missing |
ERR10000 | Access token is invalid |
ERR10036 | CSRF token is missing from request |
ERR10038 | CSRF claim is missing from JWT |
ERR10039 | Request CSRF and JWT CSRF do not match |
ERR10037 | Refresh-token response is empty |
ERR10008 | Method is not allowed for a mutation or callback endpoint |
ERR11649 | Logout CSRF cookie/header validation failed without exposing either value |
Rust Architecture
Add a shared SPA auth runtime in light-pingora and expose it through
light-gateway.
Proposed modules:
frameworks/light-pingora/src/spa_auth.rs
frameworks/light-pingora/src/stateless_auth.rs
spa_auth.rs owns the reusable mechanics:
#![allow(unused)]
fn main() {
pub struct SpaCookieConfig {
pub cookie_domain: String,
pub cookie_path: String,
pub cookie_secure: bool,
pub session_timeout: u64,
pub remember_me_timeout: u64,
pub same_site: CookieSameSite,
pub renew_before_seconds: u64,
}
pub struct SpaSessionRuntime {
pub cookies: SpaCookieConfig,
pub token_client: Arc<SpaTokenClient>,
pub jwt_verifier: Arc<SecurityRuntime>,
pub refresh_single_flight: RefreshSingleFlight,
}
pub struct SpaSessionResult {
pub access_token: Option<String>,
pub principal: Option<AuthPrincipal>,
pub response_cookies: Vec<SetCookie>,
}
}
stateless_auth.rs owns the authorization-code entrypoint:
#![allow(unused)]
fn main() {
pub struct StatelessAuthConfig {
pub enabled: bool,
pub redirect_uri: String,
pub deny_uri: Option<String>,
pub enable_http2: bool,
pub auth_path: String,
pub logout_path: String,
pub cookie_domain: String,
pub cookie_path: String,
pub cookie_timeout_uri: String,
pub cookie_secure: bool,
pub session_timeout: u64,
pub remember_me_timeout: u64,
pub bootstrap_token: Option<String>,
pub renew_before_seconds: u64,
pub google: Option<SocialProviderConfig>,
pub facebook: Option<SocialProviderConfig>,
pub github: Option<SocialProviderConfig>,
}
pub struct SocialProviderConfig {
pub path: String,
pub client_id: String,
pub client_secret: String,
pub redirect_uri: Option<String>,
pub scopes: Vec<String>,
}
pub struct StatelessAuthRuntime {
pub config: StatelessAuthConfig,
pub session: SpaSessionRuntime,
}
}
Use Java-compatible serde aliases for camel-case config fields. The primary
file should be statelessAuth.yml; accept statelessAuth.yaml as a
compatibility fallback.
The serde layer can keep the Java-compatible flat fields, such as
googlePath, googleClientId, and googleClientSecret, and normalize them
into SocialProviderConfig entries after load. This keeps config-server
compatibility while giving Rust a cleaner internal model.
Handler Registration
apps/light-gateway already reserves the stateless handler id. The runtime
loader should follow the same pattern as MCP:
#![allow(unused)]
fn main() {
let stateless_auth = load_stateless_auth_runtime(
config,
active_handlers.is_handler_active("stateless"),
)?;
}
If stateless is not active in any chain, the config does not need to be
loaded. If the config is active but enabled: false, register the disabled
module and return None.
No @alias syntax is needed. The handler id in handler.yml is the stable
Rust contract.
Example BFF chain:
handlers:
- exception
- cors
- stateless
- header
- prefix
- token
- router
chains:
default:
- exception
- cors
- stateless
- header
- prefix
- token
- router
websocket:
- exception
- stateless
- security
- websocket
paths:
- path: /authorization
method: GET
exec:
- default
- path: /google
method: GET
exec:
- google
- path: /facebook
method: GET
exec:
- facebook
- path: /github
method: GET
exec:
- github
- path: /logout
method: POST
exec:
- default
- path: /logout
method: OPTIONS
exec:
- default
The handler should normally run after CORS and before proxy/router/WebSocket.
POST /logout is the only allowed logout method. The authorization and enabled /google,
/facebook, and /github authorization-code callbacks remain permanently
GET-only. Keep the explicit OPTIONS /logout route permanently.
Login Flow
For authPath:
GET /authorization?code=...&state=...
-> validate code
-> generate csrf
-> call token endpoint with authorization_code grant
-> verify/parse returned internal access token
-> set BFF cookies
-> return { "scopes": [...], "redirectUri": "...?state=...", "denyUri": "..." }
Token request mapping should reuse client.yml:
oauth.token.server_urloroauth.token.serviceIdoauth.token.enableHttp2oauth.token.authorization_code.urioauth.token.authorization_code.client_idoauth.token.authorization_code.client_secretoauth.token.authorization_code.redirect_urioauth.token.authorization_code.scope
The form body should match Java:
grant_type=authorization_code
code=<code>
redirect_uri=<optional redirect_uri>
csrf=<generated csrf>
scope=<space separated scopes, if configured>
Logout Flow
POST /logout
Cookie: accessToken=...; csrf=...
X-CSRF-TOKEN: <csrf>
-> optionally enforce logout double-submit CSRF
-> emit deletion cookies for every cookie the runtime can set
-> return 204 No Content with no body or response content type
The logout request has no required body. A zero-length body is valid even when
a shared client declares Content-Type: application/json. A legacy GET or any
other unsupported logout method returns 405, ERR10008, and Allow: POST;
a wrong callback method returns 405, ERR10008, and Allow: GET. OPTIONS
continues to reach CORS.
Session Validation Flow
For requests that are not login/logout:
request
-> read accessToken cookie
-> verify/parse internal JWT with security.yml rules
-> extract csrf claim
-> find request CSRF from X-CSRF-TOKEN, WebSocket subprotocol, or query
-> compare csrf values
-> refresh token if exp is inside renew window
-> inject Authorization: Bearer <access-token>
-> continue handler chain
CSRF source order should match Java:
X-CSRF-TOKENheader.Sec-WebSocket-Protocolvalue starting withcsrf.when the request hasSec-WebSocket-KeyandSec-WebSocket-Version.- Query parameter
csrf.
The WebSocket subprotocol behavior is important for browser WebSocket clients
that cannot set arbitrary headers. The auth handler should run before the
websocket router so the downstream handshake receives the internal
Authorization header.
Session-Expired Response
The Java handler usually allows requests with no cookies to continue so the downstream service can decide whether the endpoint is anonymous. Rust should preserve that pass-through behavior for requests with no session evidence.
When the request does have session evidence but the session cannot be renewed,
for example an expired or rejected refresh token, Rust should clear BFF cookies
and return a structured response using cookieTimeoutUri:
{
"code": "ERR10040",
"message": "SPA session expired",
"timeoutUri": "/",
"authenticated": false
}
The status should be 401 unless a later product config explicitly asks for a
different behavior. This gives the SPA a deterministic signal to navigate to
the configured timeout or login page without scraping an Undertow-style status
string.
Internal JWT Verification
The shared SPA runtime should not call the existing verify_jwt_request
function directly. That function is designed for API requests with an
Authorization header, path skips, pass-through claims, and normal security
handler behavior.
The SPA auth runtime needs a lower-level token verifier that can:
- verify the access-token signature using the same certificates and algorithms
as
security.yml; - parse claims from a token stored in a cookie;
- optionally ignore expiration while deciding whether the token can be refreshed;
- fail hard on invalid signature, invalid algorithm, malformed JWT, and missing key;
- return an
AuthPrincipaland raw claims for CSRF, cookie metadata, and optional request-context propagation.
This can be implemented by extracting a reusable helper from security.rs,
for example:
#![allow(unused)]
fn main() {
verify_jwt_token(
runtime: &SecurityRuntime,
token: &str,
expiry_mode: JwtExpiryMode,
) -> Result<AuthPrincipal, HandlerRejection>
}
The normal security handler can keep its current request-level wrapper, while
SPA auth uses the token-level helper for cookie tokens.
Social Provider Login
Google, Facebook, and GitHub login are implemented as thin handler entrypoints that reuse the same cookie/session runtime as the authorization-code callback. The existing handler ids are kept:
chains:
google:
- exception
- correlation
- cors
- google
- stateless
- header
- prefix
- router
facebook:
- exception
- correlation
- cors
- facebook
- stateless
- header
- prefix
- router
github:
- exception
- correlation
- cors
- github
- stateless
- header
- prefix
- router
The implemented provider flow is:
- Match its configured provider path, for example
googlePath,facebookPath, orgithubPath. - For Google, exchange the authorization
codewith the Google token endpoint and use the returnedid_tokenas the subject token. If the provider does not return an ID token, fall back toaccess_token. - For Facebook, accept the Java-compatible
accessTokenquery parameter, or exchange an authorizationcodewith the Facebook token endpoint. - For GitHub, exchange the authorization
codewith the GitHub token endpoint. - Use
client.ymloauth.token.token_exchangeto exchange the provider subject token for an internal light-oauth token set with a CSRF claim. - Set the same BFF cookies as the generic stateless handler and return the same JSON shape.
Provider token endpoints default to the public provider URLs, but can be overridden for tests or regional deployments:
googleTokenEndpoint: ${statelessAuth.googleTokenEndpoint:https://oauth2.googleapis.com/token}
facebookTokenEndpoint: ${statelessAuth.facebookTokenEndpoint:https://graph.facebook.com/v19.0/oauth/access_token}
githubTokenEndpoint: ${statelessAuth.githubTokenEndpoint:https://github.com/login/oauth/access_token}
External identity mapping is intentionally delegated to the internal token-exchange implementation. Once portal-service tokenization has a final RPC contract, the subject-token exchange can map provider identities there without changing the gateway cookie/session runtime.
Refresh Flow
The Java handler refreshes 90 seconds before expiry and deduplicates concurrent
refreshes with RefreshTokenSingleFlight. Rust should keep that behavior.
Default Rust settings:
renewBeforeSeconds: ${statelessAuth.renewBeforeSeconds:90}
refreshSingleFlightWaitMs: ${statelessAuth.refreshSingleFlightWaitMs:5000}
refreshSingleFlightCacheMs: ${statelessAuth.refreshSingleFlightCacheMs:3000}
refreshSingleFlightMaxEntries: ${statelessAuth.refreshSingleFlightMaxEntries:10000}
These fields are Rust improvements. They can be omitted from config-server templates until a product needs to tune them.
Refresh-token request mapping should reuse client.yml
oauth.token.refresh_token and send:
grant_type=refresh_token
refresh_token=<cookie refresh token>
csrf=<new csrf>
scope=<space separated scopes, if configured>
Cookies
Cookie names should remain Java-compatible:
| Cookie | HttpOnly | Source |
|---|---|---|
accessToken | true | OAuth access token |
refreshToken | true | OAuth refresh token |
csrf | false | Generated CSRF value |
userId | false | JWT uid claim |
userType | false | JWT userType claim |
roles | false | Base64-encoded JWT role claim, default user |
host | false | JWT host claim |
email | false | JWT eml claim |
eid | false | JWT eid claim |
Access-token, user-info, and CSRF cookies should use the access token
expires_in value as Max-Age. Refresh-token cookie Max-Age should use
sessionTimeout unless the token response includes a remember value other than
N, in which case it should use rememberMeTimeout.
Java only clears cookies that were present on the request. Rust should improve logout by always emitting deletion cookies for the known cookie names, using the configured domain/path/secure attributes. This avoids stale browser cookies when a cookie is omitted from a particular request.
Default SameSite should remain None for Java parity. Add a Rust-only optional
cookieSameSite field with default None so deployments can choose Lax or
Strict when the SPA and BFF are same-site.
Config Server Model
The config-server should continue to resolve placeholders before startup:
statelessAuth.redirectUri: https://localhost:3000/#/app/dashboard
statelessAuth.cookieDomain: localhost
statelessAuth.cookieSecure: true
client.tokenAcClientId: ...
client.tokenAcClientSecret: ...
client.tokenRtClientId: ...
client.tokenRtClientSecret: ...
The Rust gateway should only consume the resolved statelessAuth.yml,
client.yml, security.yml, and handler.yml files. It should not need to
know whether the values came from product defaults, environment variables, or
light-portal overrides.
Implemented Surface
- Shared SPA cookie/session runtime, including cookie parser/writer, CSRF extraction, JWT claim extraction, and Java-compatible cookie names.
- OAuth token client support for authorization-code, refresh-token, and
token-exchange grant requests using
client.yml. - Refresh-token renewal with a bounded completed-result cache.
statelessAuth.ymlloader, module registry registration, active-handler gating, and runtime reload.stateless,google,facebook, andgithubrequest handling inlight-gateway.- Structured session-expired response using
cookieTimeoutUri. - Unit/runtime-load coverage for config parsing, cookie attributes, provider subject-token selection, active-handler loading, and gateway wiring.
MSAL Exchange Handler
Status
Initial Rust implementation is complete in light-pingora and
light-gateway. It includes config loading, named security-msal.yml
validation support, token-exchange handling, shared SPA session/cookie/CSRF
logic, logout, refresh-token renewal, handler wiring, config stubs, and
runtime-load tests.
Purpose
The Java light-spa-4j msal-exchange module is the on-prem BFF login bridge
for SPA deployments that use Microsoft Authentication Library SSO. The browser
uses MSAL.js to obtain a Microsoft token, sends that token to the gateway, and
the gateway exchanges it for an internal light-oauth token set. After exchange,
the browser session behaves the same as the stateless authorization-code
handler: internal tokens are stored in cookies, CSRF is validated on subsequent
requests, refresh tokens keep the session alive, and the gateway injects
Authorization: Bearer <internal-token> before routing downstream.
In light-fabric this should be a light-pingora security handler in
light-gateway. It should share most of its implementation with
stateless-auth.md; only the initial login exchange differs.
Goals
- Preserve the Java MSAL token-exchange flow.
- Keep
msal-exchange.ymlfield names recognizable for light-portal and config-server product configuration. - Validate the incoming Microsoft token with a separate
security-msal.ymlruntime before token exchange. - Exchange the Microsoft token with light-oauth using
client.ymloauth.token.token_exchange. - Store the returned internal token set in the same Java-compatible cookies as the stateless handler.
- Share CSRF validation, cookie writing, logout, refresh-token renewal, and
downstream
Authorizationinjection with the stateless handler. - Add a stable
msal-exchangehandler id tolight-gateway. - Register loaded config in
ModuleRegistryand fail startup on invalid active configuration.
Non-Goals
- Do not forward the Microsoft token to downstream services after exchange.
- Do not implement a server-side browser session store.
- Do not merge MSAL token validation into the normal downstream
securityhandler. MSAL validation applies only to the exchange endpoint. - Do not invent a REST-specific tokenization or portal-service client in this handler. The only outbound call is the OAuth token-exchange request.
- Do not require a separate BFF binary.
Resolved Decisions
- Support
subjectTokenTypein bothclient.ymlandmsal-exchange.yml. The handler-specific value takes precedence when set, andclient.ymlremains the shared OAuth token-exchange default. - Support strict Microsoft token validation in
security-msal.ymlwhen a deployment needs issuer and audience checks.
Java Behavior To Map
Java config file:
enabled: ${msal-exchange.enabled:true}
exchangePath: ${msal-exchange.exchangePath:/auth/ms/exchange}
logoutPath: ${msal-exchange.logoutPath:/auth/ms/logout}
logoutCsrfEnforced: ${msal-exchange.logoutCsrfEnforced:false}
cookieDomain: ${msal-exchange.cookieDomain:localhost}
cookiePath: ${msal-exchange.cookiePath:/}
cookieSecure: ${msal-exchange.cookieSecure:false}
sessionTimeout: ${msal-exchange.sessionTimeout:3600}
rememberMeTimeout: ${msal-exchange.rememberMeTimeout:604800}
Java also loads a separate security config named security-msal:
SecurityConfig.load("security-msal")
This config verifies the incoming Microsoft token. The normal security.yml
runtime verifies/parses internal light-oauth access tokens used in cookies.
Java request behavior:
POST exchangePath, normally/auth/ms/exchange, requiresAuthorization: Bearer <microsoft-token>.- Missing bearer token returns
ERR11647. - The handler verifies the Microsoft token with
security-msal.yml. - Verification failure returns
ERR10000. - The handler generates a CSRF value and sends an OAuth token-exchange request
with the Microsoft token as
subject_token. - Token-exchange failure returns
ERR11648. - On success, the handler sets the same BFF cookies as the stateless handler
and returns JSON containing
scopes. POST logoutPath, normally/auth/ms/logout, validates the readable CSRF cookie/header pair when enforcement is enabled, clears BFF cookies, and returns204 No Content.- Subsequent requests use the same cookie, CSRF, refresh, and downstream
Authorizationinjection flow as the stateless handler.
Error codes aligned:
| Code | Meaning |
|---|---|
ERR11647 | Microsoft bearer token is missing |
ERR11648 | Internal token exchange failed |
ERR10000 | Incoming Microsoft token or returned internal token is invalid |
ERR10036 | CSRF token is missing from request |
ERR10038 | CSRF claim is missing from JWT |
ERR10039 | Request CSRF and JWT CSRF do not match |
ERR10008 | Method is not allowed for a mutation endpoint |
ERR11649 | Logout CSRF cookie/header validation failed without exposing either value |
Rust Architecture
Use the shared SPA auth runtime described in stateless-auth.md.
Proposed modules:
frameworks/light-pingora/src/spa_auth.rs
frameworks/light-pingora/src/msal_exchange.rs
msal_exchange.rs owns only the Microsoft-token exchange entrypoint:
#![allow(unused)]
fn main() {
pub struct MsalExchangeConfig {
pub enabled: bool,
pub exchange_path: String,
pub logout_path: String,
pub cookie_domain: String,
pub cookie_path: String,
pub cookie_secure: bool,
pub session_timeout: u64,
pub remember_me_timeout: u64,
pub renew_before_seconds: u64,
pub subject_token_type: String,
}
pub struct MsalExchangeRuntime {
pub config: MsalExchangeConfig,
pub session: SpaSessionRuntime,
pub msal_security: SecurityRuntime,
}
}
Use msal-exchange.yml as the primary file name and accept
msal-exchange.yaml as a compatibility fallback.
The SecurityRuntime loader should be generalized so the MSAL handler can load
a named security config:
#![allow(unused)]
fn main() {
load_security_runtime_from_file(
runtime_config,
"security-msal.yml",
"light-pingora/security-msal",
"security-msal",
active,
)
}
That keeps normal downstream JWT behavior on security.yml while the exchange
endpoint validates Microsoft tokens against security-msal.yml.
Handler Registration
Add msal-exchange to apps/light-gateway handler descriptors as a security
handler:
#![allow(unused)]
fn main() {
("msal-exchange", PingoraHandlerKind::Security)
}
The primary handler id should be msal-exchange. No @alias syntax is
needed. An additional short alias such as msal can be added later only if a
real product config needs it.
Runtime loading should follow the existing active-handler model:
#![allow(unused)]
fn main() {
let msal_exchange = load_msal_exchange_runtime(
config,
active_handlers.is_handler_active("msal-exchange"),
)?;
}
If the handler is not active in handler.yml, no MSAL config is required. If
the handler is active and its config is invalid, startup should fail. If
enabled: false, register the disabled module and return None.
Example chain:
handlers:
- exception
- cors
- msal-exchange
- header
- prefix
- token
- router
chains:
bff:
- exception
- cors
- msal-exchange
- header
- prefix
- token
- router
websocket:
- exception
- msal-exchange
- security
- websocket
paths:
- path: /auth/ms/exchange
method: POST
exec:
- bff
- path: /auth/ms/exchange
method: OPTIONS
exec:
- bff
- path: /auth/ms/logout
method: POST
exec:
- bff
- path: /auth/ms/logout
method: OPTIONS
exec:
- bff
Exchange and logout are POST-only. Keep OPTIONS permanently and keep cors
before msal-exchange in the selected chain.
Exchange Flow
For exchangePath:
POST /auth/ms/exchange
Authorization: Bearer <microsoft-token>
-> extract bearer token
-> verify Microsoft token with security-msal.yml
-> generate csrf
-> call light-oauth token endpoint with token-exchange grant
-> verify/parse returned internal access token
-> set BFF cookies
-> return { "scopes": [...] }
The exchange request body is optional. A zero-length body is valid even when a
shared client declares Content-Type: application/json.
Logout Flow
POST /auth/ms/logout
Cookie: accessToken=...; csrf=...
X-CSRF-TOKEN: <csrf>
-> optionally enforce logout double-submit CSRF
-> emit deletion cookies for every cookie the runtime can set
-> return 204 No Content with no body or response content type
A legacy GET or any other unsupported exchange/logout method returns 405,
ERR10008, and Allow: POST before token-server, cookie, or proxy side
effects. Explicit OPTIONS routing continues to reach CORS.
The token-exchange request should use client.yml
oauth.token.token_exchange:
oauth.token.server_urloroauth.token.serviceIdoauth.token.enableHttp2oauth.token.token_exchange.urioauth.token.token_exchange.client_idoauth.token.token_exchange.client_secretoauth.token.token_exchange.scopeoauth.token.token_exchange.subjectTokenTypeas the default subject token type when the handler config does not override it
The form body should match Java and the http-client composer:
grant_type=urn:ietf:params:oauth:grant-type:token-exchange
subject_token=<microsoft-token>
subject_token_type=urn:ietf:params:oauth:token-type:jwt
csrf=<generated csrf>
requested_token_type=<optional requested token type>
audience=<optional audience>
scope=<space separated scopes, if configured>
The handler should set Authorization: Basic <client_id:client_secret> on the
outbound token-exchange request.
Session Validation Flow
After exchange, MSAL and stateless auth must use the same downstream request flow:
request
-> read accessToken cookie
-> verify/parse internal JWT with security.yml
-> validate CSRF from request against JWT csrf claim
-> refresh internal token when it is inside the renew window
-> inject Authorization: Bearer <internal-access-token>
-> continue handler chain
CSRF source order should be identical to the stateless handler:
X-CSRF-TOKENheader.Sec-WebSocket-Protocolvalue starting withcsrf.when the request hasSec-WebSocket-KeyandSec-WebSocket-Version.- Query parameter
csrf.
The MSAL handler must never inject the Microsoft token downstream. The only downstream bearer token after login is the internal light-oauth token.
Internal JWT Verification
MSAL exchange should use the same lower-level token verifier as stateless auth
for internal cookie tokens. It should not use the request-oriented
verify_jwt_request wrapper because the token source is a cookie, not an
Authorization header.
The shared verifier should validate signature and key material from
security.yml, parse claims for CSRF and user cookies, and support an
expiry-mode option so the refresh path can inspect tokens close to expiry
without treating that as a downstream API authentication success.
Cookies
MSAL exchange should use the same cookie contract as stateless auth:
| Cookie | HttpOnly | Source |
|---|---|---|
accessToken | true | Internal OAuth access token |
refreshToken | true | Internal OAuth refresh token |
csrf | false | Generated CSRF value |
userId | false | JWT uid claim |
userType | false | JWT userType claim |
roles | false | Base64-encoded JWT role claim, default user |
host | false | JWT host claim |
email | false | JWT eml claim |
eid | false | JWT eid claim |
For Java parity, keep cookieSecure defaulting to false in
msal-exchange.yml, but production config should set it to true when the BFF
is served over HTTPS.
Rust should share the logout improvement from stateless auth: always emit deletion cookies for known cookie names rather than only clearing cookies that were present on the request.
Security Config
security-msal.yml should be treated as an active handler dependency when
msal-exchange is active. Missing or invalid config should fail startup
because the gateway would otherwise accept an exchange endpoint without a
working Microsoft-token verifier.
Recommended distinction:
security-msal.yml: verifies the incoming Microsoft token onexchangePath.security.yml: verifies/parses internal light-oauth tokens in BFF cookies and is also used by normal API security handlers.
The Java code skips audience verification for MSAL in the current call path.
Rust should preserve compatibility unless security-msal.yml explicitly
configures audience validation support. That keeps on-prem deployments working
when the Microsoft token audience is the SPA client id rather than the BFF.
When a product requires stricter validation, security-msal.yml should be able
to require issuer and audience checks for the incoming Microsoft token. The
initial implementation can add these checks to the named SecurityRuntime
loader as optional fields:
issuer: ${security-msal.issuer:}
audience: ${security-msal.audience:}
Blank values preserve the Java-compatible relaxed behavior. Non-blank values must be enforced during exchange-path token verification, and invalid issuer/audience should return the same invalid-token error path as other Microsoft token verification failures.
Config Server Model
Light-portal should manage the product config values and config-server should deliver resolved files:
msal-exchange.exchangePath: /auth/ms/exchange
msal-exchange.logoutPath: /auth/ms/logout
msal-exchange.cookieDomain: localhost
msal-exchange.cookieSecure: true
msal-exchange.subjectTokenType: urn:ietf:params:oauth:token-type:jwt
client.tokenExClientId: ...
client.tokenExClientSecret: ...
client.subjectTokenType: urn:ietf:params:oauth:token-type:jwt
security-msal.issuer: https://login.microsoftonline.com/{tenant-id}/v2.0
security-msal.audience: <spa-client-id>
The gateway consumes only the resolved files:
handler.ymlmsal-exchange.ymlsecurity-msal.ymlsecurity.ymlclient.yml
Implemented Surface
- Shared SPA auth runtime from
stateless-auth.md. - Named
SecurityRuntimeloading forsecurity-msal.yml. - Token-exchange support in the shared OAuth token client.
msal-exchange.ymlparsing, module registry registration, active-handler gating, and runtime reload.msal-exchangerequest handling inlight-gateway.- Required bearer-token extraction, Microsoft token validation,
token-exchange request, Java-compatible cookie writing, logout, refresh
renewal, and downstream internal
Authorizationinjection. - Optional issuer/audience validation through
security-msal.yml. - Unit/runtime-load coverage for subject-token-type precedence and gateway wiring.
MSAL Auth Handler
Status
Initial Rust implementation is complete in light-pingora and light-gateway. It includes config loading (msal-auth.yml), standalone Microsoft Entra ID token validation through security-msal.yml, double-submit cookie CSRF handling, gateway auth-principal propagation, and downstream Authorization injection.
Purpose
The msal-auth module is an alternative to msal-exchange for Microsoft Entra ID single-page application (SPA) architectures where the frontend acts as the primary OAuth client.
In this flow:
- The SPA handles Microsoft authentication, token acquisition, and token refresh directly.
- The SPA submits the Entra ID access token to the gateway’s
/auth/ms/loginendpoint. - The gateway validates the Entra ID token with
security-msal.ymland sets theaccessTokenandcsrfcookies using the double-submit cookie pattern. - On subsequent API calls, the gateway validates the Microsoft JWT with expiry enforcement, compares the CSRF request value to the CSRF cookie, sets the gateway auth principal for later handlers, and forwards the token in the
Authorization: Bearerheader.
This eliminates the need for an internal light-oauth token exchange, reducing infrastructure dependencies while maintaining backend API security.
Configuration
handler.yml
Register msal-auth in the handler chain before handlers that need ctx.auth or the downstream Authorization header, such as access-control, router, or proxy handling.
handlers:
- cors
- msal-auth
- router
chains:
bff:
- cors
- msal-auth
- router
paths:
- path: /auth/ms/login
method: POST
exec:
- bff
- path: /auth/ms/login
method: OPTIONS
exec:
- bff
- path: /auth/ms/logout
method: POST
exec:
- bff
- path: /auth/ms/logout
method: OPTIONS
exec:
- bff
defaultHandlers:
- cors
- msal-auth
- router
msal-auth.yml
enabled: ${msal-auth.enabled:true}
loginPath: ${msal-auth.loginPath:/auth/ms/login}
logoutPath: ${msal-auth.logoutPath:/auth/ms/logout}
logoutCsrfEnforced: ${msal-auth.logoutCsrfEnforced:false}
cookieDomain: ${msal-auth.cookieDomain:localhost}
cookiePath: ${msal-auth.cookiePath:/}
cookieSecure: ${msal-auth.cookieSecure:false}
sessionTimeout: ${msal-auth.sessionTimeout:3600}
cookieSameSite: ${msal-auth.cookieSameSite:None}
security-msal.yml
msal-auth requires security-msal.yml when the handler is active and msal-auth.enabled is true. The config is loaded independently from the normal security.yml runtime.
enableVerifyJwt: ${security-msal.enableVerifyJwt:true}
ignoreJwtExpiry: ${security-msal.ignoreJwtExpiry:false}
enableRelaxedKeyValidation: ${security-msal.enableRelaxedKeyValidation:false}
issuer: ${security-msal.issuer:}
audience: ${security-msal.audience:}
jwt:
clockSkewInSeconds: ${security-msal.jwt.clockSkewInSeconds:60}
Handlers
- Login (
/auth/ms/login): Expects an Entra ID token in theAuthorization: Bearerheader. Validates it using thesecurity-msalruntime with expiry enforcement. Generates a secure CSRF token and returns bothaccessTokenandcsrfasSet-Cookieheaders. - Logout (
POST /auth/ms/logout): Validates logout CSRF when configured, clears every cookie the runtime sets (accessTokenandcsrf), and returns204 No Contentwithout a response body or content type. - Session Validation (any path with cookies): Reads the
accessTokencookie. Validates the JWT with expiry enforcement. Checks that the CSRF request value matches the CSRF cookie. If valid, it sets the gateway auth principal and forwards theaccessTokendownstream in theAuthorization: Bearerheader.
Frontend Integration
The Single Page Application (SPA) must coordinate with the gateway for session creation and destruction.
Login Request
When the SPA acquires an access token from Microsoft Entra ID (e.g., using
MSAL.js), it must send that token to the gateway’s login endpoint to establish
the secure HTTP-only cookies. Both login and logout use POST; neither needs a
request body. A zero-length body is also accepted when a shared client sets
Content-Type: application/json.
For cross-origin deployments, both examples require preflight: login sends the
non-safelisted Authorization header and CSRF-protected logout sends the
non-safelisted X-CSRF-TOKEN header. Keep the explicit OPTIONS routes and
qualify the exact origin, credentials, method, and requested headers.
async function gatewayLogin(entraIdToken) {
const response = await fetch('/auth/ms/login', {
method: 'POST',
credentials: 'include',
headers: {
'Authorization': `Bearer ${entraIdToken}`
}
});
if (!response.ok) {
throw new Error('Failed to create gateway session');
}
console.log('Gateway session established');
}
Logout Request
When the user logs out, the SPA must call the gateway’s logout endpoint to
clear the HTTP-only session cookies. Send credentials and the
X-CSRF-TOKEN header read from the readable csrf cookie; no body is needed.
// Helper to read the csrf cookie
function getCookie(name) {
const value = `; ${document.cookie}`;
const parts = value.split(`; ${name}=`);
if (parts.length === 2) return parts.pop().split(';').shift();
}
async function gatewayLogout() {
const csrfToken = getCookie('csrf');
const response = await fetch('/auth/ms/logout', {
method: 'POST',
credentials: 'include',
headers: {
'X-CSRF-TOKEN': csrfToken
}
});
if (response.status !== 204) {
throw new Error(`Unexpected logout status ${response.status}`);
}
console.log('Gateway session cleared');
}
Login and logout are POST-only. A legacy GET or any other unsupported method
returns 405, ERR10008, and Allow: POST before authentication or cookie
side effects. Keep explicit OPTIONS routes permanently with cors before
msal-auth in the selected chain.
Error Handling
| Code | Meaning |
|---|---|
ERR10008 | Method is not allowed; MSAL login/logout responses advertise Allow: POST. |
ERR10036 | Logout CSRF header is missing when enforcement is enabled. |
ERR11649 | Logout CSRF cookie/header validation failed without exposing either value. |
API Request
For standard API calls to backend services, the browser will automatically include the HTTP-only accessToken cookie. However, any request that modifies state or requires CSRF protection must include the CSRF token. The SPA must read the csrf cookie and append it as the X-CSRF-TOKEN header.
async function callBackendApi(endpoint, data) {
const csrfToken = getCookie('csrf');
const response = await fetch(endpoint, {
method: 'POST', // or PUT, DELETE, etc.
headers: {
'Content-Type': 'application/json',
'X-CSRF-TOKEN': csrfToken
},
body: JSON.stringify(data)
});
if (!response.ok) {
throw new Error('API call failed');
}
return response.json();
}
WebSocket Connection
The browser’s native WebSocket API does not allow setting custom HTTP headers. To pass the CSRF token during the WebSocket handshake upgrade, the SPA must pass it as a subprotocol string prefixed with csrf.. The gateway will extract and validate it.
function connectWebSocket(path) {
const csrfToken = getCookie('csrf');
// Create a subprotocol string that the gateway recognizes
const csrfProtocol = `csrf.${csrfToken}`;
// Note: Depending on your WebSocket server, you may also need to pass
// the actual subprotocol you intend to use (e.g., 'wamp', 'graphql-ws')
// alongside the csrf protocol.
const ws = new WebSocket(`wss://api.example.com${path}`, [csrfProtocol]);
ws.onopen = () => {
console.log('WebSocket connected securely');
};
ws.onerror = (error) => {
console.error('WebSocket connection failed (possible CSRF or Auth issue)', error);
};
return ws;
}
Double Submit Cookie CSRF
Because an Entra ID token cannot be minted with a custom CSRF claim by this gateway, msal-auth enforces CSRF protections using the double-submit cookie pattern. The SPA reads the generated csrf cookie and submits it back.
The CSRF value is accepted from the following sources, in order of precedence:
X-CSRF-TOKENheader.Sec-WebSocket-Protocolvalue starting withcsrf.(when the request hasSec-WebSocket-KeyandSec-WebSocket-Version). This provides specialized CSRF support for Websocket upgrades since browser WebSockets cannot send custom HTTP headers.- Query parameter
csrf.
The gateway compares the value from one of these sources against the csrf cookie. If they match, the session is validated.
Refresh Flow
Unlike msal-exchange or stateless-auth, msal-auth does not issue or manage refresh tokens. The SPA is responsible for using MSAL.js to silently refresh the Entra ID token and calling /auth/ms/login again to update the session cookies before they expire.
Reload Behavior
The gateway reloads msal-auth when handler.yml, msal-auth.yml, or security-msal.yml changes. Reloading security-msal.yml refreshes both msal-auth and msal-exchange because both handlers validate Microsoft tokens with that security runtime.
Unified Security Handler
Status: Phase 4 partially implemented; jwkServiceIds and sjwkServiceIds per-prefix JWK routing are wired, SJWT routing is implemented, and SWT introspection remains outstanding.
Purpose
Light Fabric’s light-gateway serves as a shared API gateway for multiple upstream
services that may belong to different organizations or security domains. In this shared
model, different request path prefixes need different authentication strategies:
- An internal
/adminroute may require HTTP Basic authentication. - A customer-facing
/api/ordersroute may require a JWT from the company’s own identity provider. - A partner
/salesforceroute may require a JWT issued by Salesforce with its own JWK endpoint. - A webhook
/webhookroute may require an API key.
The UnifiedSecurityHandler (Java) / unified-security handler (Rust) solves this by
providing a single, path-prefix-aware security dispatch point. It replaces the need to
wire separate security handlers into independent handler chains for each path family.
Java Reference
The canonical implementation lives in:
- Handler:
light-4j/unified-security/src/main/java/com/networknt/security/UnifiedSecurityHandler.java - Config:
light-4j/unified-config/src/main/resources/config/unified-security.yml
The Java handler:
- Loads
UnifiedSecurityConfigon every request (double-checked locking, hot-reload safe). - Checks
anonymousPrefixesfirst — if the path matches, all security checks are skipped. - Iterates
pathPrefixAuths; the first matching prefix wins. - For the matched rule, checks which auth methods are enabled (
basic,jwt,sjwt,swt,apikey) and dispatches to the corresponding sub-handler. - Passes
jwkServiceIds/sjwkServiceIds/swtServiceIdsto the sub-handler so it can fetch JWKs from the correct per-prefix OAuth/JWK server. - Returns
ERR10078 MISSING_PATH_PREFIX_AUTHif no rule matches any prefix.
Rust Implementation Location
frameworks/light-pingora/src/unified_security.rs
The Rust implementation is loaded in apps/light-gateway/src/main.rs when the
unified-security or unified handler IDs appear in the active handler chain:
#![allow(unused)]
fn main() {
let unified_security_config = load_unified_security_config(
&runtime_config,
handler_active(&active_handlers, &["unified-security", "unified"]),
)?;
}
Configuration
unified-security.yml
# Enable or disable this handler.
enabled: ${unified-security.enabled:true}
# Paths that bypass all security checks.
# Accepts comma-separated string, JSON array string, or YAML list.
anonymousPrefixes: ${unified-security.anonymousPrefixes:[]}
# Per-prefix authentication rules.
# Accepts comma-separated string, JSON array string, or YAML list of objects.
pathPrefixAuths: ${unified-security.pathPrefixAuths:[]}
Per-Prefix Rule Fields
| Field | Type | Purpose |
|---|---|---|
prefix | String | Path prefix to match. Longest matching prefix wins (Rust) / first wins (Java). |
basic | bool | Allow HTTP Basic authentication for this prefix. |
jwt | bool | Require Bearer JWT verification for this prefix. |
sjwt | bool | Allow Simple-JWT (no scopes) for this prefix. |
swt | bool | Allow SWT (opaque token introspection) for this prefix. |
apikey | bool | Allow API key authentication for this prefix. |
jwkServiceIds | Vec<String> | JWK service IDs (from client.yml) used to verify JWT tokens for this prefix. |
sjwkServiceIds | Vec<String> | JWK service IDs used to verify SJWT tokens for this prefix. |
swtServiceIds | Vec<String> | Introspection service IDs used to verify SWT tokens for this prefix. |
Example values.yml Entry
handler.handlers:
- correlation
- headers
- unified-security
- proxy
handler.defaultHandlers:
- default
unified-security.anonymousPrefixes:
- /health
- /server/info
unified-security.pathPrefixAuths:
- prefix: /salesforce
jwt: true
jwkServiceIds:
- com.networknt.oauth2-salesforce-1.0.0
- prefix: /blackrock
jwt: true
jwkServiceIds:
- com.networknt.oauth2-blackrock-1.0.0
- prefix: /admin
basic: true
- prefix: /webhook
apikey: true
- prefix: /internal
jwt: true
Why unified-security.yml Is Not in the light-gateway Config Folder
The config/ directory in apps/light-gateway contains only active handler
configurations that the current local development profile uses. The local profile
(defined by config/values.yml) activates only correlation, headers, and proxy.
Because unified-security is not in that handler chain, load_unified_security_config
returns None and the file is never needed.
A production or staging deployment that enables unified security would receive
unified-security.yml from config-server, populated by the light-portal product
configuration for that deployment. To use it locally, add unified-security to
handler.handlers and handler.defaultHandlers (or a path-specific chain) in
values.yml, then add a unified-security.yml to the config/ directory.
Prefix Matching: Java vs. Rust
| Behavior | Java | Rust |
|---|---|---|
| Match algorithm | First matching prefix in list order | Longest matching prefix (most specific wins) |
| Tie-breaking | Order in config list | Longest prefix.len() |
The Rust best_auth_rule function uses max_by_key(|rule| rule.prefix.len()), which
is intentionally more deterministic than Java’s iteration order. This means /api/v2
will match before /api regardless of declaration order.
Authentication Dispatch Logic
Request arrives at unified-security handler
│
├── anonymousPrefixes match? → Pass through (no auth)
│
├── No matching pathPrefixAuth rule? → 403 ERR10078
│
└── Matched rule:
├── basic=true OR jwt=true OR sjwt=true OR swt=true?
│ ├── No Authorization header → 401
│ ├── Scheme=Basic AND basic=true → BasicAuth verify
│ ├── Scheme=Bearer:
│ │ ├── jwt=true → JWT verify (using jwkServiceIds)
│ │ ├── sjwt=true → SJWT verify (using sjwkServiceIds)
│ │ └── swt=true → SWT introspect (using swtServiceIds) [⚠ Gap: not implemented]
│ └── Unknown scheme → 401
└── apikey=true (only) → API Key verify
Current Implementation Status
Implemented ✅
| Capability | Location |
|---|---|
UnifiedSecurityConfig and UnifiedPathAuth deserialization | unified_security.rs:15–55 |
anonymousPrefixes bypass | unified_security.rs:153–158 |
pathPrefixAuths parsing (YAML, JSON-string, comma-string) | via deserialize_typed_list |
Longest-prefix rule selection (best_auth_rule) | unified_security.rs:160–169 |
| Basic auth dispatch | unified_security.rs:113–123 |
| JWT/SJWT dispatch (Bearer) | unified_security.rs:126–131 |
jwkServiceIds / sjwkServiceIds JWK routing | security.rs |
| API key dispatch | unified_security.rs:145–149 |
Hot-reload via ConfigManager and UnifiedSecurityReloader | main.rs:1374–1410 |
Handler IDs: unified-security, unified | main.rs:114, 121, 128, 131, 133 |
Gaps ⚠️
Gap 1 — SWT (opaque token) introspection not implemented (Low)
When swt=true, the Rust handler returns HTTP 501. SWT introspection requires
calling an OAuth2 introspection endpoint, which needs service discovery and client
credentials.
Fix: Implement SWT introspection using the existing client.yml OAuth provider
infrastructure once service discovery is stable.
Recently Closed Gaps
jwkServiceIds and sjwkServiceIds per-prefix JWK routing
verify_unified_security now passes the matched rule’s jwkServiceIds or
sjwkServiceIds list into JWT verification. The Rust verifier tries the configured
service IDs in order for JWK lookup and accepts any matching configured audience.
SJWT routing
Java supports two SJWT modes:
sjwt=true, jwt=false— always treated as SJWT.sjwt=true, jwt=true— pre-parses the JWT to check forscope/scpclaim to distinguish SJWT (no scope) from a full JWT (with scope).
Rust now implements the same routing split. Non-JWT Bearer tokens are routed to SWT
when swt=true; otherwise they are rejected as unsupported Bearer tokens.
unified-security.yml added to light-gateway config folder
The sample config/ directory now includes an example unified-security.yml, making
the expected configuration clearer when activating the handler.
Java uses first-match; Rust uses longest-match (Design difference)
This is an intentional Rust improvement, not a bug, but it should be documented clearly so operators migrating from Java understand that configuration ordering matters less in Rust. The design difference is already captured in this document.
Interaction with security.yml
unified-security and security.yml (standalone JWT handler) are mutually exclusive
in a given handler chain. Do not include both unified-security and jwt handler IDs
in the same chain; the security check would be applied twice.
When unified-security is active:
security.ymlis still loaded to provide theSecurityRuntime(JWK cache, config).basic-auth.ymlis loaded if any rule hasbasic: true.apikey.ymlis loaded if any rule hasapikey: true.
Interaction with client.yml
JWK source resolution uses the client.yml OAuth/JWK configuration:
oauth:
token:
key:
serviceId: com.networknt.oauth2-token-1.0.0
serviceIdAuthServers:
com.networknt.oauth2-salesforce-1.0.0:
server_url: https://login.salesforce.com
uri: /id/keys
com.networknt.oauth2-blackrock-1.0.0:
server_url: https://idp.blackrock.com
uri: /.well-known/jwks.json
jwkServiceIds: [com.networknt.oauth2-salesforce-1.0.0] in a
pathPrefixAuth rule will cause the JWT verifier to fetch and cache JWKs from
https://login.salesforce.com/id/keys for that path prefix only.
Verification Plan
Existing Tests
tests::unified_security_accepts_java_style_listsinunified_security.rs— verifies YAML/JSON deserialization for anonymousPrefixes and pathPrefixAuths.
Tests to Add
-
jwkServiceIdsoverride — mock two JWK servers; configure two prefixes pointing to different service IDs; verify that JWT verification for each prefix fetches from the correct server. -
SJWT scope detection — provide a JWT with and without a
scopeclaim; verify thatsjwt=true, jwt=trueroutes to the correct verifier. -
SJWT-only rule —
sjwt=true, jwt=false; verify the handler always uses the SJWT verifier regardless of scope presence. -
SWT rule — configure
swt=truewith a mock introspection endpoint; verify the handler calls introspection with the correct service ID. -
No-match returns 403 — request a path not covered by any prefix; verify 403 with
ERR10078. -
Anonymous prefix bypass — request a path in
anonymousPrefixes; verify no auth header is required.
HMAC Webhook Authentication
Status: Implemented through Phase 4 local qualification for light-gateway,
informed by the completed Java implementation. The selected gateway-core
pre-buffer hook, configuration and policy compilation, secret resolution,
raw-body verification, replay stores, gateway integration, and protected replay
administration are implemented. Production HTTP/2 over TLS, multi-process
gateway/Redis, and deployed GitHub-to-Jenkins acceptance remain release gates.
The first provider profile is GitHub.
Tracking issue: networknt/light-4j#2772
Purpose
light-gateway needs to authenticate webhook requests whose sender proves
possession of a shared secret by signing the request body. The first use case
is a GitHub webhook that triggers a Jenkins build, but the verifier must be
configurable enough to support other providers that use the same raw-body HMAC
model.
HMAC is an authentication mechanism. A standalone hmac handler owns HMAC-only
route policy, while Unified Security owns composed route policy. The
cryptographic and body-buffering implementation is shared, and route
configuration can require:
- HMAC by itself;
- HMAC and JWT; or
- HMAC and API key.
Every required factor must pass before the request is admitted. A verified request body must be forwarded without parsing, re-encoding, or otherwise changing its entity bytes.
Document Boundary
This page is the Rust light-fabric implementation design. It also records the
external contract that Java and Rust should share, such as the GitHub headers,
signature format, body-size limit, and duplicate response behavior.
The Java implementation should have a separate design in light-4j. Java is in
maintenance mode, uses first-match prefix routing, and needs the smallest
handler-chain change compatible with its existing production deployments. This
Rust design does not require Java to adopt Rust’s longest-prefix matcher,
request lifecycle, replay-store implementation, or internal type model.
Resolved Decisions
- GitHub is the first provider profile. Tests for another real provider are not required.
- Version 1 supports HMAC-SHA-256 over the exact raw entity body.
- Provider-specific signing algorithms that include timestamps, paths, or selected headers are future strategies, not a configurable canonicalization language in version 1.
- Signature header, prefix, encoding, secret selection, replay ID header, body limit, and replay retention are configurable by profile.
- GitHub uses
X-Hub-Signature-256, thesha256=prefix, hexadecimal encoding, andX-GitHub-Deliveryfor duplicate suppression. X-GitHub-Hook-IDselects one or more candidate secrets. One shared default secret is also supported when explicitly configured.- Secrets are loaded from named environment variables. Secret values never appear in configuration snapshots, logs, metrics, or management responses.
- Secret rotation uses an ordered active/previous list. Module reload can switch among environment variables that were present when the process started. Changing an environment variable’s value requires a rolling process restart.
- The maximum request body is configurable and defaults to 16 MiB.
- Non-identity
Content-Encodingis rejected in version 1. - Duplicate deliveries return an empty
200response and are not sent to the upstream service. - A failed upstream invocation releases its replay reservation. A successful
2xxkeeps the reservation until its retention period expires. - Replay retention is configurable and defaults to seven days. GitHub currently
supports manual redelivery for deliveries from the previous three days and
reuses the original
X-GitHub-Deliveryvalue. - A purpose-specific replay-store trait supports local and distributed
implementations. Every replay-enabled profile explicitly selects a configured
provider;
type: localis an intentional deployment choice, not a fallback. - An unavailable explicitly configured distributed store fails closed with
503; it does not silently fall back to local state. - Rust retains longest-prefix route selection and adds method-aware matching. Java retains its current first-match behavior.
- HMAC-protected routes cannot also match
anonymousPrefixes. - The reusable HMAC body gate has two handler entry points: the standalone
hmachandler for HMAC-only routes, and thehmacfactor insideunified-securitywhen JWT or API-key composition is required.
GitHub’s signature and redelivery behavior are documented in:
Goals
- Authenticate a GitHub webhook before any request body reaches Jenkins.
- Preserve the exact authenticated entity-body bytes for proxy forwarding.
- Compose HMAC with the existing JWT and API-key mechanisms.
- Support one shared secret or a header-selected secret map.
- Allow an active and previous secret during rotation.
- Suppress duplicate deliveries atomically.
- Provide process-local and distributed replay-store implementations.
- Permit an authorized operator to remove one replay record before intentional redelivery.
- Hot-reload policy, profiles, and pre-provisioned secret references atomically.
- Keep logs and metrics useful without exposing signatures, secrets, bodies, or high-cardinality delivery identifiers.
Non-Goals
- Do not create an arbitrary canonical signing-expression language.
- Do not parse JSON or form data before signature verification.
- Do not support decompressed-body verification in version 1.
- Do not promise that proxy hop-by-hop headers or HTTP chunk boundaries remain byte-for-byte identical. Existing Pingora proxy normalization still applies.
- Do not implement provider-specific GitHub event filtering or Jenkins build semantics.
- Do not make a non-idempotent callee safe. Jenkins or the service that starts a build must still enforce its own idempotency key.
- Do not build a complete Rust HTTP session framework as a prerequisite. The replay-store trait is intentionally smaller and requires atomic reserve/release semantics that a general session CRUD API does not provide.
- Do not redesign Java prefix matching or Unified Security in this document.
- Do not support HMAC body gating for direct application handlers such as MCP or LLM endpoints in the first phase. Version 1 targets proxy/router chains.
Threat Model and Security Invariants
The design protects against:
- body tampering without possession of the configured secret;
- use of an unknown or missing secret selector;
- replay of the same GitHub delivery within the configured retention window;
- concurrent delivery of the same replay ID to one or more gateway instances;
- accidental authentication bypass through
anonymousPrefixes; - signature verification against parsed, normalized, decompressed, or otherwise reconstructed content; and
- partial authentication where HMAC passes but a required JWT or API key does not.
The following invariants are mandatory:
- HMAC input is exactly the entity-body byte sequence received from the downstream connection.
- HMAC comparison is constant-time.
- All configured authentication factors pass before upstream selection.
- No request-body byte is forwarded before HMAC validation and replay reservation succeed.
- The forwarded entity body is the same byte sequence that was authenticated.
- The untrusted hook ID only selects candidate secrets; it is never accepted as authenticated identity on its own.
- Replay reservation is an atomic insert-if-absent operation.
- A configured distributed replay-store outage fails closed.
- Secret material is excluded from serializable module configuration and operational output.
- A runtime reload either installs one completely validated security runtime or leaves the previous runtime active.
Replay Limitation
GitHub signs the request body, not X-GitHub-Delivery. Duplicate suppression
using that header follows GitHub’s recommendation and stops an unchanged
captured delivery, but it is not a complete cryptographic anti-replay protocol:
an attacker who possesses a valid body and signature could change an unsigned
delivery header. A finite cache also permits replay after expiry, local state is
lost on restart, and a process can fail after reserving an ID but before invoking
the upstream.
These limitations make callee idempotency mandatory. A future strict profile may additionally fence a signature/body fingerprint, with the documented tradeoff that two legitimate deliveries with identical bodies could be treated as duplicates.
Architecture
The design separates policy, HMAC mechanics, request-body gating, and replay storage:
handler.yml unified-security.yml
| |
| standalone `hmac` | `hmac` factor
v v
hmac.yml standalone rule Unified Security compiler
| |
+------------------+--------------------------+
| references profile
v
hmac.yml profile + keyring
|
v
reusable pre-upstream body gate
|
HMAC verification
|
replay reservation
|
v
normal Pingora proxy
|
retain or release replay ID
Standalone and Unified Security Entry Points
HMAC follows the existing JWT and API-key integration model. The hmac handler
can protect an HMAC-only route independently. When a route requires HMAC plus
JWT or API key, unified-security owns the composed policy and invokes the same
HMAC body gate as one required factor.
The cryptographic verifier, bounded body reader, replay reservation, and request context are shared. The two entry points differ only in policy selection:
- standalone
hmacselects a method-awarehmac.yml.pathPrefixAuthsrule; and unified-securityselects a method-awareunified-security.yml.pathPrefixAuthsrule and passes its HMAC profile to the shared gate.
A request must use only one authentication-policy entry point. Startup and
reload reject an effective chain that contains standalone hmac and
unified-security for the same path and method, even when the Unified Security
rule is JWT-only or API-key-only. Authentication composition belongs in one
Unified Security allOf; it must not emerge accidentally from two independent
handlers. They also reject a protected route whose effective chain omits or
disables the selected entry point.
The inverse is validated as well: every path and method mapped to a runnable
chain containing standalone hmac must be covered by at least one standalone
HMAC rule, from which longest-prefix matching selects one. Overlapping catch-all
and more-specific prefixes are valid; the separate duplicate-prefix rule is the
uniqueness constraint. A default chain containing hmac therefore needs a
covering rule, such as prefix / with the applicable methods, or startup/reload
fails. This turns an otherwise per-request fail-closed 503 into a configuration
error. Validation is performed against the runnable chain after handler
references and module enabled states are resolved, not against the raw
handler.yml exec list.
The standalone hmac handler retains a defensive fail-closed response if no
rule matches at runtime, but a valid compiled snapshot cannot reach that state.
For example, HMAC-only and composed routes use different handler chains while sharing the same HMAC implementation and profiles:
# handler.yml excerpt
handlers:
- correlation
- hmac
- unified-security
- limit
- router
chains:
github-webhook:
exec: [correlation, limit, hmac, router]
partner-webhook:
exec: [correlation, unified-security, limit, router]
paths:
- path: /github-webhook
method: post
exec: [github-webhook]
- path: /partner-webhook
method: post
exec: [partner-webhook]
For the HMAC-only chain, limit deliberately runs before hmac so an
unauthenticated sender cannot force a full body buffer and replay-store
round-trip for every attempt. The pre-HMAC limiter must be body-independent and
keyed only by trusted connection and route attributes, not by an unverified
selector or delivery header. A deployment may place an additional authenticated
limiter after HMAC when it needs identity-aware quotas.
The composed chain deliberately keeps limit after unified-security.
Unified Security checks the required JWT or API key before it buffers the body,
so an invalid header credential is rejected early; after authentication, the
limiter can safely apply identity-aware quotas. A deployment may additionally
place the same body-independent connection/route limiter before Unified
Security when it needs both protections.
The existing legacy boolean fields remain supported. A route uses either the
legacy fields or the new authentication object, never both:
pathPrefixAuths:
# Existing configuration: behavior remains unchanged.
- prefix: /legacy-api
jwt: true
jwkServiceIds:
- com.networknt.oauth2-token-1.0.0
# HMAC only.
- prefix: /github-webhook
methods: [POST]
authentication:
allOf:
- type: hmac
profile: github
# HMAC and JWT.
- prefix: /partner-webhook
methods: [POST]
authentication:
allOf:
- type: hmac
profile: partner
- type: jwt
jwkServiceIds:
- com.networknt.oauth2-partner-1.0.0
# HMAC and API key.
- prefix: /signed-build
methods: [POST]
authentication:
allOf:
- type: hmac
profile: build-system
- type: apiKey
Version 1 supports one HMAC factor and at most one header authentication factor
per allOf. JWT or API key is checked first so an invalid header factor can be
rejected without buffering the request body. Acceptance is still contingent on
all factors.
Route Matching
Both standalone HMAC rules and Unified Security rules select the longest
matching prefix among rules whose optional methods contains the request
method. An absent or empty method list means all methods for backward
compatibility. Matching retains Rust’s existing raw-prefix behavior; for
example, /hook also matches /hook-v2. Operators should use a trailing slash
or an exact handler path when a path-segment boundary is required.
Configuration loading rejects:
- duplicate rules with the same prefix and overlapping methods;
- a rule that mixes legacy booleans with
authentication; - an unknown factor type or HMAC profile;
- more than one HMAC factor;
- an empty
allOf; and - any overlap between an HMAC-protected route and
anonymousPrefixes.
For every pair of prefix-overlapping rules where exactly one selected policy
requires HMAC, validation proves that no path-and-method combination in the
overlap can fall through to the non-HMAC rule. This comparison crosses policy
sources: standalone hmac.yml rules are checked against legacy JWT/API-key
rules in unified-security.yml, not only against composed HMAC rules. It
rejects both:
- a more-specific non-HMAC rule whose methods override a broader HMAC rule; and
- a broader non-HMAC ancestor that matches methods omitted by a more-specific HMAC rule.
For example, an all-method JWT rule on /webhook cannot be combined with a
POST-only HMAC rule on /webhook/github, because PUT
/webhook/github would fall through to JWT-only authentication. The HMAC rule
must cover every method inherited from the ancestor, the ancestor must exclude
the uncovered methods, or the routes must not overlap. Version 1 has no implicit
security downgrade or allowHmacOverride escape hatch. Standalone and Unified
Security HMAC rules may reuse a profile, but their effective handler-chain
coverage may not overlap.
The overlap check applies only when an HMAC policy is introduced. Existing legacy configurations without HMAC retain their current anonymous-prefix behavior.
HMAC Profiles
hmac.yml contains standalone route-to-profile rules, cryptographic profiles,
and replay-store configuration. A profile itself does not contain route
prefixes. pathPrefixAuths is consulted only by the standalone hmac handler;
Unified Security continues to own its own route rules.
enabled gates the shared HMAC module, not one particular entry point. The
module is required when either an effective chain contains hmac or an enabled
Unified Security rule references an HMAC factor. A required but disabled or
missing module fails startup/reload.
enabled: true
maxBufferedBodyBytes: 268435456
pathPrefixAuths:
- prefix: /github-webhook
methods: [POST]
profile: github
profiles:
github:
signedInput: rawBody
algorithm: hmacSha256
signatureHeader: X-Hub-Signature-256
signaturePrefix: "sha256="
signatureEncoding: hex
maxBodyBytes: 16777216
bodyReadTimeoutMillis: 10000
secrets:
selectorHeader: X-GitHub-Hook-ID
bySelector:
"12345678":
- GITHUB_HOOK_12345678_CURRENT_SECRET
- GITHUB_HOOK_12345678_PREVIOUS_SECRET
"87654321":
- GITHUB_HOOK_87654321_CURRENT_SECRET
defaultEnvNames: []
replay:
enabled: true
idHeader: X-GitHub-Delivery
store: webhook-replay
retentionSeconds: 604800
shared-build-system:
signedInput: rawBody
algorithm: hmacSha256
signatureHeader: X-Build-Signature
signaturePrefix: ""
signatureEncoding: base64
maxBodyBytes: 16777216
bodyReadTimeoutMillis: 10000
secrets:
selectorHeader: ""
bySelector: {}
defaultEnvNames:
- BUILD_WEBHOOK_CURRENT_SECRET
- BUILD_WEBHOOK_PREVIOUS_SECRET
replay:
enabled: true
idHeader: X-Build-Delivery
store: webhook-replay
retentionSeconds: 604800
replayStores:
webhook-replay:
type: redis
urlEnv: WEBHOOK_REPLAY_REDIS_URL
keyPrefix: "light:hmac-replay:"
connectTimeoutMillis: 1000
operationTimeoutMillis: 1000
Supported version 1 values are:
| Field | Values and behavior |
|---|---|
maxBufferedBodyBytes | Positive gateway-wide HMAC buffering budget; default 268435456 bytes (256 MiB) |
signedInput | rawBody only |
algorithm | hmacSha256 only |
signatureEncoding | hex or base64 |
signaturePrefix | Exact optional prefix removed before decoding |
maxBodyBytes | Positive value, default and recommended maximum 16 MiB |
bodyReadTimeoutMillis | Positive bounded body-read timeout |
selectorHeader | Optional header used for exact secret-map lookup |
defaultEnvNames | Explicit shared-secret fallback; empty means no fallback |
retentionSeconds | Positive TTL; default seven days |
Each standalone rule requires a non-empty prefix and a known profile.
When methods is present and non-empty, every value is normalized and
validated; absent or empty means all methods. Duplicate prefixes with
overlapping methods are rejected. Each replay-enabled profile requires a
non-empty replay.store that references exactly one declared provider.
maxBufferedBodyBytes must be at least as large as the largest enabled
profile’s maxBodyBytes. Size it no higher than the memory the process can
safely dedicate to webhook bodies and normally near maxBodyBytes multiplied by
the intended concurrent HMAC-body admissions. The implementation acquires
budget incrementally as chunks arrive or reserves the known bounded
Content-Length; it never allocates beyond the acquired amount and releases all
budget on every terminal path.
Header names are case-insensitive according to HTTP rules. Multiple values for the signature, selector, or replay ID header are rejected. Selector values are matched exactly after trimming optional HTTP whitespace; they are not parsed as numbers.
The ordered secret list is active first, previous second. Verification computes
and compares every configured candidate rather than returning after the first
match. This bounds secret-version timing differences. Configuration limits each
selector candidate list and defaultEnvNames to at most two entries.
Missing or unknown selectors fail unless defaultEnvNames is explicitly
non-empty. If several hooks use one default secret, the selector header is not
an authenticated hook identity.
Secret Resolution and Rotation
The serializable config retains environment-variable names only. Runtime loading resolves them into a non-serializable keyring whose debug and serialization representations are redacted.
Startup and reload fail if a referenced environment variable is absent, empty, or cannot initialize HMAC. Deployments should use randomly generated secrets of at least 32 bytes, although compatibility with an existing provider secret may require accepting a shorter non-empty value.
Safe rotation is:
- Provision
CURRENTandPREVIOUSenvironment variables before process startup. - Configure the ordered list as
[CURRENT, PREVIOUS]and reload the module. - Change the provider to use
CURRENT. - After the overlap window, remove
PREVIOUSfromhmac.ymland reload. - Use a rolling restart when the bytes assigned to an environment variable must change.
An in-flight request pins the Arc<GatewaySecurityExecutionSnapshot> captured
before its handler chain and policy are selected. Reload never changes its
handlers, factors, secrets, or replay-store reference halfway through that
request.
Rust Request Lifecycle
The current verify_unified_security function executes in
GatewayProxy::request_filter, while GatewayProxy::request_body_filter
normally sees streaming chunks after upstream selection. HMAC cannot be added
only to the streaming filter: duplicate detection must be able to return a local
200 without contacting Jenkins, and no body may reach the upstream before
validation.
For a matched HMAC policy, the pre-upstream flow is:
sequenceDiagram
participant Sender as GitHub / sender
participant Filter as request_filter
participant HMAC as HMAC verifier
participant Replay as Replay store
participant Body as request_body_filter
participant Upstream as Jenkins / upstream
Sender->>Filter: Headers and request body
Filter->>Filter: Match longest prefix + method
opt composed policy has a header factor
Filter->>Filter: Verify required JWT or API key
end
Filter->>Filter: Reject non-identity content encoding
Filter->>Filter: Read bounded raw body before upstream selection
Filter->>HMAC: Verify exact raw bytes
HMAC-->>Filter: Valid
Filter->>Replay: reserve(profile, selector, delivery, TTL)
alt duplicate
Replay-->>Filter: Duplicate
Filter-->>Sender: 200, empty body
else reserved
Replay-->>Filter: Reservation handle
Filter->>Filter: Store verified bytes and handle in request context
Filter->>Body: Continue normal proxy lifecycle
Body->>Upstream: Inject exact verified bytes once, end of stream
Upstream-->>Filter: Upstream response
alt upstream status is 2xx
Filter->>Filter: Keep reservation until TTL
else transport failure or non-2xx
Filter->>Replay: release(reservation handle)
end
Filter-->>Sender: Upstream response
end
The request context needs at least:
#![allow(unused)]
fn main() {
struct PendingVerifiedBody {
bytes: bytes::Bytes,
injected: bool,
}
enum WebhookReplayState {
NotRequired,
Reserved(ReplayReservation),
Committed2xx,
Releasing,
Released,
}
}
The context also pins the Arc<GatewaySecurityExecutionSnapshot> and records
whether the HMAC gate was entered through standalone hmac or
unified-security. A second entry attempt is a fail-closed chain error.
request_filter consumes the downstream body with the configured size and time
bounds. A Content-Length above the limit is rejected immediately, but framing
metadata is never trusted as proof of the actual size. The reader continues
until end-of-stream or maxBodyBytes + 1; exactly maxBodyBytes is accepted only
after end-of-stream is observed, and the first extra byte produces 413.
Once verification and reservation pass, request_body_filter injects the
stored bytes as the single final upstream body chunk. Later body-aware filters,
including tokenization and request access control, operate on that re-injected
body. HMAC must always run before any body mutation.
In addition to the per-profile limit, the gateway enforces a configurable weighted admission budget for the total number of HMAC body bytes buffered by in-flight requests. Budget exhaustion fails closed before allocating the full body. The permit is request-owned and released on rejection, duplicate, reinjection, cancellation, or final completion.
The security feature does not modify application headers. Existing proxy
behavior may still normalize hop-by-hop headers and transfer framing. The
entity body, Content-Type, GitHub event headers, and other end-to-end headers
continue through the normal proxy path.
Phase 0 Body-Gate Proof
Before implementing the complete feature, a focused spike must prove this lifecycle against the pinned Pingora version:
request_filtercan consume a bounded HTTP/1.1 body before upstream peer selection.- The same flow works for HTTP/2 downstream requests.
request_body_filterreceives end-of-stream after the earlier read and can inject the stored bytes exactly once.- The upstream receives no connection or request when HMAC is invalid or the replay store reports a duplicate.
- The upstream receives the exact original entity bytes when validation passes.
Content-Lengthand chunked downstream requests both produce a correct upstream request.
An integration test must use a counting fake upstream; an in-memory verifier test is not sufficient. If Pingora cannot satisfy these assertions, the team must stop and choose either a gateway-core pre-buffer hook or a dedicated buffered proxy path. Streaming an unauthenticated request toward the upstream is not an acceptable fallback.
Phase 0 Result: Gateway-Core Pre-Buffer Hook Selected
Phase 0 first proved that the unmodified Pingora 0.8.1 lifecycle could not
meet assertions 3, 5, and 6 at the required 16 MiB limit. After
request_filter consumed a non-empty body, Pingora replayed it only while the
internal retry buffer remained available. The pinned HTTP/1.1 and HTTP/2
implementations hard-code that buffer to 64 KiB, so a 64 KiB plus one-byte body
could reach the upstream as headers without the authenticated entity bytes.
The selected resolution is a narrow gateway-core extension to the pinned
pingora-proxy crate. ProxyHttp::prebuffered_request_body is consulted only
after the original downstream body reports completion. When it returns bytes,
the proxy core sends them through the normal request_body_filter as one
end-of-stream chunk. The callback is repeatable for an upstream retry and its
default implementation returns None, leaving every non-participating proxy
unchanged.
The patched crate and provenance note are in patches/pingora-proxy. The
reproducer remains apps/hmac-phase0-spikes, its counting-upstream integration
test is apps/hmac-phase0-spikes/tests/body_gate.rs, and
scripts/run-hmac-phase0-gates.sh is the repeatable gate. The completed proof
covers exact 16 MiB capture and replay, 16 MiB plus one-byte rejection,
HTTP/1.1 content-length and chunked inputs, HTTP/2 above the old 64 KiB ceiling,
end-to-end header preservation, and local duplicate short-circuiting.
The hook is infrastructure, not authentication. Only a request whose compiled HMAC policy has completed bounded capture and verification may return a body from it. Phase 1 registers policy and verifier state but keeps standalone and composed HMAC traffic fail closed until the Phase 3 gateway lifecycle integration sets that verified request state.
Routes using HMAC are initially restricted to proxy/router chains. Startup
validation rejects a chain where a later direct application handler expects to
read the already-consumed body from Session. A future shared buffered-body
contract can remove that restriction.
HMAC Verification
The verifier performs these steps in order:
- Confirm the method and route were selected by either compiled standalone HMAC policy or compiled Unified Security policy, matching the active entry point.
- Reject a
Content-Encodingother than absent oridentitywith415. - Read through end-of-stream with a
maxBodyBytes + 1overflow probe; return413on the first extra byte and408on timeout. - Read exactly one signature header and remove the configured prefix.
- Decode the remaining signature as configured hex or base64 and require a SHA-256-sized result.
- Select the candidate secret list using the optional selector header.
- Compute HMAC-SHA-256 over the unmodified bytes for every candidate.
- Compare every result using the
hmaccrate’s constant-time verification. - Return one generic
401result for a missing header, unknown selector, malformed signature, or signature mismatch. - Only after HMAC and any composed header factor pass, reserve the replay key.
Parsing a GitHub JSON or form payload happens only in the upstream service, after authentication. No UTF-8 conversion is required by the verifier.
Replay Store
Replay protection needs a purpose-specific contract rather than the current
observability-oriented RuntimeCache trait:
#![allow(unused)]
fn main() {
#[async_trait::async_trait]
pub trait WebhookReplayStore: Send + Sync {
async fn reserve(
&self,
key: &WebhookReplayKey,
retention: std::time::Duration,
) -> Result<ReserveOutcome, ReplayStoreError>;
async fn release(
&self,
reservation: &ReplayReservation,
) -> Result<(), ReplayStoreError>;
async fn force_remove(
&self,
key: &WebhookReplayKey,
) -> Result<bool, ReplayStoreError>;
}
pub enum ReserveOutcome {
Reserved(ReplayReservation),
Duplicate,
}
}
The logical key contains profile, normalized selector or shared, and replay
ID. Java and Rust use the same persistent digest so replay protection remains
effective during migration or active-active operation across runtimes. The
canonical input is the three UTF-8 values in that order, each preceded by its
unsigned four-byte big-endian byte length. The persistent identity is the
lowercase hexadecimal SHA-256 digest of that byte sequence. Providers prepend
only their configured namespace; they do not invent another tuple encoding.
The reservation stores a random owner token. Normal failure release uses atomic
compare-and-delete so an old request cannot delete a newer reservation after
expiry and re-reserve. The operator-only force_remove intentionally ignores
the owner token. Shared conformance fixtures cover empty-forbidden values,
shared, non-ASCII selectors, embedded separators, the final digest, and the
provider key prefix.
Local Store
The process-local implementation is used only when a replay-enabled profile
explicitly references a store with type: local. It must:
- implement atomic insert-if-absent under concurrency;
- expire entries after the requested retention;
- enforce a configured maximum without evicting unexpired entries silently;
- return an unavailable/capacity error, producing
503, if it cannot preserve the replay guarantee; and - register a safe summary with
CacheRegistryfor operational visibility.
Local scope is selected with a referenced provider, not by omitting the store:
profiles:
development-hook:
replay:
enabled: true
idHeader: X-Delivery
store: development-local
retentionSeconds: 604800
replayStores:
development-local:
type: local
maxEntries: 100000
Process restart loses local replay history, and multiple instances do not share
it. Startup logs and metrics must clearly report scope=local. A missing store,
an unknown reference, or an enabled replay policy without replay.store fails
startup or reload; it never falls back to local state.
Redis Store
The first distributed provider should be Redis. Reservation maps to SET key owner-token NX PX retention. Release uses an atomic compare-and-delete script,
and operator removal uses DEL.
The deployment should use a Redis namespace/database whose eviction policy does
not discard unexpired replay entries. Connection or operation timeout fails the
webhook with 503. Distributed storage is recommended whenever a route is
served by more than one gateway instance.
A future general Rust cache/session abstraction may host this provider, but its contract must preserve atomic reserve and compare-and-delete semantics. A CRUD session repository or get-then-put cache adapter is insufficient.
Retention and Upstream Outcome
The default retention is 604800 seconds (seven days). Operators may reduce or increase it per profile. Documentation must explain that finite retention means finite replay protection.
Outcome handling is:
| Outcome | Reservation behavior |
|---|---|
| Duplicate before upstream selection | Return empty 200; existing reservation remains |
Upstream 2xx response header received | Atomically transition Reserved to Committed2xx; keep until TTL even if later response processing or the downstream write fails |
Upstream non-2xx | Transition to Releasing and release before completing the downstream response |
| Local rejection after reservation | Release from the final async completion phase; this includes rate limiting, access control, validation, and other later handlers |
| Body reinjection or final proxy failure before an upstream response | Release from the final async completion phase |
| Retried upstream attempt | Keep the reservation between attempts; commit on the first observed 2xx, otherwise release only after the final failed attempt |
| Release-store failure | Log and count it; return the original upstream failure; operator can remove the stale record |
| Gateway crash after reserve | Reservation remains until TTL or operator removal |
The final callback releases any state still Reserved, even when Pingora reports
no error because a later handler produced a normal local response. Release is
idempotent, and only one task may perform the Reserved to Releasing
transition. Keeping Committed2xx after an observed upstream 2xx prevents a
downstream write failure from triggering the Jenkins build again. Ambiguous
failures still require callee idempotency.
Operator Redelivery
Expose a protected runtime MCP operation rather than asking operators to know a provider-native cache key:
{
"name": "remove_webhook_replay",
"arguments": {
"profile": "github",
"selector": "12345678",
"deliveryId": "6f3f8b40-..."
}
}
The controller adds and routes by runtimeInstanceId using its existing runtime
MCP path. The response reports whether an entry was removed and whether the
store scope is local or distributed; it never returns provider keys or
stored values.
For a local store, the operation must be executed on every gateway instance that can serve the route. For a distributed store, one removal is sufficient. The operation requires administrative authorization and emits an audit event. The expected workflow is remove first, then request GitHub redelivery.
Hot Reload
Replace the current independently loaded handler and authentication values with one generation-pinned execution snapshot:
#![allow(unused)]
fn main() {
pub struct HmacRuntime {
standalone_policy: CompiledStandaloneHmacPolicy,
profiles: std::collections::BTreeMap<String, HmacProfileRuntime>,
replay_stores: std::collections::BTreeMap<String, Arc<dyn WebhookReplayStore>>,
}
pub struct UnifiedSecurityRuntime {
policy: CompiledUnifiedSecurityPolicy,
hmac: Option<Arc<HmacRuntime>>,
}
pub struct GatewaySecurityExecutionSnapshot {
generation: u64,
active_handlers: Arc<ActiveHandlerSet>,
hmac: Option<Arc<HmacRuntime>>,
unified_security: Option<Arc<UnifiedSecurityRuntime>>,
api_key: Option<Arc<ApiKeyConfig>>,
basic_auth: Option<Arc<BasicAuthConfig>>,
security: Option<Arc<SecurityRuntime>>,
}
}
The coordinated security-execution reloader is triggered by changes to
handler.yml, hmac.yml, unified-security.yml, or authentication configuration
referenced by a composed policy. It loads the effective handler set and all
required authentication inputs, validates their cross-references, resolves
secret environment names, connects configured stores, and constructs one
candidate snapshot.
Validation covers every method-aware standalone and composed HMAC rule against
the runnable handler chain. A standalone rule requires exactly one enabled
hmac entry point. A composed rule requires exactly one enabled
unified-security entry point. A covered path and method may not execute both
standalone hmac and unified-security, regardless of whether the selected
Unified Security rule itself contains HMAC. Every chain containing standalone
hmac must be covered by a standalone rule for each path and method it serves,
and HMAC must precede any terminal application handler. Disabled modules are
evaluated as disabled behavior even when their IDs remain in the expanded
handler list.
ConfigManager swaps the candidate only after every step succeeds. The request
captures one Arc<GatewaySecurityExecutionSnapshot> before resolving its chain
and uses that same generation for handler selection and every authentication
factor. A failed reload leaves the previous complete snapshot active; the
implementation must not perform a series of observable per-module stores.
Module registry output exposes the public HMAC configuration with secret environment names masked or omitted. Resolved bytes are never registered.
Failure Contract
| Condition | HTTP status | Behavior |
|---|---|---|
| Missing/malformed signature, unknown selector, or mismatch | 401 | Generic invalid webhook authentication response |
| Missing/malformed replay ID when replay is enabled | 401 | Reject without reserving or forwarding |
Unsupported Content-Encoding | 415 | Reject before reading/forwarding body |
| Body exceeds configured maximum | 413 | Stop reading and reject |
| Body read timeout | 408 | Reject and close/drain according to Pingora safety rules |
| Global HMAC body-buffer budget exhausted | 503 | Fail closed before allocating or forwarding the full body |
| Duplicate replay ID | 200 | Empty local response; do not contact upstream |
| Configured replay store unavailable or full | 503 | Fail closed; do not contact upstream |
| Missing/mismatched handler entry point or runtime generation | 503 | Fail closed and record a chain/runtime error; startup and reload validation should normally prevent this |
| Unexpected verifier/runtime failure | 503 | Fail closed and record runtime_error, not store_unavailable |
Upstream non-2xx | Upstream status | Release reservation and forward response |
| Upstream connection/proxy failure | Existing gateway error | Release reservation when outcome is not a known 2xx |
Error responses must not reveal whether the selector existed, which secret matched, or whether active versus previous material was used.
Observability
Recommended metrics are:
hmac_webhook_requests_total{profile,outcome}where outcome isaccepted,duplicate,invalid,too_large,unsupported_encoding,timeout,buffer_unavailable,store_unavailable,chain_error, orruntime_error;hmac_webhook_verification_duration_seconds{profile};hmac_webhook_body_bytes{profile};hmac_replay_operations_total{store_type,operation,outcome}; andhmac_replay_local_entriesfor the local provider.
The gateway publishes these as structured light_pingora::metrics events with
cumulative counter values or observations. This is the runtime’s operational
metric-export surface; deployments may translate the events to their metrics
backend. They are not private in-process counters only.
Logs may include profile, route, correlation ID, store type, status, and a one-way truncated hash of the selector/delivery tuple when troubleshooting is required. Do not log the raw body, signature, secret, delivery ID, selector, JWT, or API key. Do not put selector or delivery ID into metric labels.
Startup and reload logs must state whether each profile uses local or
distributed replay protection. Explicit local scope is always logged as a
warning because the runtime cannot prove that the deployment has only one
instance. Unexpected exceptions include the error chain in logs without body,
credential, selector, or delivery values. store_unavailable is reserved for
replay-provider failures so store-health alerts remain actionable.
Validation Plan
Configuration
- Legacy Unified Security rules deserialize and behave unchanged.
- Standalone HMAC rules select the longest matching prefix and method and fail
closed when the
hmachandler has no matching rule. - New
allOfrules reject mixed legacy fields, unknown factors, empty factors, missing profiles, duplicate methods, and anonymous overlap. - Reject both a more-specific non-HMAC override of a broader HMAC rule and non-HMAC ancestor-method fallthrough around a more-specific HMAC rule, including standalone-HMAC versus legacy-Unified-Security pairs.
- Reject overlapping standalone and composed HMAC coverage.
- Reject missing, unknown, duplicate, or implicit replay-store selection; accept
explicit
type: localandtype: redisproviders. - Profile parsing accepts hex and base64 encoding and rejects unsupported signed input or algorithms.
- Missing/empty secret environment variables fail startup/reload without replacing the active runtime.
- In-flight requests continue using one old handler/security snapshot after reload; no request observes mixed generations.
Cryptography and Bytes
- Use GitHub’s published secret/payload/signature test vector.
- Validate non-ASCII payload bytes, whitespace changes, empty bodies, and bodies split across different downstream chunk boundaries.
- Prove that parsing or re-serializing the same logical JSON does not validate unless its bytes match the signature.
- Reject malformed prefixes, wrong decoded lengths, duplicate signature headers, invalid hex/base64, and wrong signatures with the same generic response.
- Verify both active and previous secrets without reporting which one matched.
Authentication Composition
- A standalone
hmacchain accepts a valid signed request without requiring Unified Security. - HMAC plus JWT requires both factors.
- HMAC plus API key requires both factors.
- A valid HMAC never compensates for a missing/invalid JWT or API key.
- A valid JWT/API key never compensates for invalid HMAC.
- HMAC routes cannot bypass through
anonymousPrefixes. - Longest-prefix and method matching select the expected Rust rule.
- Missing or disabled standalone
hmacand composedunified-securityentry points fail startup/reload. - A chain containing standalone
hmacandunified-securityfor the same path and method fails startup/reload, including when Unified Security selects a JWT-only or API-key-only rule. A defensive second runtime entry fails closed. - Every path and method served by a chain containing standalone
hmachas a matching standalone rule; uncovered default-chain traffic fails validation. - Chain validation uses effective module-enabled behavior rather than only the
raw
handler.ymlexeclist.
Body Gate
- Run the Phase 0 HTTP/1.1 and HTTP/2 integration proof.
- Validate exactly 16 MiB and reject 16 MiB plus one byte.
- Validate an incorrect or absent
Content-Lengthcannot bypass themaxBodyBytes + 1end-of-stream proof. - Exhaust the aggregate HMAC body budget and verify fail-closed recovery without leaking permits.
- Reject non-identity content encoding.
- Verify a fake upstream sees no request for invalid HMAC, duplicate delivery, replay-store failure, oversized body, or timeout.
- Verify a successful upstream receives the exact authenticated bytes and expected end-to-end headers.
- Exercise interaction with request tokenization and body-aware access control.
- Reject unsupported direct application-handler chains at startup.
Replay
- Race many reservations for one key; exactly one wins in both local and Redis providers.
- A duplicate returns
200and does not increment the upstream request count. - A
2xxretains the reservation. - A non-
2xxand a pre-response transport failure release it. - A local rate-limit/access-control rejection after reservation releases it.
- A body-reinjection failure and a later handler exception release it.
- A connect failure followed by a successful Pingora retry retains one reservation; final retry exhaustion releases it.
- A downstream disconnect after an upstream
2xxretains it. - Compare-and-delete cannot remove a newer owner’s reservation.
- Local capacity exhaustion and configured Redis outage fail closed.
- Operator removal allows a subsequent intentional redelivery.
- Controller fan-out removes local entries from every selected instance.
- Java and Rust produce the same replay-key digest for shared conformance vectors, including non-ASCII and embedded-separator inputs.
End-to-End GitHub/Jenkins
- Configure one GitHub hook ID with an active secret and trigger one Jenkins build from a GitHub webhook.
- Confirm the delivered body and headers match the authenticated request.
- Resend the same delivery and confirm
200with no second Jenkins invocation. - Make Jenkins return a failure, confirm release, and then redeliver successfully.
- Rotate from previous to active secret using module reload and verify the overlap window.
- No other service provider integration test is required for version 1.
Implementation Phases
Phase 0: Prove the Pre-Upstream Body Gate
- Complete — gateway-core pre-buffer hook selected and proven.
- The initial spike records the standard-hook 64 KiB limitation; the pinned proxy extension then proves exact 16 MiB replay for HTTP/1.1 and body replay above 64 KiB for HTTP/2.
- Local short-circuiting, content-length and chunked input, exact bytes, header preservation, and 16 MiB plus one-byte rejection pass the counting-upstream integration gate.
Phase 1: Configuration and Verification Core
- Complete.
frameworks/light-pingora/src/hmac.rsprovides profile/config parsing, off-path environment-secret resolution, and raw-body HMAC-SHA-256 verification. - The standalone
hmacdescriptor and method-aware longest-prefixhmac.yml.pathPrefixAuthspolicy are registered and compiled. - Unified Security accepts method-aware
authentication.allOfpolicies for HMAC-only, HMAC plus JWT, and HMAC plus API key, while legacy rules remain compatible. - Startup and reload compile one reusable HMAC runtime, validate referenced profiles and overlap/fallthrough rules, omit secret environment names from registered public configuration, and retain the previous runtime after a failed candidate reload.
- GitHub’s published vector, raw-byte sensitivity, active/previous/default secret selection, malformed/duplicate headers, policy selection, composition, overlap validation, and reload behavior have focused tests.
- Until Phase 3 connects bounded capture to the verifier and pre-buffer hook,
both HMAC entry points return a defensive
503rather than partially authenticating traffic.
Phase 2: Replay Stores and Administration
- Complete.
WebhookReplayStorenow has capacity-safe local and Redis providers, atomic reservation, owner-checked release, and operator-only removal. - Every replay-enabled profile selects an explicit named provider; missing, invalid, or unavailable configuration fails closed without a local fallback.
WebhookReplayKeyimplements the shared Java/Rust length-prefixed SHA-256 digest and provider namespaces expose neither logical input nor owner token.- Unchanged providers are reused across HMAC reloads so local replay history is not silently reset by a valid configuration refresh.
- Local providers publish redacted, read-only summaries through
CacheRegistry; generic bulk clearing is explicitly unsupported. - The controller runtime MCP surface discovers the protected
remove_webhook_replayoperation, which removes one logical key, reports local or distributed scope, and emits a redacted audit event.
Phase 3: Gateway Integration
- Complete.
GatewayProxynow consumes the bounded exact body before upstream selection, verifies it through either standalone or composed HMAC, and exposes only verified bytes through the Phase 0 pre-buffer hook. - Each request pins one
GatewaySecurityExecutionSnapshot; handler and direct authentication reloads publish complete generations without changing an in-flight request’s chain, factors, secrets, or replay-store reference. GatewayRequestContextowns the verified body, aggregate byte permit, entry point, profile, and one-way replay state. Duplicate reservations return an empty local200without upstream selection.- An observed upstream
2xxcommits the reservation. Non-2xxresponses and final local/proxy failure paths perform owner-checked release, while retries retain the same request-owned reservation and verified body. - Materialized-chain validation covers standalone and composed entry points, duplicate entry, uncovered standalone chains, and proxy/router ordering. Metrics and logs use redacted profile/outcome/store-scope dimensions only.
scripts/run-hmac-phase3-gates.shruns the earlier phase gates plus the full gateway suite and live exact-body, duplicate, and non-2xxrelease proof.
Phase 4: Qualification
- Local qualification complete; external release qualification pending.
scripts/run-hmac-phase4-gates.shruns the cumulative focusedlight-pingora,light-runtime, andlight-gatewaysuites, including the Phase 0 HTTP/1.1 and h2c body-hook counting-upstream matrix. That h2c spike is not a full HMAC request through the production TLS listener; HTTP/2 over TLS with the complete HMAC chain remains an external release gate. - A disposable Redis 7 instance qualifies atomic concurrent reservation from two independently connected provider objects in one test process, cross-connection duplicate visibility, owner-checked release, and stale-owner protection. Two concurrently running gateway processes against one Redis deployment remain an external release gate.
- Live gateway tests cover standalone and API-key-plus-HMAC composition,
generation-pinned reload, exact upstream bytes and GitHub event headers,
duplicate suppression, non-
2xxrelease, later local-router rejection, and final upstream connection failure followed by successful redelivery. fixtures/hmac-webhook-conformance-v1.jsonis the Rust version-1 conformance mirror of Java’s language-neutral fixture contract. It covers published and binary raw-request signatures plus length-prefixed replay keys containing non-ASCII and embedded-separator inputs; the repositories do not consume one physical fixture file.- The in-process counting HTTP upstream simulates the Jenkins lifecycle and proves one call for a successful GitHub delivery, no second build for its duplicate, release after a failed build, and successful redelivery. The secret-rotation exercise reloads active plus previous secrets atomically and proves that an older pinned generation keeps its original secret set.
- A deployed GitHub webhook reaching a real Jenkins target, including duplicate,
non-
2xxredelivery, and secret rotation, remains an external release gate.
Java Implementation Lessons and Cross-Runtime Conformance
The completed Java implementation keeps its maintenance-oriented architecture:
first-match prefix rules, request interception for exact-body verification, a
materialized handler-chain validator, service.yml replay-store dependency
injection, and a completion listener for replay release. Rust does not copy
those internals, but incorporates the correctness lessons:
- validate HMAC structure in always-loaded policy configuration so missing optional wiring cannot hide a protected rule;
- compare prefix overlaps in both directions and account for Rust’s longest-prefix and method-aware selection;
- validate the effective runnable chain, including disabled modules, rather than the raw configured chain;
- pin one immutable runtime generation per request and avoid repeated startup-grade validation on the request hot path;
- distinguish replay-store failures from generic runtime failures;
- release reservations from every final non-success path, including later local handler rejection; and
- require an explicit replay-store implementation when replay is enabled.
Java and Rust share raw request/signature fixtures, duplicate behavior, header
preservation, maximum-body boundary cases, the exact length-prefixed replay-key
digest, and the status cases both implementations expose: invalid 401, body
limit 413, unsupported encoding 415, fail-closed 503, and duplicate empty
200. Rust-only lifecycle controls such as body-read timeout 408 and aggregate
buffer-budget exhaustion are not cross-runtime status fixtures. The runtimes are
also not required to share prefix selection, request lifecycle,
dependency-injection mechanism, or internal type model.
PII Tokenization
Status
Proposed design for migrating the light-tokenization capability into
light-fabric as light-pingora handlers used by light-gateway.
Purpose
PII tokenization protects sensitive employee/customer data when a request is sent from inside the organization to an external cloud service through the gateway. The outbound request replaces configured PII fields with generated tokens. When the cloud response returns, the gateway replaces those tokens with the original cleartext values so internal employees can complete their work.
This is a request/response hot-path concern. The first Rust implementation
should therefore run inside light-gateway and access PostgreSQL directly
instead of making a network call to a tokenization service for every field.
Current Java Behavior
The current light-tokenization service exposes REST endpoints:
POST /v1/token: body{ "schemeId": <int>, "value": "<cleartext>" }; returns a token string. If the value already exists, it returns the existing token.GET /v1/token/{token}: returns the cleartext value.DELETE /v1/token/{token}: deletes the token mapping.GET /v1/schemeandGET /v1/scheme/{id}: return token format schemes.
Startup loads multiple JDBC pools from datasource.yml. One database is named
tokenization; the others are vault databases such as vault000. The
tokenization database maps client_id to a vault database through
client_database. Each vault database has a token_vault table.
Java tokenization flow:
- Read
client_idfrom the JWT audit info. - Resolve
client_id -> db_name. - Select a vault datasource by
db_name. - For tokenization, look up by cleartext
value; return existingidif found. - If not found, generate a token with the configured
schemeId, insert(id, value), cachetoken -> value, and return the token. - For detokenization, check the cache first, then query by token
id.
The current Java MCP router also uses tokenization through token-client.
Tool input schemas can mark fields with x-tokenize; the router extracts
JsonPath rules from the schema and calls the tokenization service.
Design Direction
Use direct PostgreSQL access for the initial light-fabric implementation.
Reasons:
- It removes one HTTP hop per tokenized field in the gateway hot path.
- It avoids running and scaling another service only to perform local database lookups.
- PostgreSQL connection pooling is already used in nearby light-fabric apps
with
sqlx. - The same database will also support other gateway handlers that need local data access, such as vector search for MCP routing.
- Multi-tenancy is cleaner with
host_idin the schema than with one vault database per tenant.
If this capability is later exposed as a standalone service, prefer gRPC over MCP for the hot-path service API. gRPC gives a strongly typed protobuf contract, HTTP/2 multiplexing, compact binary payloads, deadlines, and well-understood client pooling. MCP is useful when tokenization is exposed as an agent tool or administrative capability, but it adds JSON-RPC/tooling semantics that are not needed for a low-latency service-to-service data-plane call.
Goals
- Implement
TokenizeHandlerandDetokenizeHandlerinlight-pingora. - Activate handlers only through
handler.yml. - Use one PostgreSQL database with
host_idtenant isolation. - Integrate schema into
portal-db/postgres/ddl.sqland future patch files. - Preserve the Java token schemes and stable tokenization behavior.
- Avoid storing/indexing cleartext PII directly in PostgreSQL.
- Support request-body tokenization before proxy/router sends to the external service.
- Support response-body detokenization before the gateway returns to the internal caller.
- Reuse the same runtime for MCP tool argument tokenization.
Non-Goals
- Do not preserve multiple vault databases.
- Do not preserve MySQL or SQLite runtime support in light-fabric.
- Do not make tokenization an MCP-only service.
- Do not require a separate tokenization service for the first implementation.
- Do not try to tokenize arbitrary binary payloads in the first pass.
Handler Model
Use two public handler ids:
tokenize: request-phase handler that replaces cleartext fields with tokens.detokenize: response-phase handler that replaces configured token fields with cleartext.
Both handlers share one runtime:
frameworks/light-pingora/src/pii_tokenization.rs
Primary types:
#![allow(unused)]
fn main() {
pub struct PiiTokenizationConfig {
pub database: PiiDatabaseConfig,
pub host_id_claim: String,
pub max_body_size: usize,
pub cache: PiiTokenCacheConfig,
pub crypto: PiiTokenCryptoConfig,
pub rules: Vec<PiiTokenizationRule>,
}
pub struct PiiTokenizationRuntime {
pub config: Arc<PiiTokenizationConfig>,
pub pool: PgPool,
pub tokenizers: TokenizerRegistry,
pub value_cache: TokenCache,
pub token_cache: TokenCache,
pub keyring: PiiKeyring,
}
pub struct PiiTokenizationRule {
pub path_prefix: String,
pub methods: Vec<String>,
pub request: Vec<PiiFieldRule>,
pub response: Vec<PiiFieldRule>,
}
pub struct PiiFieldRule {
pub path: String,
pub scheme: String,
pub required: bool,
}
}
The handler should fail startup if an active config references an unknown scheme, has invalid field paths, cannot initialize the keyring, or cannot connect to PostgreSQL within the configured startup timeout.
Resolved Decisions
- Handler ids are
tokenizeanddetokenizeto align with other light-fabric handler names. - Encrypt stored cleartext with AES-256-GCM. Resolve key material from environment variables first, with direct config values allowed only as a local-development fallback.
- Detokenization fails closed by default when a configured token field cannot be resolved.
- Field selection uses a constrained compiled JsonPath subset rather than full dynamic JsonPath evaluation.
- Cleartext reverse caching is configurable through
cache.cacheCleartext. - Request/response mutation buffers are bounded by configurable
maxBodySize.
Handler Chain
For a BFF or gateway that calls an external cloud service:
handlers:
- correlation
- security
- tokenize
- router
- detokenize
chains:
external-cloud:
- correlation
- security
- tokenize
- router
- detokenize
paths:
- path: /claims
method: POST
exec:
- external-cloud
tokenize must run after authentication so it can resolve host_id from the
verified JWT principal. It must run before router or proxy so the external
service never receives cleartext PII. detokenize must run after the upstream
response body is available and before response delivery.
This likely requires extending the existing gateway handler model with a response-body filter phase:
#![allow(unused)]
fn main() {
pub trait PingoraBodyHandler {
async fn request_body_filter(&self, ctx: &mut GatewayRequestContext, body: Bytes)
-> Result<Bytes, HandlerRejection>;
async fn response_body_filter(&self, ctx: &mut GatewayRequestContext, body: Bytes)
-> Result<Bytes, HandlerRejection>;
}
}
The first implementation can wire this directly in light-gateway; later it
can be generalized for other body-mutating handlers.
Configuration
Primary file: pii-tokenization.yml.
enabled is not needed. If neither tokenize nor detokenize appears in
handler.yml, this config is not loaded. If either handler is active, the
config is required and invalid config fails startup.
Example:
database:
url: ${pii-tokenization.database.url:${database.url:}}
maxConnections: ${pii-tokenization.database.maxConnections:8}
minConnections: ${pii-tokenization.database.minConnections:1}
connectTimeoutMs: ${pii-tokenization.database.connectTimeoutMs:2000}
hostIdClaim: ${pii-tokenization.hostIdClaim:host_id}
maxBodySize: ${pii-tokenization.maxBodySize:1048576}
crypto:
algorithm: ${pii-tokenization.crypto.algorithm:AES-256-GCM}
keyId: ${pii-tokenization.crypto.keyId:default}
valueEncryptionKeyEnv: ${pii-tokenization.crypto.valueEncryptionKeyEnv:PII_TOKENIZATION_VALUE_ENCRYPTION_KEY}
valueHashKeyEnv: ${pii-tokenization.crypto.valueHashKeyEnv:PII_TOKENIZATION_VALUE_HASH_KEY}
valueEncryptionKey: ${pii-tokenization.crypto.valueEncryptionKey:}
valueHashKey: ${pii-tokenization.crypto.valueHashKey:}
cache:
enabled: ${pii-tokenization.cache.enabled:true}
maxEntries: ${pii-tokenization.cache.maxEntries:10000}
ttlSeconds: ${pii-tokenization.cache.ttlSeconds:86400}
cacheCleartext: ${pii-tokenization.cache.cacheCleartext:true}
rules:
- pathPrefix: /claims
methods: [POST]
request:
- path: $.claimant.ssn
scheme: LN
required: false
- path: $.payment.cardNumber
scheme: CC4
required: false
response:
- path: $.claimant.ssn
scheme: LN
required: false
- path: $.payment.cardNumber
scheme: CC4
required: false
Field paths should support the Java-compatible JsonPath subset used by
mcp-router tokenization rules: object fields and [*] arrays. For
performance and predictable mutation, the Rust implementation should compile
rules at startup and avoid dynamic path parsing on every request.
For MCP tools, keep supporting x-tokenize in input schemas. The MCP router
can convert schema annotations into the same compiled field rules and call the
shared PiiTokenizationRuntime directly.
PostgreSQL Schema
Replace the old split between tokenization and vault databases with
tenant-scoped tables in portal-db.
Recommended DDL:
CREATE TABLE pii_token_scheme_t (
scheme_id SMALLINT PRIMARY KEY,
scheme_code VARCHAR(16) NOT NULL UNIQUE,
description TEXT NOT NULL,
active BOOLEAN DEFAULT TRUE NOT NULL,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
update_user VARCHAR(126) DEFAULT SESSION_USER NOT NULL
);
CREATE TABLE pii_token_vault_t (
host_id UUID NOT NULL,
token TEXT NOT NULL,
scheme_id SMALLINT NOT NULL,
value_hash BYTEA NOT NULL,
value_ciphertext BYTEA NOT NULL,
value_nonce BYTEA NOT NULL,
key_id VARCHAR(128) NOT NULL,
created_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
expires_ts TIMESTAMP WITH TIME ZONE,
active BOOLEAN DEFAULT TRUE NOT NULL,
update_ts TIMESTAMP WITH TIME ZONE DEFAULT CURRENT_TIMESTAMP NOT NULL,
update_user VARCHAR(126) DEFAULT SESSION_USER NOT NULL,
PRIMARY KEY(host_id, token),
FOREIGN KEY(scheme_id) REFERENCES pii_token_scheme_t(scheme_id)
);
CREATE UNIQUE INDEX pii_token_vault_value_uk
ON pii_token_vault_t(host_id, scheme_id, value_hash)
WHERE active = TRUE;
CREATE INDEX pii_token_vault_expiry_idx
ON pii_token_vault_t(expires_ts)
WHERE expires_ts IS NOT NULL;
Seed schemes:
| Id | Code | Meaning |
|---|---|---|
0 | UUID | UUID v4 token |
1 | GUID | URL-safe base64 UUID token |
2 | LN | Luhn compliant numeric token |
3 | N | Random numeric token, length preserving |
4 | LN4 | Luhn numeric token retaining last four digits |
5 | AN | Random alpha-numeric token, length preserving |
6 | AN4 | Alpha-numeric token retaining last four characters |
7 | CC | Credit-card-shaped Luhn token retaining first digit |
8 | CC4 | Credit-card-shaped Luhn token retaining first and last four digits |
The old database_owner and client_database tables are not needed. Tenant
isolation is by host_id, resolved from the authenticated request. If a
legacy client only has client_id, handle that with normal portal auth/client
metadata rather than recreating tokenization-specific vault routing.
Cleartext Storage
The Java schema stores cleartext PII in token_vault.value and indexes it.
The Rust schema should not.
Use:
value_hash: deterministic HMAC-SHA-256 of(host_id, scheme_id, canonical_value)withvalueHashKey; used for idempotent token lookup.value_ciphertextandvalue_nonce: encrypted cleartext value, for example AES-GCM or ChaCha20-Poly1305 withvalueEncryptionKey.key_id: identifies which key encrypted the row so key rotation is possible.
This keeps tokenization idempotent without indexing cleartext PII.
Tokenization Algorithm
Shared runtime operation:
tokenize(host_id, scheme_id, value)
-> canonicalize value
-> compute value_hash
-> cache lookup by (host_id, scheme_id, value_hash)
-> SELECT token WHERE host_id, scheme_id, value_hash, active
-> if found, cache and return
-> generate scheme-specific token
-> encrypt cleartext
-> INSERT row
-> on token collision, retry generation
-> on value_hash conflict, SELECT existing token and return it
Use PostgreSQL uniqueness instead of application locks:
INSERT INTO pii_token_vault_t (...)
VALUES (...)
ON CONFLICT DO NOTHING;
If no row is inserted, determine whether the conflict was on
(host_id, token) or (host_id, scheme_id, value_hash). Token collision means
retry with a new token. Value conflict means another request already inserted
the mapping; select and return the existing token.
Detokenization:
detokenize(host_id, token)
-> cache lookup by (host_id, token)
-> SELECT encrypted value WHERE host_id, token, active
-> decrypt cleartext
-> cache and return
If token is not found, the handler fails the response with a handler error. For gateway response-body detokenization, fail closed so employees do not see partial or incorrect data without a signal.
Runtime Caching
Use bounded in-process caches:
(host_id, scheme_id, value_hash) -> token(host_id, token) -> cleartext
The cache must be tenant-scoped and bounded by count and TTL. Because the reverse cache contains cleartext PII, make it configurable and register it with the runtime cache registry only with masked summaries. A clear-cache operation should be available through the runtime control plane.
The cache is an optimization only. PostgreSQL remains the source of truth.
Request And Response Mutation
Only mutate supported structured content:
application/jsonin phase 1.- JSON arrays and nested objects through compiled path rules.
- Missing optional fields are ignored.
- Missing required fields reject the request or response with a handler error.
For outbound request tokenization:
- Buffer the JSON request body within a configured max body size.
- Parse to
serde_json::Value. - Apply matching request rules.
- Replace every string value with a token.
- Serialize JSON, update
Content-Length, and forward upstream.
For inbound response detokenization:
- Buffer the JSON response body within a configured max body size.
- Parse to
serde_json::Value. - Apply matching response rules.
- Replace every string token with cleartext.
- Serialize JSON, update
Content-Length, and return downstream.
For very large or streaming payloads, skip mutation and fail closed by default. Streaming tokenization can be considered later only if a real product requires it.
Security
- Require a verified JWT principal before tokenization.
- Resolve
host_idfrom a configured claim, defaulthost_id. - Reject active tokenization if
host_idis missing. - Do not log cleartext values, generated tokens, value hashes, ciphertext, or keys.
- Mask crypto keys in module registry summaries.
- Use least-privilege PostgreSQL credentials: only select/insert/update on the tokenization tables.
- Prefer encrypted cleartext storage, not plaintext
value. - Keep tokens scoped by
host_id; the same token string in another tenant does not detokenize.
Future Service API
The direct database implementation should be the first production path. However, keep the core API independent from Pingora:
#![allow(unused)]
fn main() {
#[async_trait]
pub trait PiiTokenVault: Send + Sync {
async fn tokenize(&self, host_id: Uuid, scheme_id: i16, value: &str)
-> Result<String, PiiTokenError>;
async fn detokenize(&self, host_id: Uuid, token: &str)
-> Result<String, PiiTokenError>;
}
}
Then a future service can wrap the same trait.
Protocol recommendation:
- gRPC for request-path service-to-service tokenization if a standalone service becomes necessary.
- MCP only as an optional tool surface for agents or administrative workflows.
- REST/JSON-RPC only for compatibility or operational simplicity, not the preferred low-latency path.
The gRPC API can be very small:
service PiiTokenization {
rpc Tokenize(TokenizeRequest) returns (TokenizeResponse);
rpc Detokenize(DetokenizeRequest) returns (DetokenizeResponse);
rpc BatchTokenize(BatchTokenizeRequest) returns (BatchTokenizeResponse);
rpc BatchDetokenize(BatchDetokenizeRequest) returns (BatchDetokenizeResponse);
}
Batch operations are important if a future remote service is used; otherwise per-field network calls will dominate latency.
Implementation Phases
- Add portal-db DDL and seed data for
pii_token_scheme_tandpii_token_vault_t. - Add a
light-pingorashared tokenization runtime withsqlx::PgPool, scheme registry, value hashing, encryption, token generation, and tests. - Add
pii-tokenization.ymlloader, module registry registration, and runtime reload. - Add gateway request-body and response-body filter support.
- Implement
tokenizeanddetokenizehandler wiring inlight-gateway. - Integrate MCP
x-tokenizewith the same runtime so MCP tools do not call a hardcoded tokenization service. - Add optional gRPC service wrapper only if deployment needs a separate tokenization service.
Remaining Considerations
- KMS or light-portal managed keys can be added later, but the first implementation should read the configured environment variables before any resolved config fallback.
- Products that disable
cache.cacheCleartextwill still use PostgreSQL as the source of truth, with higher detokenization latency.
Token Handler
Status
Proposed design for migrating the Java egress-router TokenHandler into
light-fabric as the token handler used by light-pingora and
light-gateway.
A baseline Rust token runtime already exists in light-pingora. This document
captures the Java behavior, the compatibility contract, and the design direction
for hardening it for gateway and sidecar deployments.
Purpose
The token handler obtains an OAuth 2.0 client credentials access token on behalf
of the backend service in the sidecar or gateway egress path. The token is then
attached to the outbound request before router or proxy sends the request to
the downstream API.
This is different from the PII tokenize and detokenize handlers. The
token handler deals only with service-to-service OAuth tokens.
Java Behavior To Map
The Java implementation is centered on:
egress-router/.../TokenHandler.javasidecar/.../SidecarTokenHandler.javarouter-config/.../TokenConfig.javaclient-config/.../client.yamlsidecar-config/.../sidecar.yml
token.yml controls whether the handler is active and which request paths need
token injection:
enabled: ${token.enabled:false}
appliedPathPrefixes: ${token.appliedPathPrefixes:}
The OAuth provider, client credentials, cache, timeout, proxy, HTTP/2, and
single-vs-multiple-auth-server settings live in client.yml:
oauth:
multipleAuthServers: ${client.multipleAuthServers:false}
token:
cache:
capacity: ${client.tokenCacheCapacity:200}
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
expiredRefreshRetryDelay: ${client.expiredRefreshRetryDelay:2000}
earlyRefreshRetryDelay: ${client.earlyRefreshRetryDelay:30000}
server_url: ${client.tokenServerUrl:}
serviceId: ${client.tokenServiceId:}
proxyHost: ${client.tokenProxyHost:}
proxyPort: ${client.tokenProxyPort:}
enableHttp2: ${client.tokenEnableHttp2:true}
client_credentials:
uri: ${client.tokenCcUri:/oauth2/token}
client_id: ${client.tokenCcClientId:}
client_secret: ${client.tokenCcClientSecret:}
scope: ${client.tokenCcScope:}
serviceIdAuthServers: ${client.tokenCcServiceIdAuthServers:}
pathPrefixServices: ${client.pathPrefixServices:}
request:
connectTimeout: ${client.connectTimeout:2000}
timeout: ${client.timeout:4000}
The Java request flow is:
- Reload
token.ymlfor the request. - Check
appliedPathPrefixeswith a string prefix match. - Read
service_idfrom the request. This header is expected to be set byPathPrefixServiceHandlerorServiceDictHandler. - Resolve the auth server configuration from
client.yml. - Get or refresh a cached client credentials JWT for the service.
- If the request has no
Authorizationheader, setAuthorization: Bearer <token>. - If the request already has
Authorization, preserve it and setX-Scope-Token: Bearer <token>. - Continue to the next handler, usually
router.
For multiple auth servers, Java reads
oauth.token.client_credentials.serviceIdAuthServers[service_id] and enriches
that entry with the global token defaults. For a single auth server, it uses the
global oauth.token.client_credentials section.
The Java cache is a static map keyed by service_id. The cached Jwt stores
the access token and its exp claim in milliseconds. OauthHelper refreshes
synchronously after expiry and attempts async refresh while the token is in the
renewal window.
SidecarTokenHandler adds an egress gate before calling TokenHandler:
sidecar.egressIngressIndicator: headerruns the token handler only when the request hasservice_idorservice_url.sidecar.egressIngressIndicator: protocolruns the token handler for HTTP requests, which is the usual in-pod sidecar egress protocol.- Any other value skips token injection.
The base Java TokenHandler still needs service_id to choose the service
token. A request with only service_url can identify egress traffic, but it
does not by itself select a service-specific token.
Goals
- Preserve the Java configuration files:
token.ymlandclient.yml. - Activate the handler with the existing
tokenid inhandler.yml. - Support config-server injection for
token.enabled,token.appliedPathPrefixes,client.multipleAuthServers,client.tokenCcServiceIdAuthServers,sidecar.egressIngressIndicator, and the rest of theclient.ymltoken fields. - Support single auth server and per-service auth server configurations.
- Support token endpoint discovery through
oauth.token.serviceIdwhen a directserver_urlis not configured. - Preserve the Java header behavior for
AuthorizationandX-Scope-Token. - Keep token retrieval fast and safe for request-path execution.
- Register configuration and token cache state with the module registry and runtime cache registry without exposing token or secret values.
- Keep the design usable by
light-gateway, future sidecar products, and BFF deployments that need to call downstream APIs.
Non-Goals
- Do not use
inventoryor dynamic plugins. Handler availability is compiled into the binary; handler activation is controlled byhandler.yml. - Do not implement authorization code, refresh token, or token exchange in this
handler. This handler only performs
client_credentials. - Do not migrate Java
SAMLTokenHandleras part of this design. - Do not use the PII tokenization table or handlers.
token,tokenize, anddetokenizeare separate concerns. - Do not send the generated access token to logs, metrics labels, module registry output, or cache summaries.
Resolved Decisions
- Use
sidecar.ymlto differentiate inbound proxy traffic from outbound router traffic before applying token injection. - Implement refresh with the same concurrency model as Java
http-client: synchronize refresh per cached token, refresh expired tokens synchronously, refresh valid tokens in the renewal window asynchronously, and use retry windows to prevent repeated failed refresh attempts.
Handler Chain
The token handler must run after service resolution and before egress routing:
handlers:
- correlation
- security
- path-prefix-service
- token
- router
chains:
sidecar-egress:
- correlation
- security
- path-prefix-service
- token
- router
paths:
- path: /v1/pets
method: GET
exec:
- sidecar-egress
path-prefix-service sets service_id from path configuration. token uses
that service id to resolve and cache the client credentials token. router
uses the same service id to select the downstream API target and should remove
routing-only headers before forwarding.
For products where only some outbound APIs need a scope token, keep one chain
with token and another without it, or use token.appliedPathPrefixes to
limit token injection inside a shared chain.
Rust Architecture
Keep the implementation in light-pingora because token injection is a
request-path gateway handler. light-gateway wires the handler into the
existing chain execution model.
Primary Rust module:
frameworks/light-pingora/src/token.rs
Primary types:
#![allow(unused)]
fn main() {
pub struct TokenHandlerConfig {
pub enabled: bool,
pub applied_path_prefixes: Vec<String>,
}
pub struct ClientTokenConfig {
pub tls: ClientTlsConfig,
pub oauth: ClientOauthConfig,
pub path_prefix_services: BTreeMap<String, String>,
pub request: ClientRequestConfig,
}
pub struct TokenRuntime {
handler: TokenHandlerConfig,
sidecar: SidecarTrafficConfig,
client: ClientTokenConfig,
cache: Arc<TokenCache>,
registry_client: Option<Arc<PortalRegistryClient>>,
}
}
apps/light-gateway should load TokenRuntime only when the matched handler
configuration contains token. For Java compatibility, token.yml still has
enabled; therefore the handler is effective only when both conditions are
true:
handler.yml contains token
token.yml enabled is true
If token.yml enables the handler, client.yml is required and invalid
configuration should fail startup. sidecar.yml is also loaded into the token
runtime so the same handler chain can distinguish inbound proxy requests from
outbound router requests. Invalid reloads should be rejected while the last
valid runtime keeps serving traffic.
Request Flow
The Rust request flow should be:
- Resolve the active handler chain for the path and method.
- When
tokenis encountered, checkTokenHandlerConfig.enabled. - Evaluate
sidecar.ymland skip token injection for inbound proxy traffic. - Check
appliedPathPrefixeswith boundary-aware matching./v1/addressshould match/v1/address/123, but not/v1/address2. - Resolve the token service id:
- first from the
service_idrequest header, - then from
client.yml pathPrefixServices, - then from
oauth.token.serviceIdfor single-auth-server token endpoint discovery when applicable.
- first from the
- Resolve the token endpoint:
- use direct
server_urlfirst, - otherwise discover
oauth.token.serviceIdthrough portal registry.
- use direct
- Select client credentials:
- for single auth server, use
oauth.token.client_credentials, - for multiple auth servers, require
client_credentials.serviceIdAuthServers[service_id]and merge it with global token defaults.
- for single auth server, use
- Look up the token cache.
- Fetch a new token when the cache is missing, expired, or inside the refresh window.
- Add
AuthorizationorX-Scope-Tokenusing the Java-compatible rule.
The outbound token request should be Java-compatible:
POST {server_url}{uri}
Content-Type: application/x-www-form-urlencoded
Accept: application/json
Authorization: Basic base64(client_id:client_secret)
grant_type=client_credentials&scope=...
The response must contain access_token. Expiry should be derived from the JWT
exp claim when available, with expires_in as a fallback for non-JWT token
servers.
Cache And Refresh
Use a bounded async cache owned by TokenRuntime.
The cache key should include both service id and scope:
#![allow(unused)]
fn main() {
pub struct TokenCacheKey {
pub service_id: Option<String>,
pub scope: Option<String>,
}
}
This is stricter than the Java Map<String, Jwt> keyed only by service_id
and avoids collisions when the same service uses multiple scope sets.
Refresh policy:
- If the token is valid and outside the renewal window, use the cached token.
- If the token is expired, synchronize on that cache entry and refresh synchronously. Concurrent requests for the same service and scope should wait on the same per-entry lock, then re-check the refreshed token instead of making duplicate token endpoint calls.
- If the token is expired but another failed refresh attempt is still inside
expiredRefreshRetryDelay, fail closed with a token-not-available rejection. - If the token is in the renewal window but not expired, return the current
token and start one background refresh for that cache entry when no refresh is
already running and
earlyRefreshRetryDelayhas elapsed. - Keep refresh state per cached token: token string, expiry, scope,
renewing,expired_retry_timeout, andearly_retry_timeout.
This intentionally mirrors Java OauthHelper.populateCCToken. The Rust
implementation should use tokio locks/tasks instead of Java synchronized
and ScheduledExecutorService, but the observable behavior should stay the
same: expired tokens block the current request, early refresh does not block the
current request, and multiple concurrent requests for the same token are
coordinated through one cache entry.
On token.yml or client.yml reload, build a new TokenRuntime and discard
the old cache. This prevents tokens issued with old client credentials or old
scope configuration from being reused after a config change.
Sidecar Egress Gate
The token handler must use sidecar.yml to decide whether the current request
is outbound router traffic or inbound proxy traffic. This allows one gateway or
sidecar process to host both directions while applying token injection only to
egress calls.
Use the Java sidecar.yml contract:
egressIngressIndicator: ${sidecar.egressIngressIndicator:header}
Rust behavior:
header: runtokenonly whenservice_idorservice_urlis present.protocol: runtokenfor HTTP requests entering the sidecar listener.- any other value: skip token injection.
Even with this gate, token selection should still require either a resolved
service id or a single-auth-server configuration that can use a direct
server_url.
The sidecar config should be registered in the module registry as a framework config. Invalid values should fail startup or reject reload.
Configuration Examples
Single auth server:
# sidecar.yml
egressIngressIndicator: ${sidecar.egressIngressIndicator:header}
# token.yml
enabled: ${token.enabled:true}
appliedPathPrefixes: ${token.appliedPathPrefixes:/v1}
# client.yml
oauth:
multipleAuthServers: false
token:
server_url: ${client.tokenServerUrl:https://oauth.example.com}
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
client_credentials:
uri: ${client.tokenCcUri:/oauth2/token}
client_id: ${client.tokenCcClientId:gateway-client}
client_secret: ${client.tokenCcClientSecret:}
scope: ${client.tokenCcScope:petstore.r petstore.w}
Multiple auth servers:
# client.yml
oauth:
multipleAuthServers: true
token:
tokenRenewBeforeExpired: ${client.tokenRenewBeforeExpired:60000}
client_credentials:
uri: /oauth2/token
serviceIdAuthServers: ${client.tokenCcServiceIdAuthServers:}
pathPrefixServices: ${client.pathPrefixServices:}
The config server can inject client.tokenCcServiceIdAuthServers as YAML or a
JSON string:
com.networknt.petstore-1.0.0:
server_url: https://oauth-petstore.example.com
client_id: petstore-client
client_secret: ${PETSTORE_CLIENT_SECRET}
scope:
- petstore.r
- petstore.w
Rust Improvements Over Java
- Use boundary-aware path prefix matching instead of raw
startsWith. - Include scope in the cache key.
- Mask
client_secretand token values in module registry and cache output. - Fail startup for enabled but invalid token configuration.
- Use Rust async primitives to implement the same per-token synchronized refresh behavior as Java without spawning a dedicated executor per refresh attempt.
- Support direct
server_urland portal-registry discovery with the same runtime path. - Keep all config-server injected values in the normal module registry and reload model.
Observability
Record metrics and logs around the token operation, but never include the token or client secret:
- handler duration for
token, - cache hit, miss, refresh, and failure counts,
- token endpoint latency and HTTP status,
- service id and provider selection,
- refresh retry suppression counts,
- module registry entry for loaded
token.ymland maskedclient.yml, - runtime cache entry count and expiry summaries without access token strings.
Failure Behavior
Fail closed when token injection is required but cannot be completed:
- missing
service_idfor multiple auth servers, - missing
serviceIdAuthServers[service_id], - missing
client_idorclient_secret, - no direct
server_urland failed token service discovery, - token endpoint returns non-2xx,
- token response has no
access_token, - token response has neither JWT
expnorexpires_in, - invalid proxy, URL, or TLS configuration.
Requests outside appliedPathPrefixes should bypass the handler without error.
Test Plan
Unit tests in light-pingora:
- parse Java-compatible
token.ymlandclient.yml, - parse and validate Java-compatible
sidecar.yml, - parse
appliedPathPrefixesas YAML list, JSON string list, and comma list, - parse
serviceIdAuthServersas YAML map and JSON string map, - verify boundary-aware prefix matching,
- verify
sidecar.ymlheader mode applies token only to outbound requests withservice_idorservice_url, - verify
sidecar.ymlprotocol mode applies token to HTTP egress traffic, - verify single auth server option resolution,
- verify multiple auth server option merging,
- verify
AuthorizationversusX-Scope-Tokenheader selection, - verify cache key includes service id and scope,
- verify token cache summaries never include token strings,
- verify expired token refresh is synchronized across concurrent requests,
- verify early-window refresh returns the current token and starts only one background refresh.
Gateway tests in light-gateway:
- chain with
path-prefix-service -> token -> router, - inbound proxy request skips token injection according to
sidecar.yml, - outbound router request applies token injection according to
sidecar.yml, - missing service id for multiple auth servers returns a handler rejection,
- existing caller
Authorizationis preserved and scope token is added toX-Scope-Token, - token runtime reload swaps config and clears old cache,
- inactive
tokenhandler does not requiretoken.ymlorclient.yml.
Integration tests:
- mock OAuth token endpoint with client credentials Basic auth,
- mock discovered token service through portal registry,
- mock downstream service and assert the final outbound headers,
- refresh behavior with expired and near-expiry tokens.
Service Discovery
Status
Implemented baseline.
light-runtime, portal-registry, light-pingora, and light-gateway
already have the main pieces needed for controller-backed service discovery.
This document captures the supported invocation path, the configuration
contract, and the intended hardening direction for gateway, sidecar, BFF, MCP,
WebSocket, and token-handler deployments.
Purpose
light-gateway should be able to discover downstream service instances from
the Light Controller through portal-registry instead of relying only on static
host lists in router.yml, proxy.yml, mcp-router.yml, or handler-specific
configuration.
The same mechanism should work with both controller implementations:
- Rust
controller-rs - Java
light-controller
The gateway should use one portal-registry connection for registration,
runtime control-plane callbacks, and service discovery lookup. A separate
discovery client connection is not required for a registered runtime.
Goals
- Reuse the existing
portal-registryJSON-RPC WebSocket client. - Keep service discovery available to all
light-pingorahandlers throughRuntimeConfig.registry_client. - Support controller-backed lookup for:
- REST/router outbound calls
- WebSocket routing
- MCP tool routing
- OAuth token-server resolution
- SPA auth token-server resolution
- Keep direct URL configuration as an explicit override when a handler supports it.
- Keep static target configuration as a fallback where it already exists.
- Preserve Java-compatible discovery data names such as
serviceId,envTag,protocol,address, andport. - Let
light-portaland config-server manage product-specific registry and handler configuration. - Work with one
light-gatewaybinary and different product config sets.
Non-Goals
- Do not add a second discovery protocol for
light-gateway. - Do not require dynamic Rust plugins,
inventory, or reflection for discovery. - Do not make each handler own a separate controller connection.
- Do not require
/ws/discoveryfor registered gateway instances. - Do not remove static fallback configuration from router-style deployments.
- Do not make service discovery hide invalid product configuration. Startup validation and runtime errors should remain explicit.
Controller Protocol
The controller exposes two WebSocket endpoints:
/ws/microservice
/ws/discovery
light-gateway uses /ws/microservice.
The flow is:
light-gateway
-> connect /ws/microservice
-> JSON-RPC service/register
<- registered runtimeInstanceId
-> JSON-RPC discovery/lookup on the same websocket
<- DiscoverySnapshot
The dedicated /ws/discovery endpoint is still useful for non-service clients
that only need discovery. It is not needed by the gateway because both
controller-rs and light-controller accept discovery JSON-RPC methods on the
registered microservice socket after service/register succeeds.
The lookup request uses a DiscoverySubscription shape:
{
"serviceId": "com.networknt.petstore-1.0.0",
"envTag": "dev",
"protocol": "https"
}
envTag and protocol are optional. When protocol is omitted, the
controller can return all matching protocols and the caller decides which nodes
are usable.
The response is a DiscoverySnapshot:
{
"serviceId": "com.networknt.petstore-1.0.0",
"envTag": "dev",
"protocol": "https",
"nodes": [
{
"runtimeInstanceId": "...",
"serviceId": "com.networknt.petstore-1.0.0",
"envTag": "dev",
"environment": "dev",
"version": "1.0.0",
"protocol": "https",
"address": "petstore",
"port": 8443,
"tags": {},
"connectedAt": "...",
"lastSeenAt": "...",
"connected": true
}
]
}
Only connected nodes with a non-zero port should be used as upstream targets. Handlers should ignore protocols they cannot proxy.
Runtime Configuration
Registry participation is controlled by server.yml:
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
environment: ${server.environment:dev}
Controller connection settings come from portal-registry.yml:
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
Current light-gateway discovery uses the microservice registration token from
LIGHT_PORTAL_AUTHORIZATION or portalToken. The token is sent in the
service/register payload. controllerDiscoveryToken is reserved for clients
that use the dedicated /ws/discovery endpoint and is not part of the current
gateway lookup path.
The runtime converts portalUrl to /ws/microservice, strips any query string,
and starts the shared PortalRegistryClient when registry is enabled. The
client must be connected and registered before discovery lookup can succeed.
Gateway Invocation Path
Startup path:
config-server/local config
-> light-runtime loads server.yml, client.yml, portal-registry.yml
-> RuntimeConfig.service_identity is built from server/bootstrap config
-> RuntimeConfig.registry_client is created when registry is enabled
-> runtime startup registers the gateway with controller
-> light-gateway builds Pingora proxy state from RuntimeConfig
Request-time path:
incoming request
-> handler.yml selects a handler chain
-> handler resolves direct target, serviceId, or static target
-> handler calls PortalRegistryClient.lookup_discovery when serviceId discovery is needed
-> controller returns DiscoverySnapshot
-> handler converts nodes to Pingora ProxyTarget or base URL
-> Pingora proxies the request
PortalRegistryClient.lookup_discovery sends JSON-RPC method
discovery/lookup over the registered websocket and waits for a response. If
the websocket is not connected, lookup fails with a registry client connection
error.
Handler Usage
Router
The router handler supports both direct routing and service discovery.
Resolution order:
service_urlrequest routing, when configured and present.service_idfrom query/header/path-prefix logic.- Controller discovery with
serviceIdand optionalenvTag. direct-registry.directUrlsusingserviceId|envTag, thenserviceId.
Direct registry is the standard static fallback service map when the controller cannot resolve the target.
WebSocket Router
The websocket handler resolves the target service from header, query, or
pathPrefixService. It checks direct-registry.directUrls first, then passes
serviceId, optional envTag, and protocol to discovery. Connected http and
https nodes are converted to upstream WebSocket targets and Pingora handles
the upgrade proxying.
MCP Router
The mcp handler can route tools by direct targetHost or discovered
serviceId.
Resolution order:
- Tool
targetHost. direct-registry.directUrlsusingserviceId|envTag, thenserviceId.- Tool
serviceIdthrough controller discovery.
When a tool uses serviceId, portal registry is only required if no direct URL
mapping exists. The tool can also specify envTag and protocol to constrain
direct URL and discovery lookup.
Token Handler
The token handler can resolve the OAuth token server by direct
oauth.token.server_url or by oauth.token.serviceId.
Resolution order:
- Direct token server URL.
direct-registry.directUrlsusing token serverserviceId.- Token server
serviceIdthrough controller discovery.
The selected node prefers https and then falls back to http. If discovery is
required and portal registry is not enabled, token injection fails explicitly.
SPA Auth
The stateless SPA auth and MSAL exchange token clients use the same token-server resolution model as the token handler:
- Direct token server URL.
direct-registry.directUrlsusing token serverserviceId.- Token server
serviceIdthrough controller discovery.
This keeps BFF deployments independent from fixed OAuth hostnames when the token service is registered with the controller.
Direct URLs And Fallbacks
Discovery should not override an explicit direct URL selected by a handler.
Direct URLs are operator intent and should remain authoritative. The standard
shared direct URL map is direct-registry.directUrls.
Static fallback is handler-specific:
- The router checks portal-registry discovery before
direct-registry.directUrls. - Other service-id paths can check
direct-registry.directUrlsbefore controller discovery when they need local/static overrides. - MCP, token, SPA auth, JWK, and WebSocket service-id routing can use
direct-registry.directUrlswithout per-handler duplicate maps.
Base Path For Path-Based Ingress
A direct URL may carry a path, which is the base path of the target service. It is used when the service runs in a Kubernetes cluster behind an ingress that routes on a path prefix:
direct-registry.directUrls:
com.networknt.petstore-1.0.0: https://api.example.com/namespace1/service1
The router, proxy, WebSocket, A2A, and MCP handlers prepend that path to the
upstream request path, so a request for /v1/pets is sent to
https://api.example.com/namespace1/service1/v1/pets. The ingress selects the
pod from /namespace1/service1 and strips it, so the pod receives the original
path. The token, JWK, and SPA auth clients append their configured URI to the
direct URL, which keeps the base path as well.
The Host header and the TLS SNI come from the host of the URL, not the base
path.
A discovered node expresses the same thing with the reserved basePath
registration tag. A service sets basePath in its server.yml, the runtime
registers it as a tag, the controller round-trips it, and every consumer of a
discovery node prepends it:
# server.yml of the service behind the ingress
basePath: ${server.basePath:/namespace1/service1}
DiscoveryNode::base_path() normalizes the tag (leading slash, no trailing
slash) and DiscoveryNode::base_url() includes it, so the router, the WebSocket
router, MCP tools, the token handler, SPA auth, and the JWK client all route
through the ingress prefix without their own mapping. A node that does not
advertise the tag has an empty base path and is reached by address and port as
before, which keeps every existing deployment unchanged.
Only a path is accepted. A value that carries a query, a fragment, a relative
segment, or an empty segment is rejected, because the base path is concatenated
with the path of the request and with the uri of an endpoint. A percent encoded
dot segment such as /namespace/%2e%2e/admin is rejected as well, since a url
parser decodes it before it resolves the segment and the prefix would route to a
path other than the one validated. An accepted base path is one that survives
url parsing unchanged. A basePath in
server.yml that is not a path fails the startup. A node that advertises one is
skipped by every consumer, rather than reached at the root, because a service
behind an ingress serves nothing useful without its prefix and the request would
land on whatever else that ingress serves. Discovery then yields no usable node
for the service, so the router falls back to direct-registry or answers with a
502, which is a visible failure instead of a request sent to another backend.
basePath is a reserved identity tag, which means the runtime is its only
authority. The transport metadata of an application is merged into the
registration, so a reserved key it carries is dropped and the identity tags are
applied last. A metadata update publishes the complete tag map of the application,
so send_metadata_update restores a reserved tag that an update leaves out and
drops one that an update carries but the runtime never registered. Without the
first, a service that publishes operational tags would lose its base path
shortly after it registers and on every reconnect. Without the second, an
application could route its own traffic elsewhere, or drop the prefix with a
blank value, without the operator changing any configuration.
This keeps failure behavior predictable. Product configs that require dynamic discovery should fail requests loudly when the controller connection is down instead of silently choosing an unrelated target.
Load Balancing
The controller returns a list of matching nodes. The handler is responsible for choosing one.
Current behavior is intentionally simple:
- drop disconnected nodes
- drop nodes with port
0 - drop unsupported protocols
- prefer
httpsfor token-server resolution - round-robin or index-based selection where the handler already has an index
Future hardening can add weighted selection, zone preference, health score, least-connections, or sticky routing. Those policies should live in the handler or a shared target-selection helper, not in the controller protocol.
Failure Semantics
Startup behavior is controlled by startOnRegistryFailure:
true: the gateway can start if initial controller registration times out; the registry client keeps retrying in the background.false: initial controller registration timeout fails startup.
Request-time behavior depends on handler fallback:
- with direct URL: discovery is bypassed
- with usable static fallback: handler may continue
- with discovery-only config: return an explicit gateway error
The runtime should continue reconnecting the registry websocket. Once the client is registered again, new discovery lookups can succeed without restarting the gateway.
Security
The gateway registers through /ws/microservice with the portal registry token.
The controller validates the registration token and then allows discovery RPCs
on that registered socket.
Security expectations:
- Use TLS for controller connections outside local development.
- Keep hostname verification enabled outside local development.
- Prefer environment-provided token values over static config files.
- Mask
portalTokenandcontrollerDiscoveryTokenin module-registry output. - Do not pass registry tokens to downstream services.
- Do not trust discovery data from an untrusted controller.
Discovery returns transport endpoints. Authentication, authorization, rate limit, CORS, header mutation, token injection, and access-control decisions remain normal handler-chain responsibilities.
Config Server Model
In production, light-portal owns product configuration and config-server
delivers resolved files at startup.
A product that needs controller-backed discovery should include:
server.ymlwithenableRegistry: trueportal-registry.ymlwithportalUrland a valid portal token sourcedirect-registry.ymlorvalues.ymlentries underdirect-registry.directUrlsfor transition services that are not registered in the controller yet- handler-specific config that uses
serviceIdinstead of direct host URLs handler.ymlchains that include the relevant handler IDs
For local Docker Compose, the Rust gateway must not keep the default
https://localhost:8438 controller URL because localhost is the gateway
container. Use portalRegistry.portalUrl: https://controller:8438, pass
LIGHT_PORTAL_AUTHORIZATION, and keep static transition mappings in
direct-registry.directUrls.
The same binary can therefore run as:
- gateway
- sidecar
- proxy server
- proxy client
- balancer
- BFF
The product identity comes from config, not from a separate executable.
Compatibility Notes
The current Rust and Java controllers are compatible with the gateway discovery path because both support:
/ws/microserviceservice/register- discovery lookup on the registered microservice socket
serviceId,envTag, andprotocolfiltersDiscoverySnapshot.nodes- connected-node metadata with
address,port, andprotocol
The gateway does not currently depend on /ws/discovery, although that endpoint
can remain available for external discovery clients.
Future Work
- Add optional discovery subscriptions for handlers that benefit from a local in-memory discovery cache.
- Add shared target-selection policies for weighted, sticky, or zone-aware routing.
- Expose discovery health through the module registry or an admin endpoint.
- Add an integration test that starts a controller, registers a backend, starts light-gateway, and verifies an end-to-end proxied request through discovery.
- Decide whether
controllerDiscoveryTokenshould be used by any standalone discovery-only client in light-fabric. - Document operational examples for gateway, sidecar, WebSocket, MCP, token handler, and BFF product profiles.
Tracing
Light-Fabric uses Rust tracing for application logs and runtime diagnostics.
The same tracing events must support two different consumers:
- operators and developers reading live logs from the console or control plane
- log platforms such as Splunk that ingest structured JSON
The logging design should keep one source of truth for emitted events and make the output format configurable at the edge of the process.
Goals
- Preserve the current human-readable console format for local development and controller-streamed logs.
- Support newline-delimited JSON logs for Splunk and other log ingestion systems.
- Allow deployments to choose text or JSON console output without changing application code.
- Allow authorized control-plane users to change log levels and logger targets without restarting the service.
- Avoid coupling Light-Fabric services directly to Splunk availability, credentials, retry policy, or backpressure handling.
- Keep log fields stable enough for portal-view, controller, and Splunk queries.
Non-Goals
- Implement a Splunk HTTP Event Collector client inside every Light-Fabric service.
- Mix human text logs and JSON logs on the same stream.
- Use
values.ymlto mutate process environment variables. Environment variables are startup inputs; runtime changes should use an explicit logging configuration model.
Current State
The application binaries initialize tracing_subscriber locally. The current
format is text-oriented and is easy to read in a terminal, Docker logs, or a
controller stream. Some binaries also support an ANSI toggle so container logs
can avoid escape sequences.
This works well for humans, but it is less reliable for Splunk field extraction. Splunk can ingest text logs, but structured JSON gives predictable fields for filtering, dashboards, alerts, and correlation.
Output Formats
Light-Fabric should support the following output formats:
| Format | Intended Consumer | Notes |
|---|---|---|
text | humans, local development, controller live log stream | Existing behavior. Best for direct reading. |
json | Splunk, OpenTelemetry Collector, Kubernetes log collectors | Newline-delimited JSON. Best for machine ingestion. |
The output should be selected with an environment variable:
LIGHT_LOG_FORMAT=text
or:
LIGHT_LOG_FORMAT=json
If the variable is absent, the default should remain text to preserve existing
operator behavior.
RUST_LOG should continue to provide the startup filter:
RUST_LOG=info
RUST_LOG=light_gateway=debug,info
RUST_LOG=light_workflow=debug,info
Single Console Stream
For most deployments, the preferred model is a single console stream with a configurable format:
application tracing event
|
v
tracing_subscriber fmt layer
|
+-- stdout/stderr as text or JSON
This has the lowest runtime overhead because each event is formatted and written once. It also keeps container logging simple: the platform captures the process console stream, and the customer chooses whether that stream is text or JSON.
When LIGHT_LOG_FORMAT=json, the console output should be newline-delimited
JSON:
{"timestamp":"2026-06-03T14:12:41.233Z","level":"INFO","target":"light_gateway","fields":{"message":"proxy request completed","method":"GET","path":"/api/customer","status":200,"elapsed_ms":18,"correlation_id":"abc-123"}}
Raw JSON is readable, but it is not as pleasant as the text format. For the control plane, portal-view should parse JSON log lines and render a human projection:
14:12:41.233 INFO light_gateway proxy request completed
method=GET path=/api/customer status=200 elapsed_ms=18 correlation_id=abc-123
This lets Splunk receive structured logs while portal-view remains readable for operators.
Portal-View Rendering
The controller should stream log lines without needing to understand every field. Portal-view can detect whether a line is JSON:
- Trim the line.
- If it starts with
{, try to parse it as JSON. - If parsing succeeds, render common fields in a stable layout.
- If parsing fails, render the original line as plain text.
The renderer should treat JSON parsing as an enhancement, not a hard requirement. This keeps mixed historical output, startup messages, and unrelated tool output usable.
Recommended display fields:
| JSON Field | Display Use |
|---|---|
timestamp | leading timestamp |
level | severity badge/text |
target | module or service source |
fields.message | main message |
fields.correlation_id | request correlation |
fields.request_id | request identifier, when present |
fields.status | HTTP or operation status |
fields.elapsed_ms | latency |
Unknown fields can be shown in an expandable details view or appended as
key=value pairs.
Splunk Ingestion
A log file is not the only option for Splunk ingestion.
Console JSON in Containers
For Kubernetes and container deployments, console JSON is usually the best default. The service writes JSON to stdout/stderr, and the platform logging agent collects the container log stream. Splunk Connect for Kubernetes, OpenTelemetry Collector, or an equivalent customer-managed collector can parse the JSON and send it to Splunk HTTP Event Collector.
This avoids application-level Splunk credentials and keeps retry, batching, and backpressure in the collector.
JSON Log File
For VM or bare-metal deployments where the customer already uses Splunk Universal Forwarder, a JSON log file is also valid. In that mode the application would write newline-delimited JSON to a rotating file, and the forwarder or OpenTelemetry filelog receiver would tail it.
This mode is useful when stdout is reserved for human-readable controller logs, but it formats and writes each event through an additional sink if text console output remains enabled.
Direct Splunk HEC
Direct HTTP Event Collector delivery from the application is possible but should not be the default. It adds Splunk endpoint configuration, token management, retry policy, buffering, and failure handling to every service. A collector or forwarder is a cleaner boundary for production deployments.
Dual Sink Option
If a deployment must keep text console logs and produce JSON at the same time, Light-Fabric can use multiple tracing layers:
application tracing event
|
v
tracing subscriber registry
|
+-- text layer -> stdout/stderr
|
+-- JSON layer -> rolling file
This preserves the current control-plane stream and gives Splunk a clean JSON source. The tradeoff is extra formatting and I/O work per event.
Use this mode only when a single JSON console stream is not acceptable for the operator experience.
Configuration
The design supports both single-stream and dual-sink logging through configuration. The two common deployment profiles are:
| Deployment | Console Output | JSON File | Typical Splunk Path |
|---|---|---|---|
| Kubernetes/container | json | disabled | container log collector to Splunk HEC |
| Bare metal/VM with human console | text | enabled | Splunk Universal Forwarder or filelog receiver tails the JSON file |
| Local development | text | disabled | terminal or controller stream only |
The minimal configuration should be:
LIGHT_LOG_FORMAT=text
LIGHT_LOG_ANSI=false
RUST_LOG=info
JSON console mode:
LIGHT_LOG_FORMAT=json
LIGHT_LOG_ANSI=false
RUST_LOG=info
Optional dual-sink file mode:
LIGHT_LOG_FORMAT=text
LIGHT_LOG_ANSI=false
LIGHT_LOG_JSON_FILE_ENABLED=true
LIGHT_LOG_JSON_FILE_DIR=/var/log/light-fabric
LIGHT_LOG_JSON_FILE_NAME=light-gateway.jsonl
LIGHT_LOG_JSON_FILE_ROTATION=daily
RUST_LOG=info
In this dual-sink mode, the application emits the same tracing event to both sinks: text to stdout/stderr for humans and controller-streamed logs, and JSON to the configured file for Splunk ingestion.
Service-specific aliases such as GATEWAY_LOG_ANSI, AGENT_LOG_ANSI, or
WORKFLOW_LOG_ANSI can remain during migration, but the long-term interface
should converge on LIGHT_LOG_* variables shared by all Light-Fabric binaries.
Runtime Logging Control
Light-Fabric should support the Java control-plane behavior where an authorized operator changes log levels and logger targets from portal-view without restarting the service.
Rust can support this through tracing_subscriber::reload. Instead of installing
a fixed EnvFilter, the runtime should wrap the filter in a reloadable layer and
keep a reload handle in a shared logging controller:
application tracing event
|
v
reloadable EnvFilter
|
v
text/json formatting layers
The reloadable part is the filter only. A filter can change the global level and individual logger targets:
info
debug
info,light_gateway=debug
info,light_gateway=debug,light_pingora::security=trace
info,light_pingora::security=off
This matches the practical Java use case: enable debug or trace for one logger
while keeping the rest of the service at info.
Dynamic Versus Restart-Only Settings
| Setting | Dynamic | Reason |
|---|---|---|
| Global log level | yes | Updates the reloadable EnvFilter. |
| Per-target logger level | yes | Updates the reloadable EnvFilter. |
Disable a target with target=off | yes | Updates the reloadable EnvFilter. |
Console format text/json | no | Requires rebuilding formatter layers. |
| JSON file enabled/disabled | no | Requires adding or removing a writer layer. |
| JSON file directory/name/rotation | no | Requires replacing the appender and guard. |
| ANSI setting | no | Formatter setting; treat as startup-only. |
Startup Precedence
The startup filter should use this precedence:
RUST_LOG, when present.logging.filterfromvalues.yml.- The service default, such as
infoorlight_workflow=debug,info.
This preserves existing RUST_LOG behavior for local and container deployments
while allowing managed deployments to define a persistent default filter in
config.
Example values.yml:
logging.filter: info
More targeted example:
logging.filter: info,light_gateway=debug,light_pingora::security=trace
values.yml should not overwrite environment variables and should not be the
normal path for day-to-day control-plane log-level changes. It should provide the
baseline filter that the logging module reads at startup. If an operator wants to
restore that baseline after a live debugging change, reload_modules can reload
runtime/logging from the latest resolved values.
Changing config server values and then triggering reload is therefore a persistence/reset workflow, not the primary live-control workflow.
MCP Tools
The runtime MCP tool surface should expose logging control alongside existing
runtime tools such as get_service_info, get_modules, and reload_modules.
Recommended tools:
| Tool | Purpose |
|---|---|
get_logging_filter | Return the current effective filter and startup source. |
set_logging_filter | Validate and apply a new live filter immediately. This is the normal portal-view control path. |
reload_modules with runtime/logging | Reset the live filter from the configured baseline in values.yml or remote values. |
Example live filter update:
{
"name": "set_logging_filter",
"arguments": {
"filter": "info,light_gateway=debug"
}
}
Example reset from the configured baseline:
{
"name": "reload_modules",
"arguments": {
"modules": ["runtime/logging"]
}
}
The service response should include the active filter and status:
{
"status": "success",
"filter": "info,light_gateway=debug"
}
Invalid filters should be rejected without changing the current filter:
{
"status": "error",
"message": "invalid logging filter: ..."
}
Portal-View Flow
The portal-view control plane should follow the same route used for other runtime management tools:
portal-view
-> controller
-> portal-registry/runtime instance connection
-> service runtime MCP handler
-> logging control
The UI can offer:
- a global level selector:
off,error,warn,info,debug,trace - per-target rows for Rust targets such as
light_gatewayorlight_pingora::security - an advanced filter text box for the full
EnvFilterexpression - an apply action that calls
set_logging_filter - a reset action that reloads
runtime/loggingfrom the configured baseline - an optional “save as default” action that persists the filter to config server
The advanced filter is important because Rust logger targets are module paths, and operators may need precise target-level control during incident debugging.
The default portal-view workflow should be:
operator changes filter
-> portal-view calls set_logging_filter
-> service updates the reloadable EnvFilter immediately
Portal-view should not require this slower path for a temporary debug change:
operator changes filter
-> portal-view updates config server
-> portal-view calls reload_modules
-> service reloads values.yml
That slower path is still useful when the operator intentionally wants the new filter to survive service restart or redeploy.
JSON Field Shape
JSON logs should be stable enough for both portal-view rendering and Splunk searches. Recommended fields include:
| Field | Meaning |
|---|---|
timestamp | event time in UTC |
level | ERROR, WARN, INFO, DEBUG, or TRACE |
target | Rust module or logical component |
fields.message | human message |
fields.service | logical service name, such as light-gateway |
fields.instance_id | runtime instance, when known |
fields.host_id | tenant/host context, when safe to log |
fields.correlation_id | cross-service request correlation |
fields.request_id | request identifier |
fields.method | HTTP method, when applicable |
fields.path | request path without sensitive query string |
fields.status | response or operation status |
fields.elapsed_ms | operation duration |
Sensitive values must not be logged in either format. This includes tokens, API keys, session cookies, full authorization headers, raw secrets, and request or response payload fields that may contain PII.
Implementation Notes
Use tracing_subscriber as the formatting boundary. The JSON format requires
the json feature:
tracing-subscriber = { version = "0.3", features = ["env-filter", "fmt", "json"] }
File output should use tracing_appender:
tracing-appender = "0.2"
If non-blocking file output is used, the returned WorkerGuard must be kept
alive until process shutdown so buffered log lines are flushed.
The implementation should move per-binary init_tracing() logic into a shared
runtime helper so light-gateway, light-agent, light-workflow, and
light-deployer expose the same behavior.
For dynamic filtering, the shared helper should:
- Build the initial
EnvFilterfromRUST_LOG,logging.filter, or the service default. - Install the filter through
tracing_subscriber::reload. - Keep the reload handle in a
LoggingControlvalue. - Register a reloadable module named
runtime/loggingwithModuleRegistry. - Add runtime MCP handlers for
get_logging_filterandset_logging_filter. - Reject invalid filter expressions before swapping the active filter.
Recommendation
Start with configurable single-stream console output:
- default
LIGHT_LOG_FORMAT=text - production/Splunk option
LIGHT_LOG_FORMAT=json - portal-view JSON parsing and human-friendly rendering
- no direct Splunk dependency in the application
Add dual-sink JSON file output only for customers who cannot change the console stream to JSON but still require structured Splunk ingestion.
Graceful Service Shutdown
Status: Implemented; deployment qualification pending
Purpose
Light Fabric services must stop promptly when an orchestrator asks a container to terminate, while still protecting requests and durable background work that are already in progress.
The required behavior is:
- install shutdown signal handlers before the service becomes ready
- react to both
SIGINTandSIGTERMon Unix - stop accepting new work immediately
- allow in-flight work to drain
- release service-owned resources and unregister where required
- exit as soon as draining and cleanup finish
- enforce a configured application deadline for work that does not finish
When a service has no in-flight work, shutdown should normally complete in less than one second. A container stop timeout is a last-resort safety boundary, not a delay that the application should consume on every stop.
Problem Statement
Container engines normally stop a container by sending SIGTERM to PID 1,
waiting for a configured timeout, and then sending SIGKILL if the process is
still running.
Several current Rust applications wait only for:
#![allow(unused)]
fn main() {
tokio::signal::ctrl_c().await?;
}
That future handles SIGINT, but not the SIGTERM sent by Docker, Podman,
Kubernetes, and most process supervisors. Other applications run an Axum server
or a set of background tasks forever without installing either handler.
This is especially visible in the local Portal deployment. The external
lightapi/portal-config-loc repository invokes Compose with
down --timeout 30 in scripts/deploy-local.sh inside
stop_docker_compose(). A Rust binary running as container PID 1 does not exit
through its intended shutdown path when it does not handle SIGTERM; the
engine waits for the entire timeout and then kills the process. The delay is
therefore unrelated to request volume. This motivating deployment setting does
not live in the light-fabric repository.
The current code also has inconsistent graceful-shutdown behavior:
light-runtimeapplications callRunningRuntime::shutdown()only after their application-level signal future completes.light-pingoraappliesserver.shutdownGracefulPeriodas a maximum drain time, but the application must first initiate runtime shutdown.light-axuminitiates graceful shutdown without a deadline and currently does not applyserver.shutdownGracefulPeriod.RunningRuntime::shutdown()awaits transport shutdown and each module hook without a wall-clock backstop.PingoraTransport::stop()joins a blocking server thread without an outer bound. The pinnedpingora-core 0.8.0graceful path also sleeps for its full configured runtime shutdown timeout, even if no work remains.- standalone Axum applications such as
controller-rs,light-oauth, andconfig-serverdo not share a shutdown contract. - task-oriented applications such as
light-workflowneed cancellation and task-join behavior in addition to HTTP request draining.
Goals
- Provide one cross-platform shutdown-signal implementation in
light-runtime. - Make
SIGTERMthe canonical orchestrator signal and retainSIGINTfor interactive use. - Make graceful shutdown the default path for Light Runtime transports.
- Use
server.shutdownGracefulPeriodconsistently as a maximum drain period. - Enforce one wall-clock deadline around the complete application shutdown sequence, including deregistration, transport drain, and module cleanup.
- Exit immediately after the listener, in-flight work, and cleanup complete.
- Give background workers an explicit cancellation and join contract.
- Preserve a larger orchestrator timeout as protection against process bugs.
- Make forced termination observable in automated tests and operations.
Non-Goals
- Do not guarantee completion of arbitrary work after the graceful deadline.
- Do not use an orchestrator timeout as an application sleep period.
- Do not change the meaning of readiness, liveness, or startup timeouts.
- Do not make
stop_signal: SIGINTthe permanent deployment solution. - Do not treat
docker compose down --timeout 0orSIGKILLas graceful shutdown. - Do not require all applications to adopt the same HTTP framework.
Shutdown Contract
Accepted signals
On Unix, every long-running service must handle both:
SIGTERM, used by container engines and orchestratorsSIGINT, used by an interactive Ctrl-C and local development tools
On non-Unix platforms, the shared implementation waits for the platform’s Ctrl-C event. Platform-specific service-manager integration can be added behind the same API later.
The first accepted signal starts graceful shutdown. A second SIGINT or
SIGTERM collapses the remaining drain budget to zero, starts only the
mandatory cleanup floor described below, and then terminates with the
deadline-exceeded exit status if the process has not already stopped. This
makes Ctrl-C twice useful during interactive development without making the
first signal destructive. The second signal sets the hard exit deadline to the
earlier of the existing hard deadline and now + MANDATORY_CLEANUP_FLOOR; it
never extends shutdown.
Shutdown phases
The service lifecycle gains the following terminal phases:
Starting -> Ready -> Quiescing -> Draining -> CleaningUp -> Stopped
|
+-------> AbortingStartup ---------------------------> Stopped
- AbortingStartup: cancel the active startup phase, keep admission closed,
unwind resources recorded by
StartupGuard, and apply only the mandatory cleanup floor. - Quiescing: atomically mark the instance unready, close the admission gate to new application work, and send a bounded deregistration request so peers stop advertising the instance.
- Draining: wait for accepted HTTP requests, WebSockets, streams, and claimed background work according to their component policy.
- CleaningUp: run module shutdown hooks, flush durable buffers, close clients, and release leases or registrations.
- Stopped: return from
mainwith exit code zero only if the sequence completed before its hard deadline.
The application deadline begins when shutdown is accepted. Components receive the same shutdown context and remaining deadline rather than each receiving the full configured period sequentially.
Readiness and admission semantics
There is no fixed pre-drain dwell. Shutdown changes the runtime readiness state
and closes a shared admission gate synchronously before awaiting network I/O.
The initial implementation uses one cloneable AdmissionGate, created closed
by the runtime and opened exactly once at the Ready transition:
#![allow(unused)]
fn main() {
pub enum AdmissionKind {
Application,
Control,
}
#[derive(Clone)]
pub struct AdmissionGate { /* atomic state and in-flight counters */ }
impl AdmissionGate {
pub fn open(&self);
pub fn close(&self);
pub fn try_enter(
&self,
kind: AdmissionKind,
) -> Result<AdmissionPermit, AdmissionClosed>;
}
}
AdmissionPermit increments the relevant in-flight counter before dispatch and
decrements it from Drop. Application admission fails whenever the gate is
closed. Control admission remains available during Quiescing only for
framework-declared liveness, readiness, and shutdown-status handlers; it cannot
claim work, mutate application state, or start an unbounded operation. Readiness
reads the same gate and reports not ready as soon as close() returns.
The default classification is Application. An application migration may
declare an exact method-and-path route as Control only in its reviewed route
inventory; prefix and wildcard bypasses are forbidden. Until such an inventory
exists, existing health routes also receive the shutdown 503, which is a
valid not-ready response. This makes the initial behavior fail closed and keeps
the admission exception set auditable.
Axum installs the admission layer outside the application router. Pingora calls
the same gate before handler dispatch. A rejected HTTP application request gets
503 Service Unavailable, Connection: close, and Retry-After: 0. A worker
must acquire an Application permit before claiming a unit; failure means it
stops its claim loop. WebSocket and stream upgrades retain the permit for their
full lifetime. No application may implement a second, unsynchronized shutdown
flag.
The runtime then sends the bounded deregistration request. Once it is acknowledged or its small bound expires, transport drain begins. Upstream readiness propagation delay is not modeled as a sleep and does not create a second grace period. Time used by deregistration is charged against the one application deadline.
For Axum, Handle::graceful_shutdown combines listener close and connection
drain. The observable Quiescing phase therefore comes from the runtime state,
admission gate, deregistration event, and phase metrics, not a distinct Axum
transport state. During the bounded deregistration step the socket may still
accept a connection, but new application work receives the defined 503.
After deregistration is acknowledged or bounded, the handle is invoked and new
TCP connections are refused. Transport tests distinguish these two observable
boundaries instead of treating admission rejection and listener closure as the
same event.
Deadline behavior
server.shutdownGracefulPeriod is the application-level maximum, in
milliseconds. It is not a minimum wait.
The top-level runtime supervisor is the deadline enforcer. It creates one
absolute graceful deadline and wraps the entire normal shutdown sequence with
tokio::time::timeout_at. Transport-local timeouts are cooperative component
bounds, not substitutes for this backstop.
For example, with a value of 2000:
- zero in-flight work should stop immediately
- a 500 ms request may finish normally
- work still active at two seconds is cancelled or disconnected according to the component contract
- cleanup uses only the time remaining in the same deadline
A graceful deadline expiry cancels the shared shutdown context. The internal
sequence then allows only MANDATORY_CLEANUP_FLOOR for emergency cleanup that
was prepared in advance and returns ShutdownOutcome::DeadlineExceeded. The
production run_until_shutdown supervisor emits the final deadline-exceeded
record to stderr and calls std::process::exit(1); the lower-level API returns
the corresponding error. Calling process::exit at the production boundary is
intentional: merely returning an error can still hang while Tokio drops a
runtime that owns an unbounded spawn_blocking task such as Pingora’s
server-thread join. Exit code 1 means application shutdown failure; exit code
137 still means the container engine had to send SIGKILL and is a stronger
qualification failure.
shutdownGracefulPeriod: 0 skips request and worker drain. It does not remove
the emergency cleanup budget. Define a non-configurable initial
MANDATORY_CLEANUP_FLOOR of 250 ms for readiness change, admission closure,
best-effort deregistration/socket close, and already-prepared checkpoints. The
hard process deadline is therefore:
signal time + shutdownGracefulPeriod + MANDATORY_CLEANUP_FLOOR
Ordinary module cleanup remains inside shutdownGracefulPeriod; it is not
added again. The floor is exclusively for emergency cleanup after cancellation
and does not authorize starting a new durable flush after the graceful
deadline. Deployments should use zero only for tests or an explicitly
documented emergency policy.
Shared Runtime Design
Signal API
Add a small signal module to light-runtime. Handler installation and waiting
must be separate operations so the process cannot become ready before it owns
the handlers:
#![allow(unused)]
fn main() {
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ShutdownReason {
Interrupt,
Terminate,
Programmatic,
}
pub struct ShutdownWatcher { /* platform signal streams */ }
impl ShutdownWatcher {
pub fn install() -> std::io::Result<Self>;
pub async fn recv(&mut self) -> ShutdownReason;
}
}
ShutdownWatcher::install() synchronously creates the platform signal streams.
On Unix it installs streams for both interrupt and terminate. It must not defer
registration until the first poll of recv(). The non-Unix implementation
constructs the available platform Ctrl-C stream behind the same API.
ShutdownWatcher::recv() returns only Interrupt or Terminate;
Programmatic is reserved for the lower-level embedding and test API.
Tokio’s Unix signal stream requires an active reactor and panics when created
outside a runtime context. ShutdownWatcher::install() must therefore be the
first lifecycle statement inside the asynchronous body created by
#[tokio::main], not a call made in synchronous code before entering the Tokio
runtime:
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let watcher = ShutdownWatcher::install()?;
// Logging and configuration follow handler installation.
let runtime = build_runtime()?;
runtime.run_until_shutdown(watcher).await?;
Ok(())
}
The API documentation and a subprocess panic test must call out this reactor precondition explicitly. Installing before logging is acceptable; a signal received during later startup remains pending in the watcher.
The production lifecycle is therefore:
#![allow(unused)]
fn main() {
let watcher = ShutdownWatcher::install()?;
runtime.run_until_shutdown(watcher).await?;
}
LightRuntime::run_until_shutdown owns cancellable startup, readiness
publication, signal receipt, and shutdown so the ordering is enforced by
construction. The lower-level start() API remains available for embedding
and tests, but its documentation must require an installed watcher or another
programmatic cancellation owner before start() can publish readiness.
The supervisor retains the watcher after the first signal and selects between
shutdown completion and watcher.recv() again. A second accepted signal
cancels the drain context and moves directly to mandatory cleanup.
Runtime convenience API
Add a production convenience method that owns the complete lifecycle. The following pseudocode shows the required concurrency; helper types may package the select and deadline handling differently:
#![allow(unused)]
fn main() {
impl<T: TransportRuntime> LightRuntime<T> {
pub async fn run_until_shutdown(
self,
mut watcher: ShutdownWatcher,
) -> Result<(), RuntimeError> {
let startup_cancel = CancellationToken::new();
let mut startup = Box::pin(
self.start_cancellable(startup_cancel.child_token()),
);
let running = tokio::select! {
biased;
reason = watcher.recv() => {
startup_cancel.cancel();
return finish_startup_abort(
reason,
startup,
&mut watcher,
).await;
}
result = &mut startup => result?,
};
let reason = watcher.recv().await;
tracing::info!(?reason, "shutdown signal received");
running.shutdown_with_watcher(reason, &mut watcher).await
}
}
}
Applications using LightRuntimeBuilder then use:
#![allow(unused)]
fn main() {
let watcher = ShutdownWatcher::install()?;
runtime.run_until_shutdown(watcher).await?;
}
This replaces app-local ctrl_c() calls. RunningRuntime::shutdown() remains
available for tests, embedding, and programmatic lifecycle management. It
delegates to the same internal shutdown sequence with
ShutdownReason::Programmatic; it has no second-signal branch. The production
watcher path calls shutdown_with_watcher, which selects between that sequence
and another accepted signal. Neither public entry point duplicates the
shutdown implementation. The internal sequence returns a structured
ShutdownOutcome. RunningRuntime::shutdown() converts a deadline outcome to
RuntimeError::ShutdownDeadlineExceeded and never terminates its caller’s
process. LightRuntime::run_until_shutdown() is the production policy boundary:
after emergency cleanup and the final stderr record, it converts that same
outcome to std::process::exit(1). An embedding caller that uses the lower-level
API owns its own escalation policy.
Cancellable startup
Startup is cancellable. Retaining a signal until start() finishes is not
sufficient because remote bootstrap and controller registration perform network
I/O, and registration alone is allowed five seconds by default.
start_cancellable() owns a StartupGuard from its first operation. As startup
progresses, the guard records every resource that requires asynchronous unwind:
- remote bootstrap/config fetch and any staged config-cache write
- registered lifecycle participants and application resources
- a partially or fully bound transport handle
- controller registration state, socket, and reconnect task
- readiness/admission state
Ownership must reach the guard before the next cancellation point. This is an API invariant, not a convention. In particular, controller startup is split into two operations:
#![allow(unused)]
fn main() {
let registry_session = registry_client.start_session(/* ... */)?;
startup_guard.set_registry_session(registry_session);
startup_guard
.registry_session()
.wait_until_registered(startup_cancel.child_token())
.await?;
}
RegistrySession owns the client, socket-generation state, reconnect task, and
task join handle. start_session() may spawn the task, but after spawning it
must return the owning session without another .await. Dropping a registration
wait therefore cannot detach the task; startup abort calls the session’s
deadline-aware shutdown operation through StartupGuard.
The same rule applies to transport binding. A bind() implementation owns an
internal BindingGuard until it returns BoundTransport. A listener, thread,
or task created inside bind() must either remain owned by that guard across
every .await, or be created as the final non-awaiting operation immediately
before the handle is returned. Cancelling bind() must synchronously close any
listener and cancel any task that has not been handed to StartupGuard; if a
resource requires asynchronous unwind, bind() must expose a staged owned
handle before beginning that operation. A transport implementation that can
detach work when its future is dropped does not satisfy TransportRuntime.
Each startup phase selects between its work and the startup cancellation token. Configuration-cache writes and other persistent startup effects must use an atomic stage-and-commit pattern so cancellation cannot expose a partial file or half-published state.
If the watcher wins before Ready, the runtime:
- cancels the in-progress bootstrap or registration future
- keeps readiness false and seals admission closed
- seals the lifecycle registry against new participants
- asks
StartupGuardto close any bound listener and partial registration - invokes already-registered participants in startup-abort mode
- exits zero if unwind completes within
MANDATORY_CLEANUP_FLOOR
There is no request-drain period because the service never reached Ready.
The mandatory cleanup floor is the complete startup-abort budget. If unwind
does not finish inside it, the supervisor uses the same final stderr record and
std::process::exit(1) policy as a graceful-deadline failure. A second signal
while aborting startup collapses to the remaining portion of that floor and
never extends it.
Dropping the start() future alone is not the abort mechanism: that would lose
ownership of partially bound resources without awaited cleanup. The
resource-owning StartupGuard and cancellation-aware phase boundaries are
required implementation mechanisms.
Shutdown context and module contract
The deadline requires changing the Module trait.
Wrapping today’s on_shutdown(&RuntimeConfig) future in timeout() would stop
waiting but would cancel that future at an arbitrary await point. That is not a
safe contract for a durable flush or checkpoint.
Add a shared context and pass it into every hook:
#![allow(unused)]
fn main() {
pub enum ShutdownMode {
Graceful,
StartupAbort,
Emergency,
}
pub struct ShutdownContext {
pub reason: ShutdownReason,
pub mode: ShutdownMode,
pub deadline: tokio::time::Instant,
pub cancellation: tokio_util::sync::CancellationToken,
}
impl ShutdownContext {
pub fn remaining(&self) -> Duration;
pub async fn cancelled(&self);
}
#[async_trait]
pub trait Module: Send + Sync {
// Existing lifecycle methods omitted.
async fn on_shutdown(
&self,
config: &RuntimeConfig,
context: &ShutdownContext,
) -> Result<(), RuntimeError> {
Ok(())
}
}
}
The change is source-breaking in type-system terms, but its known migration set
is currently empty. A workspace-wide source audit finds no impl Module for ...
and no .with_module(...) call site. This is the lowest-cost point to correct
the signature. Phase 4 must repeat the audit in external consumers; if it finds
an implementation, that repository is named and migrated explicitly. The
design does not assume an external coordination cost without such evidence.
Making cleanup real
Today the modules vector is always empty, so the CleaningUp loop has no
participants. Most resources currently rely on Rust Drop behavior, including
application state that owns database pools; other tasks are aborted or left to
process teardown. Drop is useful as a final safety net but does not provide an
awaited pool close, durable flush, checkpoint acknowledgement, or observable
deadline outcome.
Adopting lifecycle participants for durable cleanup is in scope for this
design, not a prerequisite assumed to exist. Lifecycle registration is
transport-neutral. Add LifecycleRegistry, a cloneable registration-only
LifecycleRegistrar, and one object-safe participant contract to
light-runtime:
#![allow(unused)]
fn main() {
#[async_trait]
pub trait LifecycleParticipant: Send + Sync {
fn name(&self) -> &'static str;
async fn shutdown(
&self,
config: &RuntimeConfig,
context: &ShutdownContext,
) -> Result<(), RuntimeError>;
}
impl LifecycleRegistrar {
pub fn register(
&self,
participant: Arc<dyn LifecycleParticipant>,
) -> Result<(), RuntimeError>;
}
}
Participant names are unique within a runtime; duplicate registration is a startup error. The registrar can add a participant but cannot enumerate, invoke, or seal the set. The registry invokes participants sequentially in reverse registration order, which is also reverse resource-construction order. Every hook is attempted even after an earlier error, and the runtime returns an aggregate error after the bounded sequence. Phase 1a does not run hooks in parallel and does not add dependency declarations; a later optimization may add explicit parallel groups without changing the default ordering.
The runtime wraps each builder-supplied Arc<dyn Module> in an internal
ModuleParticipantAdapter; name() delegates to the module and shutdown()
calls its deadline-aware on_shutdown(). Module does not extend
LifecycleParticipant. Modules are inserted at their construction position in
the same registry, while application-owned resources implement
LifecycleParticipant directly.
Transport handles themselves retain explicit transport ownership and are not
also registered as participants, which prevents double shutdown.
The reviewed initial ownership inventory is:
| Owner | Resource | Shutdown owner | Delivery phase |
|---|---|---|---|
light-runtime | portal-registry socket, terminal state, reconnect task, and join handle | RegistrySession::shutdown before transport drain | Phase 1b |
light-axum | listener handle and server task | AxumBoundHandle through TransportRuntime::stop | Phase 1c |
light-pingora | controlled-shutdown sender and Pingora server thread | PingoraBoundHandle through TransportRuntime::stop | Phase 1c |
| builder modules | resources explicitly owned by each module | reverse-order lifecycle participant | Phase 1a and consumer migration |
| application pools, buffers, leases, and task supervisors | resource identified in that application’s migration inventory | application lifecycle participant | Phases 2 and 3 |
Each application migration PR must add a checked inventory table naming every
pool, durable buffer, lease owner, and spawned-task supervisor and either name
its participant or state why synchronous Drop is sufficient. Phase 1a is not
blocked on undiscovered application resources, and a later application phase
cannot claim completion without its reviewed table.
The initial application migration inventory is:
| Service | Owned asynchronous resource | Shutdown ownership |
|---|---|---|
light-agent | SQLx application pool | light-agent-database participant closes and awaits the pool |
light-knowledge | SQLx application pool | light-knowledge-database participant closes and awaits the pool |
light-workflow | SQLx pool; consumer, executor, reconciler, rule API, scheduler, lease, fixed-action, and retention tasks | light-workflow-database participant closes the pool after the task supervisor cooperatively cancels and joins every task; abort is deadline-only |
light-gateway | transport-owned listener/server thread; in-memory configuration and bounded caches | transport stop owns the thread; cache/configuration owners require only synchronous Drop |
light-deployer | transport-owned listener/server task; in-memory service state | transport stop owns the task; service state requires only synchronous Drop |
light-workflow-runner | execution supervisor, transport, health/watchdog/reconciler tasks, SQLite journal | its standalone shutdown path drains the supervisor and transport, joins or deadline-aborts tasks, and returns failure on timeout; SQLite cleanup is synchronous Drop |
light-knowledge-worker | command-scoped SQLx pool and bounded command tasks | its standalone shutdown path bounds each command and awaits PgPool::close; command tasks do not outlive the selected command |
Both transport construction paths receive the registrar alongside
&RuntimeConfig:
#![allow(unused)]
fn main() {
pub trait TransportRuntime {
async fn bind(
&self,
config: &RuntimeConfig,
lifecycle: &LifecycleRegistrar,
admission: &AdmissionGate,
startup_cancel: CancellationToken,
) -> Result<BoundTransport<Self::Handle>, RuntimeError>;
async fn stop(
&self,
handle: &mut Self::Handle,
context: &ShutdownContext,
) -> Result<(), RuntimeError>;
}
pub trait PingoraApp: Send + Sync + 'static {
type Proxy: ProxyHttp + Send + Sync + 'static;
fn proxy(
&self,
config: &RuntimeConfig,
lifecycle: &LifecycleRegistrar,
admission: &AdmissionGate,
) -> Result<Self::Proxy, RuntimeError>;
}
}
Light Axum’s ServerContext re-exposes clones of the light-runtime registrar
and admission gate to AxumApp::router(). It does not own or define either
contract. Light Pingora passes the same values to PingoraApp::proxy(), which
closes the construction-order gap for light-gateway proxies that create durable
buffers, pools, or clients. Standalone applications can construct the same
light-runtime registry and gate directly without depending on either framework
context type.
The successful startup publication order is fixed: run all on_ready hooks
while admission remains closed, seal the registry, transition the state to
Ready, and open admission as the final synchronous step. Registration after
sealing is an error. Startup cancellation seals the registry against new
participants before invoking the already-registered participants’ abort
cleanup.
Each participant owns its concrete resource and implements the deadline-aware
hook. Standalone applications use the same ShutdownContext and participant
contract even when they do not use LightRuntimeBuilder. The phase is complete
only when the resource inventory is explicit; an empty module loop is not
accepted as successful cleanup.
Hooks must use context.remaining() for their own I/O bounds, observe
context.cancellation, and leave durable work committed, checkpointed, or
recoverable before returning. The top-level timeout_at remains the hard
backstop for defective or legacy components. At expiry it may cancel a
mid-flight hook; the nonzero process exit and component timeout telemetry make
that failure explicit rather than reporting a graceful stop.
A participant is registered only after its owned resource is internally
consistent. It must handle both Graceful and StartupAbort; the latter may be
called before the overall service reaches readiness. Emergency permits only
the prearranged bounded cleanup described by the mandatory floor.
Runtime shutdown ordering
RunningRuntime::shutdown() should perform a single deadline-aware sequence:
- create the absolute graceful and hard deadlines and shared
ShutdownContext - transition runtime state to
Quiescing, mark readiness false, and close the admission gate synchronously - atomically put the registry session in terminal mode so it can never reconnect
- send an explicit bounded deregistration/goodbye on the current WebSocket, wait for acknowledgement, close the socket, and join the reconnect task
- ask the transport to stop accepting connections and drain existing work
- invoke deadline-aware module hooks with the same context
- log the duration and return the structured outcome; the production
run_until_shutdownboundary enforces process exit on expiry
The current registration_task.abort() is not sufficient. RegistrySession
uses an atomic Running -> Stopping -> Stopped state. shutdown(context) wins
the Running -> Stopping transition before sending anything. The reconnect
loop observes Stopping in connection attempts, the active connection loop,
and retry sleeps. It may finish the current goodbye exchange, but after that
connection ends it exits instead of sleeping or registering again. Concurrent
or repeated shutdown calls join the same terminal operation.
The codec-neutral logical request is frozen as:
{
"jsonrpc": "2.0",
"id": "shutdown-generated-request-id",
"method": "service/deregister",
"params": {
"runtimeInstanceId": "019...",
"reason": "terminate"
}
}
The successful result is:
{
"runtimeInstanceId": "019...",
"status": "deregistered"
}
reason uses the lowercase shutdown reason names interrupt, terminate, or
programmatic. The controller rejects a runtimeInstanceId that does not
match the authenticated session with JSON-RPC -32602. For the negotiated
binary profile, add ClientGoodbyeV1 { request_id, runtime_instance_id, reason } and ServerGoodbyeV1 { request_id, runtime_instance_id } to
controller-wire; assign new message-kind values without changing any existing
v1 discriminant. The legacy JSON and binary adapters map to the same
SessionInput::Deregister and SessionOutput::Deregistered values.
Controller handling is idempotent for a repeated request on the same session. On the first valid request it marks the session terminal, removes the instance only when the connection id still matches, fails pending commands, emits the discovery and MCP removal notifications, records the disconnect event, and then queues the acknowledgement. The route must flush that acknowledgement before sending the WebSocket close frame; it cannot abort the writer task first. The existing connection-id comparison remains the stale-socket protection. A normal socket close without goodbye continues to use the same cleanup routine, so an old client remains safe.
RegistrySession::shutdown returns Acknowledged, Disconnected, or
TimedOut. Only Acknowledged proves the controller removed the instance
before transport drain. Disconnected and TimedOut are logged and shutdown
continues because the controller’s ordinary socket cleanup remains the
fallback. The operation’s bound is
min(context.remaining(), registration_timeout), where the existing builder
registration timeout defaults to five seconds. For the normal two-second
shutdown setting, the remaining application deadline is therefore the tighter
bound. Deregistration never creates an additional deadline.
Errors from one cleanup hook must not silently prevent the remaining hooks from running. The runtime should collect cleanup failures and return a combined error after all bounded cleanup attempts finish.
Framework Integration
Light Axum
AxumTransport should store the configured shutdown duration in its bound
handle and pass it to axum_server::Handle:
#![allow(unused)]
fn main() {
handle.graceful_shutdown(Some(Duration::from_millis(
shutdown_graceful_period,
)));
}
The listener stops accepting new connections when transport drain starts,
after bounded deregistration. From the earlier admission-close boundary until
then, new application requests receive the defined 503. Existing accepted
connections drain until they complete or the deadline expires. With no active
connections, the server task should join immediately.
The transport receives the shared absolute deadline and derives its remaining
duration immediately before calling the handle. It must not restart the full
configured period after deregistration. The top-level runtime timeout remains
the enforcer if Handle or its task join fails to return.
Tests must include ordinary requests, streaming bodies, keep-alive connections, and WebSockets. A connection being idle must not hold shutdown open indefinitely.
Light Pingora
PingoraTransport uses a controlled shutdown channel. The shared runtime first
closes admission and drains the gateway’s application permits against the
absolute shutdown deadline. Pingora’s internal graceful timeout is therefore
zero: it must not restart the original configured period after the shared drain.
The transport joins the Pingora thread using ShutdownContext::remaining(). The current
implementation uses the crates.io 0.8.1 release. That release contains a
redundant sleep after Runtime::shutdown_timeout for nonzero internal timeout
values. Light-Fabric does not exercise that path: the shared admission gate
owns application draining, and both Pingora internal shutdown periods are set
to zero before the server starts. The redundant upstream sleep is therefore
zero-duration, while the outer thread join remains bounded by the shared
absolute deadline. No vendored Pingora source or Cargo override is required.
Returning Pingora’s FastShutdown is not an acceptable normal-path workaround
because it forfeits request draining.
The migration must verify that:
- the controlled signal stops listener acceptance immediately
- with no active downstream exchange, transport stop completes in less than one second
- active proxy requests can finish inside the deadline
- WebSockets and streaming exchanges cannot exceed the deadline
- the Pingora thread is joined before module cleanup completes
For upstream pingora-core 0.8.1, Some(0) is a zero-duration runtime shutdown;
it does not mean wait forever. A transport configuration test must pin both
zero values so a dependency upgrade cannot silently restart an internal grace
period after the shared drain.
The separate grace_period_seconds setting is equally load-bearing. Pingora
performs another unconditional sleep before runtime shutdown and defaults a
missing value to EXIT_TIMEOUT, currently five minutes. light-pingora
explicitly sets grace_period_seconds = Some(0); Phase 1c must preserve that
assignment and pin it in the same configuration and latency tests. Removing it
must fail a test rather than turn a normal stop into a five-minute pre-drain
sleep.
Standalone Axum services
Services that do not use LightRuntimeBuilder must still use the shared signal
API. They should create an axum_server::Handle, run the server and shutdown
future concurrently, then invoke graceful_shutdown with their configured
deadline.
Migration to LightRuntimeBuilder is preferred when it does not introduce an
unrelated architectural change, but signal correctness must not wait for that
migration.
Background workers
Long-running loops must accept a CancellationToken or equivalent cancellation
receiver. On shutdown they must stop claiming new work before joining already
spawned tasks.
Each worker must classify its work as one of:
- drain: finish the current unit inside the deadline
- checkpoint: persist progress and release the unit for retry
- cancel: abort work that is side-effect free or transactionally safe
Dropping a Tokio JoinHandle does not stop its task and is not an acceptable
shutdown implementation. Every service must cancel, join, or deliberately
abort each owned task.
light-workflow needs a service-level cancellation token shared by its event
consumer, executor, reconcilers, rule API, scheduler, reaper, and fixed-action
workers. Lease-backed work must be released or left in a state that its normal
reconciliation policy can recover.
Existing prior art
apps/light-workflow-runner/src/main.rs is the closest in-repository precedent.
After Ctrl-C it calls supervisor.drain(), publishes shutdown through a watch
channel, and wraps the transport join in
timeout(config.shutdown_grace, transport). Its application-prefixed
shutdownGraceMs demonstrates the intended supervisor-drain-bound shape.
It is a starting point, not yet the completed contract: it handles only Ctrl-C, installs no eager watcher, discards the timeout result, does not use the shared exit policy, and does not account for every spawned health/watchdog task. Its migration should preserve the drain-first behavior while adopting the shared signal, context, deadline outcome, and task-ownership rules.
Container And Orchestrator Contract
The image and deployment contract remains SIGTERM. A service-specific
stop_signal: SIGINT may be used only as a temporary compatibility measure
while an older image is being migrated.
The outer stop timeout must be greater than the application graceful period:
orchestrator stop timeout
>= application graceful period
+ mandatory cleanup floor
+ scheduling allowance
Deregistration, drain, and normal asynchronous cleanup all share the configured application graceful period. Only the fixed emergency cleanup floor sits outside it, so ordinary cleanup is not double-counted.
If termination arrives during startup, cancellation begins immediately and the service gets only the mandatory cleanup floor; it does not finish the remaining bootstrap or registration timeout and does not add a graceful drain period. Therefore the normal running-service inequality above is also the worst-case bound for startup termination. If a future startup phase cannot honor cancellation, its maximum remainder must be added explicitly to this inequality until that phase is repaired.
For a two-second application grace period, a 5-10 second container timeout is
usually sufficient. Retaining a 30-second Compose timeout is also safe once the
application handles SIGTERM, because the engine stops waiting as soon as the
process exits.
Kubernetes deployments should set terminationGracePeriodSeconds using the
same inequality. A future preStop hook must not duplicate the application
grace period or merely sleep.
If the application deadline expires during a deliberate Pod deletion, the
container’s terminated state records exit code 1. That is the expected
representation of an application-level graceful-shutdown failure, not exit
code 137 from an orchestrator kill. Kubernetes dashboards and alerts must
correlate the nonzero exit with Pod deletion/termination context: record and
trend it, but do not page solely on that exit code during an intentional
rollout. Repeated deadline expiry or exit 1 outside termination remains
actionable.
As a preventive rule, shell entrypoints must end with exec so the Rust binary
becomes PID 1. If an init or wrapper process is required, it must forward
SIGTERM and reap child processes. This is not a currently identified
light-fabric image defect: in-repository Dockerfiles use exec-form
CMD/ENTRYPOINT, and the workflow runner uses tini -- in
apps/light-workflow-runner/docker/Dockerfile.
Configuration
The existing setting remains canonical:
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
No separate shutdownSleep, Compose-specific, Axum-specific, or
Pingora-specific duration should be introduced.
Validation rules:
- warn when a value is zero outside tests
- retain millisecond precision in the global runtime deadline
- round a positive Pingora component timeout up to a whole second and log the configured and effective values
- log the configured and effective duration at startup
- never log that shutdown was graceful if the deadline expired
Task-oriented applications not yet using ServerConfig may temporarily expose
an application-prefixed duration, but they should converge on the shared
setting when adopting light-runtime lifecycle management.
light-workflow-runner currently rejects shutdownGraceMs: 0, while the shared
server contract permits zero with an emergency cleanup floor. That stricter
runner rule is deliberate for a lease-owning worker: it must preserve some
cooperative drain/checkpoint opportunity. Convergence means sharing signal,
deadline, outcome, and outer-timeout semantics; it does not require every
service class to permit zero. The runner validator and example documentation
must state this exception explicitly.
Observability
Emit structured events for:
- accepted signal and shutdown reason
- transition into each shutdown phase
- number of active requests, streams, sockets, and worker tasks at drain start
- configured deadline and remaining time
- component completion or timeout
- total shutdown duration
- final graceful, deadline-exceeded, or failed outcome
Active-work cardinality is not currently available for free from
axum_server or Pingora. Reporting active requests, streams, sockets, and
worker tasks requires framework-owned admission/in-flight counters that are
incremented before dispatch and decremented by a drop guard. Counter delivery
is implementation work in the framework phases, not merely a logging change.
Recommended metric families are:
service_shutdown_total{reason,outcome}service_shutdown_duration_secondsservice_shutdown_active_work{kind}service_shutdown_component_duration_seconds{component}
Normal termination should exit with code zero. Startup failure, cleanup
failure, and graceful-deadline expiry exit nonzero; deadline expiry specifically
uses exit code 1. Exit code 137 in the container qualification test indicates
forced SIGKILL and fails the graceful-shutdown gate.
Migration Plan
Phase 1a: Shared primitives and compile surface
- add eagerly installed
ShutdownWatcher,ShutdownReason,ShutdownContext,AdmissionGate,LifecycleRegistry, andLifecycleRegistrartolight-runtime - make the deadline-aware
Module::on_shutdownchange and add the fixed reverse-registrationLifecycleParticipantbehavior - update all five known
TransportRuntimeimplementations in the same change:AxumTransport,PingoraTransport, thelight-runtimetest transport, and the headless transports inlight-workflowandlight-knowledge-worker; the headless implementations may ignore registrar/admission arguments until their Phase 3 behavioral migration, but the workspace must remain compiling - pass registrar and admission capabilities through
ServerContextandPingoraApp, update the in-repositoryGatewayAppimplementation in the same compile change, invokeon_ready, seal lifecycle registration, transition toReady, and open admission in the specified order - add cancellation-aware startup phases, binding ownership guards,
StartupGuard, andLightRuntime::run_until_shutdown; preserveRunningRuntime::shutdown()throughShutdownReason::Programmatic - add signal, admission, startup-abort, lifecycle-order, aggregate-error, and public-API tests
Phase 1b: Registry terminal protocol
This is an explicitly coordinated light-fabric plus controller-rs phase,
not deferred external adoption:
- add the codec-neutral deregister values and append-only v1 goodbye message
kinds to
controller-wire, including legacy JSON, rkyv, golden-fixture, and invalid-instance tests - split registry startup into owned
RegistrySessioncreation followed by a cancellable registration wait - implement terminal/no-reconnect state, bounded goodbye, socket close, and
task join in
portal-registry - implement idempotent connection-matched removal, acknowledgement flush, and
close ordering in
controller-rs - run cross-repository tests for acknowledged shutdown, stale sockets, disconnect fallback, timeout, startup abort, and proof that no registration occurs after terminal state is entered
Phase 1c: Runtime transports
- pass the shared remaining deadline into
TransportRuntime::stop - apply admission and the remaining deadline in
light-axum - upgrade the Pingora crate family to the crates.io
0.8.1release and pin Pingora’s internal grace and runtime-shutdown periods to zero - verify Pingora rounded timeout,
grace_period_seconds = Some(0), zero-duration behavior, active drain, and no-load fast return - add Axum and Pingora request, stream, WebSocket, join, and global-backstop integration tests
Phase 2: Light Runtime applications
Replace app-local ctrl_c() handling in:
light-gatewaylight-agentlight-knowledgelight-deployer- any example application built with
LightRuntimeBuilder
Each migration must prove both SIGINT and SIGTERM paths.
Phase 3: In-repository standalone services and workers
Adopt the shared signal API and explicit drain behavior in:
light-workflowlight-workflow-runner, preserving its existingsupervisor.drain()/watch/timeout structure as prior art; replace theshutdownGraceMs: 30000examples with a validated value strictly below the external 30-second container timeout after reserving the 250 ms emergency floor and scheduling allowancelight-agent-channel, including its HTTP server and three spawned delivery, trigger, and attachment-recovery loopslight-github-action-providerlight-knowledge-workerbuild and projection loops; its shippedconfig/server.ymlalready declaresshutdownGracefulPeriod: 2000, which currently is not consumed by the worker- long-running Rust example and MCP-server applications that ship from this workspace
light-agent-worker is deliberately excluded because its stdio lifecycle ends
on EOF from its supervisor. If it becomes an independently orchestrated
long-running service, it enters this contract.
Phase 4: External repository adoption
The following services are not implemented in light-fabric, so this phase
cannot be marked complete by a light-fabric change alone:
controller-rsin the externalcontroller-rsrepositoryportal-service,config-server, andlight-oauthin the externalportal-servicerepository- demo APIs and MCP servers in the external
light-example-rsrepository
These repositories consume the exported signal API without moving their
application code. Phase 1b already delivers the controller-rs deregistration
protocol; this phase migrates the controller process’s own server lifecycle.
Phase 5: Deployment qualification
- keep the current outer timeout as a safety boundary
- publish images containing the signal-handling changes
- recreate containers so the new images are active
- measure no-load and in-flight shutdown behavior under Docker and Podman
- reduce local outer timeouts only if faster forced-failure feedback is useful
- align Kubernetes termination grace values with the proven application bound
Verification Strategy
Signal tests
Use a subprocess fixture rather than sending termination signals to the test runner itself. The fixture must:
- report ready only after both handlers are installed
- exit zero after
SIGINT - exit zero after
SIGTERMon Unix - record exactly one accepted shutdown reason
- collapse the remaining drain budget after a second accepted signal
- prove a signal delivered between watcher installation and runtime readiness immediately cancels startup rather than waiting for startup completion or taking the default signal disposition
- document and test that watcher installation outside a Tokio reactor is an
invalid call, while the first statement inside
#[tokio::main]succeeds
Startup cancellation tests
Inject a controllable future into each startup phase and deliver SIGTERM
while it is pending:
- remote bootstrap is cancelled without publishing a partial cache file
- a bound Axum or Pingora listener is closed by
StartupGuard - partial controller registration is explicitly closed or deregistered
- lifecycle participants registered before cancellation receive startup-abort cleanup, while later registration is rejected after sealing
- the process exits zero within the 250 ms floor when cleanup cooperates
- a stuck startup resource triggers exit code
1at the floor rather than waiting for the five-second registration timeout - a startup signal and readiness completion in the same scheduler turn choose cancellation because the supervisor select is biased toward the signal
Transport tests
For both Axum and Pingora:
- with no active request, transport stop completes in less than one second
- a request completing inside the deadline returns its normal response
- a request exceeding the deadline is terminated at the bound
- new application requests receive
503as soon as quiescing begins - new TCP connections are refused after bounded deregistration starts transport drain
- streaming and WebSocket connections obey the bound
- module shutdown hooks run after listener quiescence
- the whole runtime reaches its explicit deadline outcome even if transport stop or an underlying thread join does not cooperate
The Pingora no-load assertion is a release gate, not an aspirational timing
description. The shared admission drain and remaining-budget join must prevent
a dependency upgrade from introducing a fixed poll or sleep near one second.
Configuration tests assert both
grace_period_seconds == Some(0) and
graceful_shutdown_timeout_seconds == Some(0).
Module and deregistration tests
- every module receives the same absolute deadline and cancellation token
- participants run sequentially in reverse registration order; duplicate names fail startup and one hook error does not skip later hooks
- the production cleanup participant registry matches the reviewed resource inventory and is empty only for an explicitly resource-free service
- a cooperative hook checkpoints before cancellation and returns
- a deliberately stuck hook triggers the global backstop and exit-code-1 path
- cleanup errors are aggregated without skipping later bounded hooks
- the runtime closes admission before sending deregistration
- entering registry terminal state before goodbye prevents reconnect during send, acknowledgement, close, retry sleep, and concurrent shutdown calls
- legacy JSON and binary goodbye requests map to the same logical operation
- a mismatched runtime instance id fails closed and a stale connection cannot remove the replacement instance
- the controller removes the instance from routing and flushes acknowledgement before the client closes and transport drain starts
- an already disconnected socket returns the fallback outcome without trying to reconnect
- deregistration failure consumes only its share of the global remaining time
- Axum and Pingora resources created during router/proxy construction register
through the same light-runtime lifecycle registry and are sealed at
Ready
Worker tests
- cancellation prevents new work from being claimed
- drainable work completes inside the deadline
- retryable work preserves or releases its durable lease correctly
- all owned tasks are joined, cancelled, or explicitly aborted
- database pools and durable buffers close without data loss
- runner example grace values leave the documented emergency and scheduling margin below the deployment’s outer stop timeout
Container qualification
Run each production image as PID 1, wait for readiness, send SIGTERM, and
assert:
- the service logs receipt of
Terminate - no-load shutdown completes within the target fast-path threshold
- in-flight work follows its drain policy
- a normal drain exits zero before the engine timeout
- a deliberate application deadline expiry exits
1, not0or137 - container inspection does not report an out-of-memory kill or exit code 137
Run the matrix with Docker and Podman because signal forwarding and wrapper entrypoints are deployment concerns, not only Rust unit-test concerns.
Acceptance Criteria
The design is complete when:
- every production Rust service handles orchestrator
SIGTERM - signal handlers are installed before readiness is published
- a signal received during startup cancels and unwinds startup within the mandatory cleanup floor rather than waiting for bootstrap or registration
- no-load shutdown normally completes in less than one second
- application admission closes synchronously and returns the defined
503before deregistration performs network I/O; TCP refusal follows bounded deregistration when transport drain starts - in-flight HTTP work drains up to
server.shutdownGracefulPeriod - long-lived streams and background workers cannot delay exit beyond the bound
SIGINTremains functional for interactive development- a second signal skips remaining drain and enters mandatory cleanup
- registry terminal state is entered before controller goodbye, no reconnect or re-registration occurs afterward, and deregistration is acknowledged or bounded before transport drain
- every inventoried durable resource owner is registered as a cleanup participant; an empty set is accepted only when the service inventory explicitly proves it owns no asynchronous cleanup
- lifecycle participants execute in deterministic reverse registration order and cleanup errors are aggregated
- the Pingora dependency resolves to crates.io
0.8.1, both internal shutdown periods remain zero, and the outer join observes the shared remaining budget - application logs distinguish graceful exit from deadline expiry
- normal container tests prove exit code zero without forced termination;
deadline-expiry tests prove exit code
1 - Compose and Kubernetes retain an outer timeout greater than the application deadline
Operational Guidance
If a service consistently consumes the full container stop timeout, treat it as a shutdown defect. Check, in order:
- whether the Rust binary is PID 1 or receives forwarded signals
- whether it installed a
SIGTERMhandler - whether the shutdown path stopped listener acceptance
- which request, stream, task, or cleanup hook remains active
- whether the application deadline is actually connected to that component
Lowering the container timeout can shorten the symptom, but it does not repair the lifecycle. The correct steady state is a cooperative application that exits as soon as its real work is safe, with the orchestrator deadline unused during normal shutdown.
Release Workflow
Light-Fabric already has a release.sh script that builds Linux binaries,
packages release archives, and creates or updates a GitHub release. The current
release page uses a static note string, so operators can download artifacts but
cannot easily see what changed between tags.
This design introduces a cascading polyrepo release orchestrated by light-workflow.
It automates release-notes, changelog flow, binary generation, Docker image pushes,
and downstream dependency propagation across both public (light-fabric, light-example-rs)
and private (controller-rs, portal-service) repositories.
The implementation should start with a small dependency-free git-log script and leave room to adopt a more structured changelog generator later. It should also centralize Docker image publishing so binary archives and container images use the same release version.
Goals
- Generate release notes from commits between the previous release tag and the current release tag.
- Use the same generated notes for GitHub release creation and release updates.
- Maintain a checked-in
CHANGELOG.mdso release history is visible without opening GitHub. - Preserve the current
release.sh VERSION [-l|--local] [--skip-build]operator workflow. - Keep binary and Docker release commands independent initially, while allowing them to converge on the same version tag and compiled Linux binaries later.
- Support Apple Silicon and Windows binary artifacts through CI runners that match those operating systems.
- Add one repo-root
build.shfor all Docker images while preserving app-level build script compatibility. - Allow manual edits before publishing when release notes need customer-facing cleanup.
- Avoid requiring Conventional Commit messages on day one.
Non-Goals
- Replace GitHub releases as the artifact distribution point.
- Require every commit message to follow
feat:,fix:, or another convention immediately. - Generate perfect marketing release notes without review.
- Upload changelog files as separate release artifacts.
- Remove existing app-level
build.shentrypoints immediately. - Build macOS binaries from a normal Linux Docker builder. Apple toolchains and SDKs require a macOS build runner.
- Build Windows MSVC binaries from a normal Linux Docker builder. Use a Windows runner for the official Windows artifacts.
- Publish Windows container images as part of the first release flow. Windows container images require Windows base images and a Windows container builder.
Current State
release.sh currently performs these steps:
- Parse release options and target version.
- Build
light-agent,light-deployer,light-gateway,light-workflow,light-workflow-runner,light-knowledge, andlight-knowledge-workerfor Linux GNU and Linux musl targets. - Package the binaries into
dist/light-fabric-${VERSION}-${TARGET}.tar.gz. - If
--localis not set, create a GitHub release or upload artifacts to an existing release.
When creating a new GitHub release, the script uses a static note body:
Light-Fabric Linux release binaries
When the release already exists, the script uploads artifacts but does not update the release notes.
Docker image builds are handled by the repo-root build.sh. App-level scripts
remain as compatibility wrappers:
apps/light-agent/build.sh
apps/light-deployer/build.sh
apps/light-gateway/build.sh
apps/light-workflow/build.sh
apps/light-workflow-runner/build.sh
apps/light-knowledge/build.sh
apps/light-knowledge-worker/build.sh
The root script uses this shape:
./build.sh 0.3.0
./build.sh 0.3.0 --local
./build.sh 0.3.0 --no-cache
It builds and optionally pushes networknt/<app>:${VERSION} and
networknt/<app>:latest for an explicit release-app allowlist. All app-level
entrypoints delegate to the same implementation.
release.sh intentionally remains binary-only. Its version and the Docker image
version may differ during the current transition; the future orchestrated flow
can pass the same version to both commands.
Options
Option 1: GitHub Generated Notes
GitHub CLI can generate release notes:
gh release create "$VERSION" --generate-notes --notes-start-tag "$PREVIOUS_TAG"
This is the least code, and it works well for the GitHub release page. The
tradeoff is that it does not update CHANGELOG.md in the repository unless an
additional script calls the GitHub API and copies the generated notes back into
the repo.
This option is useful as a fallback, but it should not be the primary design if the repo changelog is a required output.
Option 2: Dependency-Free Git-Log Script
A local script can generate release notes from the git history:
git log "${PREVIOUS_TAG}..${TARGET_REF}" --pretty=format:"- %s (%h)"
The script can write a markdown file and use that same file for both
CHANGELOG.md and gh release create --notes-file.
This option is simple, reviewable, and fits the current Bash release script. It does not require new tooling or commit-message conventions. The initial output will be commit-oriented rather than category-oriented, but it can be improved incrementally.
Option 3: git-cliff
git-cliff can generate structured changelogs from Conventional Commit
messages and custom templates. It can group entries into sections such as
features, fixes, documentation, and breaking changes.
This gives the best long-term release notes, but it adds a release-tool dependency and works best only after the team consistently writes conventional commit messages.
This can be adopted later without changing the overall release flow: replace the
internal git-log generator with a git-cliff invocation that writes the same
release notes file.
Proposed Design
Start with Option 2.
Add a helper script:
scripts/release-notes.sh
The script should generate:
dist/release-notes-${VERSION}.md
It should optionally update:
CHANGELOG.md
release.sh should call the helper before publishing the GitHub release. The
generated notes file becomes the release page source:
gh release create "$VERSION" "${ARCHIVES[@]}" \
--title "$VERSION" \
--notes-file "$NOTES_FILE"
For an existing release, the script should update the release body as well as uploading artifacts:
gh release edit "$VERSION" --notes-file "$NOTES_FILE"
gh release upload "$VERSION" "${ARCHIVES[@]}" --clobber
Use Docker as the official Linux release builder. The controlled Docker builder
environment should compile Linux binaries once per Linux platform, export those
binaries into dist/, and use the same binaries when assembling runtime Docker
images. Local host builds remain useful for development, but they should not be
the official release source for Linux artifacts.
Add a repo-root Docker image script:
build.sh
The root script should become the source of truth for building and publishing all Light-Fabric app images:
./build.sh 0.3.0
./build.sh 0.3.0 --local
./build.sh 0.3.0 --app light-agent
./build.sh 0.3.0 --app light-agent --no-cache
./build.sh 0.3.0 --skip-latest
The script should build these images by default:
networknt/light-agent:0.3.0
networknt/light-deployer:0.3.0
networknt/light-gateway:0.3.0
networknt/light-workflow:0.3.0
networknt/light-workflow-runner:0.3.0
networknt/light-knowledge:0.3.0
networknt/light-knowledge-worker:0.3.0
Unless --skip-latest is set, it should also tag and push:
networknt/light-agent:latest
networknt/light-deployer:latest
networknt/light-gateway:latest
networknt/light-workflow:latest
networknt/light-workflow-runner:latest
networknt/light-knowledge:latest
networknt/light-knowledge-worker:latest
Existing app-level build scripts should remain, but they should become thin wrappers around the root script:
../../build.sh "$@" --app light-agent
This preserves established operator muscle memory and removes duplicated Docker publish logic.
release.sh should call the root build.sh with the same VERSION. For Linux
targets, the release should build once per platform and reuse the output:
Docker/BuildKit Linux builder
|
+-- dist/linux/<target>/bin/<app> -> GitHub release tarballs
|
+-- dist/linux/<target>/bin/<app> -> Docker runtime images
This makes one command release both binary artifacts and Docker images without compiling the same Linux binaries twice.
Changelog Format
CHANGELOG.md should use reverse chronological release sections:
# Changelog
## 0.3.0 - 2026-06-03
- Add JSON file logging support to `light-runtime` (abc1234)
- Wire runtime logging control into `light-gateway` (def5678)
- Document Splunk ingestion options for tracing (123abcd)
## 0.2.0 - 2026-05-20
- ...
The generated release notes file should contain the same section body:
## 0.3.0 - 2026-06-03
### Changes
- Add JSON file logging support to `light-runtime` (abc1234)
- Wire runtime logging control into `light-gateway` (def5678)
- Document Splunk ingestion options for tracing (123abcd)
### Artifacts
- `light-fabric-0.3.0-x86_64-unknown-linux-gnu.tar.gz`
- `light-fabric-0.3.0-x86_64-unknown-linux-musl.tar.gz`
- `light-fabric-0.3.0-aarch64-unknown-linux-gnu.tar.gz`
- `light-fabric-0.3.0-aarch64-unknown-linux-musl.tar.gz`
- `light-fabric-0.3.0-aarch64-apple-darwin.tar.gz`
- `light-fabric-0.3.0-x86_64-pc-windows-msvc.zip`
- `networknt/light-agent:0.3.0`
- `networknt/light-deployer:0.3.0`
- `networknt/light-gateway:0.3.0`
- `networknt/light-workflow:0.3.0`
- `networknt/light-workflow-runner:0.3.0`
- `networknt/light-knowledge:0.3.0`
- `networknt/light-knowledge-worker:0.3.0`
The release notes file can include artifact names because it is used directly
for the GitHub release page. CHANGELOG.md should focus on changes and can
omit artifact details.
Docker images should be listed in the GitHub release body even though they are published to Docker Hub instead of attached to the release page. This gives operators one place to see every artifact produced by a release.
Docker image platform variants should also be visible:
networknt/light-agent:0.3.0 linux/amd64, linux/arm64
networknt/light-deployer:0.3.0 linux/amd64, linux/arm64
networknt/light-gateway:0.3.0 linux/amd64, linux/arm64
networknt/light-workflow:0.3.0 linux/amd64, linux/arm64
networknt/light-workflow-runner:0.3.0 linux/amd64, linux/arm64
networknt/light-knowledge:0.3.0 linux/amd64, linux/arm64
networknt/light-knowledge-worker:0.3.0 linux/amd64, linux/arm64
Tag Range Selection
The release-notes script needs a deterministic commit range.
Inputs:
VERSION: target tag, for example0.3.0orv0.3.0- optional
--from PREVIOUS_TAG - optional
--target-ref TARGET_REF
Default behavior:
- Unless
--fromis supplied, fetch tags fromoriginand stop if they cannot be synchronized. - If
--target-refis supplied, use it as the end of the range. - Else if the
VERSIONtag exists locally, useVERSION. - Else use
HEAD. - If
--fromis supplied, use it as the start of the range. - Else find the newest semver-like tag before
VERSION. - If no previous tag exists, use the first commit as the start.
For existing releases, this allows regenerating the notes for the exact tag. For new releases, this allows generating notes before the tag exists.
Recommended git command:
git log --no-merges --pretty=format:"- %s (%h)" "${PREVIOUS_TAG}..${TARGET_REF}"
If merge commits are important for the team, the script can add a
--include-merges option.
Release Script Flow
The binary-only release.sh flow is:
- Parse release options.
- Validate build and publish dependencies.
- Generate release notes into
dist/release-notes-${VERSION}.md. - Build Linux GNU and musl binaries unless
--skip-buildis set. - Package binary release archives.
- Print generated archive names and the release-notes path.
- If
--localis set, stop before GitHub publishing. - Create or update the GitHub release and upload archives with
--clobber.
Docker images are a separate operation. The repo-root build.sh builds all
selected images before publishing, pushes versioned tags, and only then pushes
latest tags. The binary version passed to release.sh and the image version
passed to build.sh may differ for now.
The release notes should be generated before publishing, but the changelog
update should be explicit. A release engineer may want to review and commit
CHANGELOG.md before publishing.
Recommended flags:
--update-changelog prepend the generated section to CHANGELOG.md
--notes-only generate notes and optionally update changelog without building
--from TAG override previous tag selection
--target REF override release notes target ref
--include-merges include merge commits in generated commit list
--skip-build package existing binary outputs without rebuilding
--no-target-add do not install Rust targets automatically
--dist DIR write release output to a different directory
--local still builds and packages locally, but it does not call gh.
Automated Polyrepo Release Workflow
Because controller-rs, portal-service, and light-example-rs depend on light-fabric crates, they must be released sequentially in a Cascading Release Pipeline. Attempting to release them manually is error-prone.
We will dogfood light-workflow as our Release Orchestrator to automate this across the public and private repository boundaries.
The Release-Train Workflow Template
The light-workflow template acts as the overarching controller:
-
Step 1: Upstream Release (
light-fabric)- Task A: The workflow runs
cargo release(or equivalent) to bump versions, tag, and publish the publiclight-fabriccrates tocrates.io. - Task B: The workflow invokes the
build.shscript to compile Linux binaries and pushlight-fabricDocker images. - Task C: The workflow calls
release.shto generate the changelog and publish the GitHub Release page.
- Task A: The workflow runs
-
Step 2: The Sync Barrier (Wait Step)
- The workflow pauses for a short duration (e.g., 2 minutes) to ensure
crates.ioindexing has completed, preventing downstream builds from failing to find the new crate versions.
- The workflow pauses for a short duration (e.g., 2 minutes) to ensure
-
Step 3: Downstream Dependency Propagation
- The workflow clones
controller-rs,portal-service, andlight-example-rs. - It runs
cargo update -p light-fabricto point the downstream repositories to the newly published version. - It pushes these changes to their respective
mainbranches.
- The workflow clones
-
Step 4: Parallel Downstream Releases
- The workflow uses a parallel execution pattern to trigger releases for the downstream repositories simultaneously:
- Branch 1 (
controller-rs): Build private binaries, push private Docker images, and tag the private repo. - Branch 2 (
portal-service): Build private binaries, push private Docker images, and tag the private repo. - Branch 3 (
light-example-rs): Publish any downstream public crates, push public Docker images, and create the GitHub Release.
- Branch 1 (
- The workflow uses a parallel execution pattern to trigger releases for the downstream repositories simultaneously:
By wrapping the individual release.sh and build.sh scripts in a light-workflow execution, we gain stateful retries, full pipeline visibility, and automated propagation without exposing secure tokens on developer workstations.
Root Docker Build Script
The repo-root build.sh should own Linux Docker image build and push behavior
for all apps.
Recommended app metadata:
| App | Image | Dockerfile |
|---|---|---|
light-agent | networknt/light-agent | apps/light-agent/docker/Dockerfile |
light-deployer | networknt/light-deployer | apps/light-deployer/Dockerfile |
light-gateway | networknt/light-gateway | apps/light-gateway/docker/Dockerfile |
light-workflow | networknt/light-workflow | apps/light-workflow/docker/Dockerfile |
light-workflow-runner | networknt/light-workflow-runner | apps/light-workflow-runner/docker/Dockerfile |
light-knowledge | networknt/light-knowledge | apps/light-knowledge/docker/Dockerfile |
light-knowledge-worker | networknt/light-knowledge-worker | apps/light-knowledge-worker/docker/Dockerfile |
The Docker build context should remain the workspace root because the
Dockerfiles copy workspace-level Cargo.toml, Cargo.lock, crates,
frameworks, and app directories.
The script should support:
build.sh [VERSION] [-l|--local] [--no-cache] [--app APP] [--skip-latest]
Default behavior:
- Build all explicitly configured release-app images.
- Tag each image as
networknt/${APP}:${VERSION}. - Tag each image as
networknt/${APP}:latestunless--skip-latestis set. - Complete every local build before pushing any image.
- If
--localis set, stop after local image builds. - Otherwise push every versioned tag, followed by every
latesttag.
The script should print the full list of image tags it built and pushed. This
list should be available to release.sh so the GitHub release notes can include
the Docker image artifacts.
When build.sh is called from release.sh, it should receive the exported
binary directory explicitly:
./build.sh "$VERSION" --binary-dir "dist/build"
When build.sh is called directly without --binary-dir, it can either invoke
the Docker release builder for the requested platforms or fall back to the
current Dockerfile builder stages. The preferred direct behavior is to invoke
the same Docker release builder so local and CI image builds stay aligned.
Recommended implementation:
- Add a release builder Dockerfile, for example:
docker/Dockerfile.release
- Add a builder target that compiles all apps for one Linux target and exports binaries:
docker buildx build \
--target export-binaries \
--platform linux/amd64 \
--output type=local,dest=dist/build/linux-amd64 \
.
- Repeat for
linux/arm64if multi-architecture Linux images are enabled. - Package the exported binaries into GitHub release tarballs.
- Build runtime images from those exported binaries, not from another
cargo build.
The runtime image Dockerfiles can use a binary-only context or a release target that copies prebuilt binaries:
COPY dist/build/linux-amd64/bin/light-gateway /app/light-gateway
For multi-platform images, docker buildx build --platform linux/amd64,linux/arm64
can publish one image tag with a manifest list. The important point is that
each platform-specific image must use the binary built for that platform.
Cross-Platform Binary Strategy
“Build once” means build once per target platform, then reuse that output everywhere that platform can run. It does not mean one binary can serve every operating system and CPU architecture.
Recommended artifact matrix:
| Artifact | Target | Builder |
|---|---|---|
| Linux x86_64 binary archive | x86_64-unknown-linux-gnu or x86_64-unknown-linux-musl | Docker/BuildKit Linux builder |
| Linux arm64 binary archive | aarch64-unknown-linux-gnu or aarch64-unknown-linux-musl | Docker/BuildKit Linux builder |
| Linux Docker image for Intel/AMD | linux/amd64 | Docker/BuildKit Linux builder |
| Linux Docker image for Apple Silicon Docker Desktop | linux/arm64 | Docker/BuildKit Linux builder |
| Apple Silicon macOS binary archive | aarch64-apple-darwin | macOS arm64 runner |
| Windows binary archive | x86_64-pc-windows-msvc | Windows runner |
Apple Silicon has two different release meanings:
- Docker image support for Apple Silicon machines is a Linux
arm64container image. Docker Desktop on Apple Silicon runs Linux containers, solinux/arm64is the right image platform. - Native Apple Silicon binaries are macOS binaries targeting
aarch64-apple-darwin. These should be built on a macOS runner, not inside a normal Linux Docker build.
Windows binaries and Windows container images are also separate concerns:
- Windows binary archives should target
x86_64-pc-windows-msvcand should be built on a Windows runner for the official release. - Windows container images require Windows base images and a Windows container builder. They should be treated as a later phase unless customers explicitly need Windows containers.
In CI, these builds can run at the same time as separate jobs:
linux-release:
Docker/BuildKit builds Linux binaries and Linux Docker images.
macos-release:
macOS runner builds aarch64-apple-darwin binaries.
windows-release:
Windows runner builds x86_64-pc-windows-msvc binaries.
The release publish job should collect all artifacts and update the same GitHub release page. Docker Hub publishing should remain in the Linux release job because the Docker images are Linux container images.
CHANGELOG Update Strategy
The changelog update should be idempotent.
Rules:
- If
CHANGELOG.mddoes not exist, create it with# Changelog. - If a section for
VERSIONalready exists, replace that section. - If no section for
VERSIONexists, insert the new section immediately after the# Changelogheading. - Preserve older release sections as-is.
- Never rewrite unrelated content below older release sections.
This makes rerunning the release script safe during release preparation.
Manual Review Workflow
For a normal release:
./release.sh 0.3.0 --notes-only --update-changelog
git diff CHANGELOG.md dist/release-notes-0.3.0.md
The release engineer reviews and edits CHANGELOG.md if needed, commits it,
then publishes:
./release.sh 0.3.0 --skip-build
If binaries also need to be rebuilt:
./release.sh 0.3.0
By default, the official Linux binaries and Linux Docker images should be built from Docker and published together. If a developer needs the old host-build path for local troubleshooting:
./release.sh 0.3.0 --host-build --local
If CI is producing all OS artifacts, the release job should collect the platform-specific archives before publishing:
dist/light-fabric-0.3.0-x86_64-unknown-linux-gnu.tar.gz
dist/light-fabric-0.3.0-aarch64-unknown-linux-gnu.tar.gz
dist/light-fabric-0.3.0-aarch64-apple-darwin.tar.gz
dist/light-fabric-0.3.0-x86_64-pc-windows-msvc.zip
If only Docker images need to be rebuilt and pushed with the same release tag:
./release.sh 0.3.0 --docker-only
If only one Docker image needs to be rebuilt locally:
./build.sh 0.3.0 --app light-gateway --local
If the release page already exists and only the notes need refreshing:
./release.sh 0.3.0 --notes-only
gh release edit 0.3.0 --notes-file dist/release-notes-0.3.0.md
The final implementation can make the last command part of release.sh when
--local is not set.
GitHub Release Body
The GitHub release body should be generated from the same release notes file. For new releases:
gh release create "$VERSION" "${ARCHIVES[@]}" \
--title "$VERSION" \
--notes-file "$NOTES_FILE"
For existing releases:
gh release edit "$VERSION" --notes-file "$NOTES_FILE"
gh release upload "$VERSION" "${ARCHIVES[@]}" --clobber
This keeps release reruns predictable. Re-uploading artifacts should not leave stale release notes behind.
Future Conventional Commit Mode
If the team later adopts Conventional Commits, the helper script can switch from
plain git log output to grouped output:
### Features
- add JSON tracing output
### Fixes
- preserve ANSI toggle in demo services
### Documentation
- document Splunk ingestion options
At that point, git-cliff is a good fit. The public contract can remain the
same:
scripts/release-notes.sh VERSION --update-changelog
Only the internals of the generator change.
Risks And Mitigations
| Risk | Mitigation |
|---|---|
| Commit messages are too noisy for customer-facing notes | Generate notes early, then review and edit before publishing. |
| Previous tag detection picks the wrong tag | Support --from TAG override and print the selected range. |
| Release script rerun duplicates changelog sections | Replace existing VERSION section instead of blindly prepending. |
| Existing GitHub release has stale notes after artifact upload | Always call gh release edit --notes-file for existing releases. |
Local builds unexpectedly modify CHANGELOG.md | Require explicit --update-changelog for file mutation. |
| Binary archives publish but Docker push fails | Build and push images before or immediately after GitHub release publication, print clear recovery commands, and support --docker-only reruns. |
| Docker image tags drift from GitHub release version | Have release.sh call root build.sh with the same VERSION; do not ask operators to type the image version separately. |
| Full release builds take longer because Dockerfiles rebuild Rust | Use Docker/BuildKit as the release builder and make runtime images copy exported binaries instead of running another cargo build. |
| App-level build scripts diverge again | Convert them to wrappers around repo-root build.sh. |
| Apple Silicon image support is confused with macOS binary support | Document that Docker Desktop on Apple Silicon needs linux/arm64 images, while native macOS binaries need aarch64-apple-darwin. |
| Windows artifacts are expected from a Linux Docker build | Build official Windows MSVC binaries on a Windows runner; treat Windows container images as a separate later phase. |
Implementation Plan
- Add
CHANGELOG.mdwith a short heading and no release entries. - Add
scripts/release-notes.shwith dependency-free git-log generation. - Add idempotent changelog insertion or replacement.
- Add
docker/Dockerfile.releaseor equivalent release-builder targets for Linux binaries. - Add repo-root
build.shfor all app Docker images and Linux image platforms. - Convert app-level build scripts into compatibility wrappers.
- Update
release.shto generatedist/release-notes-${VERSION}.md. - Update
release.shto call rootbuild.shwith the sameVERSION, unless--skip-dockeris set. - Update runtime image builds to copy binaries exported by the Docker release builder instead of compiling Rust again.
- Add CI matrix jobs for macOS Apple Silicon and Windows binary archives.
- Update
publish_release()to use--notes-filefor both new and existing releases. - Add README release documentation for the new flags and review workflow.
- Validate changelog generation locally with:
./release.sh 0.3.0 --notes-only --update-changelog --local
git diff --check
- Validate Docker image builds locally with:
./build.sh 0.3.0 --local
./build.sh 0.3.0 --app light-gateway --local
- Validate combined local release packaging with:
./release.sh 0.3.0 --local
- Validate CI artifact collection for Linux, macOS, and Windows archives.
- Validate GitHub and Docker Hub publishing on a test tag or draft release before using it for a production release.
Light-Workflow Runner
Status
Proposed design.
light-workflow-runner is a tenant-side execution agent primarily introduced
for workflow tasks that must run near tenant systems, tenant repositories,
private tools, local gateways, sidecars, or sandboxed release workspaces. The
same controller, lease, fencing, and backend substrate can also execute
standalone agent turns or actions submitted by light-agent. It is not a
second workflow or agent engine and it must not consume workflow start events
or own interactive agent sessions.
The SaaS-owned light-workflow instance remains authoritative for workflow
subjects. light-agent remains authoritative for standalone agent sessions,
turns, and actions. Tenant runners register with controller-rs, receive
server-issued fenced execution leases, execute only the leased attempt, and
report normalized results for the authenticated origin service to reconcile.
For effectful work, the runner uses a capability-described ExecutionBackend
as defined in the
Execution Backends And Sandbox Execution design.
The backend may be a microVM sandbox, shared-kernel container, Kubernetes Job,
dedicated VM, host-integrated environment, or fixed external action. Backend
credentials, lifecycle calls, logs, and artifact transfer belong to the runner.
light-workflow must not implement competing direct backend protocols.
The interactive agent ownership and placement model is defined in Light-Agent Execution.
Problem
For SaaS deployments, Light owns the main workflow control plane. Tenants may run APIs, gateways, sidecars, deployers, and other services in their own networks. Some workflow tasks need to execute inside those tenant environments instead of inside the SaaS control plane.
Examples:
- release workflows running in a prepared VM or sandbox with many repositories checked out,
- command-line tasks that need local files or private repository access,
- build and test tasks that need tenant-specific toolchains,
- deployment tasks that need access to private clusters,
- MCP servers or sidecars running only in the tenant network,
- AI repair tasks that need to inspect and patch a local sandbox workspace.
Running multiple full light-workflow instances would create control-plane
ambiguity:
- more than one instance may see the same workflow start event,
- tenant-side config can be changed through environment variables or local
values.yml, - a tenant runtime could claim work outside its intended scope,
- workflow definition loading and event consumption become hard to audit,
- duplicate workflow starts require more complex idempotency and broker ACLs.
The platform needs a runner model that lets tenant-side services execute approved tasks without letting them own workflow orchestration.
Goals
- Keep one authoritative SaaS
light-workfloworchestrator for workflow start events and workflow state. - Add a tenant-side
light-workflow-runnerexecutable for command, sandbox, deployment, MCP, and local tool execution. - Register tenant runners through
controller-rs. - Enforce task visibility with server-side leases, not runner-side local config.
- Support release runners in prepared VMs or sandboxes with checked-out repos and approved toolchains.
- Support per-tenant runner pools, execution profiles, capabilities, and network placement.
- Let
controller-rsperiodically audit effective runtime configuration. - Reuse
workflow-coretask models and result contracts where possible. - Keep the runner transport origin-neutral so workflow tasks and standalone agent turns can share execution infrastructure without sharing domain ownership.
Non-Goals
- Do not create a second workflow orchestrator that consumes workflow start events.
- Do not let tenant runners load arbitrary workflow definitions from local config.
- Do not trust tenant-side environment variables or local
values.ymlas the enforcement boundary. - Do not expose all workflow tasks to all registered runners.
- Do not let AI or command tasks bypass publish, signing, or human approval gates.
- Do not turn a standalone agent turn into a fake workflow task merely to use a runner.
- Do not let
controller-rsor a runner advance workflow or agent domain state.
Current Runtime Boundary
The current light-workflow executable starts the workflow event consumer, task
executor, and rule API in one process. The executor actively handles
control-plane task types such as ask, assert, call, set, and switch.
workflow-core already models run.container, run.script, run.shell, and
run.workflow. These task definitions are the right surface for runner-backed
execution, but they still need a runtime executor boundary.
This design keeps the workflow model shared and adds a separate runner executable for effectful execution.
Recommended Architecture
Domain Event Interactive Client
| |
v v
light-workflow light-agent
| workflow task | agent turn/action
| origin-owned policy and attempt | origin-owned policy and attempt
+------------------+-------------------+
|
v
controller-rs
|
| registration, fenced leases, heartbeat, quarantine
v
light-workflow-runner
|
| approved execution and backend interaction
v
ExecutionBackend
|
| declared isolation, lifecycle, resource, network, workspace,
| credential, log, and artifact policy
v
Tenant Runtime Environment
The split is:
light-workflow: Authoritative orchestrator. It sees workflow start events, loads workflow definitions, persists immutable policy snapshots, creates task attempts, owns retry and cancellation decisions, and records state.light-agent: Authoritative orchestrator for standalone authenticated sessions, turns, and agent actions. It owns model-loop and memory state, creates action intent, and reconciles results without advancing workflow state.controller-rs: Runtime control plane. It authenticates runners, records runner capabilities, issues and renews fenced execution leases, rejects stale reports, audits runtime config, and quarantines mismatched runners.light-workflow-runner: Tenant-side execution agent. It claims only leased attempts, validates the effective policy and command template, executes in the approved environment, streams bounded logs, safely exports artifacts, and reports normalized results.ExecutionBackend: Backend-specific adapter used by the runner for capability discovery, effective-configuration validation, idempotent prepare and execute, inspection, cancellation, log cursors, artifact copy, optional checkpoints, and cleanup.- Execution environment: A microVM sandbox, shared-kernel container, Kubernetes Job, dedicated VM, host-integrated environment, or fixed external action selected for the task’s purpose and minimum isolation requirement.
The runner can run beside tenant APIs, gateways, sidecars, and deployers. It may also run in a prepared release VM or sandbox with approved tools and repository workspaces.
A local deployment may colocate these components, but it must retain the same durable attempt, policy, lease, fencing, result, and audit contracts. Colocation must not create a second execution model.
Execution Origin And Subject
The runner wire contract and controller capacity queue are origin-neutral. They carry a generic execution subject instead of requiring every execution to be a workflow task:
executionId
origin.service
origin.instance
subject.kind = workflow-task | agent-turn | agent-action
subject.id
subject.attempt
optional workflow or agent correlation
The authenticated origin service owns the domain state:
light-workflowmay create and reconcileworkflow-tasksubjects;light-agentmay create and reconcileagent-turnandagent-actionsubjects;controller-rsreserves capacity, transports leases, and stores fenced execution observations, but cannot complete a workflow task or agent turn;- the runner executes a lease and cannot change origin-owned state.
Origin authorization is server-owned. A caller cannot select another origin kind in the payload. Tenant, host, origin service, and allowed subject kinds come from validated identity and registration.
Workflow and agent domain tables remain separate. Common scheduling, execution-attempt, lease, backend, session, artifact, and runtime-audit records may share the generic subject identity.
A runner-backed agent_action_attempt_t references the shared
execution_attempt_t row. Agent-domain tool, model-iteration, approval, budget,
and conversation fields remain outside the common runner table, matching the
separation between task_info_t and runner execution state.
Origin Result Wakeup
The common execution_attempt_t row is the durable source of truth. The
controller transaction that conditionally stores a newly terminal result also
emits a versioned PostgreSQL execution_result_ready_v1 notification. Its
bounded payload contains only attempt ID, authenticated origin, subject kind,
and correlation ID—never result bytes, tenant content, or authorization.
The named origin uses the notification only to wake its reconciler, reloads the authoritative row, verifies origin/subject/fencing bindings, and conditionally accepts the result into its own domain transaction. Every origin also performs an indexed startup and periodic scan of unaccepted terminal attempts because notifications can be missed, duplicated, or reordered. A later typed callback may be another wakeup, but neither a callback nor the runner may directly update workflow or agent domain tables.
The listener uses a dedicated PostgreSQL connection. On initial startup and
reconnect it establishes LISTEN first, then runs the catch-up scan, so a
terminal commit in that handoff window is either found by the query or queued
as a notification.
Event Visibility
Workflow start events should be visible only to the SaaS-owned
light-workflow orchestrator.
Recommended flow:
- A domain event is published.
- The SaaS
light-workflowconsumer evaluates matching workflow definitions. - It creates one workflow instance per matching definition.
- It creates tasks with runner requirements.
controller-rsexposes only eligible task leases to registered runners.- Runners execute leased tasks and return results.
This avoids duplicate starts and avoids tenant-side event subscription authorization problems.
If a future deployment requires separate workflow clusters, route start events by lane and enforce broker ACLs:
workflow.start.main
workflow.start.release
workflow.start.deployment
workflow.start.tenant.<tenantId>
Even with event lanes, the workflow database should enforce idempotency on a source-event key such as:
tenant_id + source_event_id + workflow_definition_id
For the SaaS model, task leases are the cleaner boundary than exposing start events to tenant runtimes.
Runner Registration
A runner must register before it can claim work.
Registration should include:
{
"runnerId": "release-runner-01",
"tenantId": "tenant-a",
"hostId": "host-a",
"runnerKind": "release",
"runnerPools": ["release"],
"executionProfiles": ["release-sandbox"],
"capabilities": [
"git",
"maven",
"cargo",
"rootless-buildkit",
"event-importer"
],
"executionBackends": [
{
"backendId": "cube-prod-east",
"kind": "microvm",
"implementation": "cubesandbox",
"version": "approved-version",
"capabilityDigest": "sha256:...",
"isolationBoundary": "microvm",
"supportsUntrustedCode": true,
"sessionScopes": ["task", "workflow"],
"workspaceModes": ["ephemeral", "copy-on-write", "workflow"],
"networkEnforcement": ["deny-by-default", "http-l7"],
"credentialDelivery": ["brokered", "proxy-injected"],
"lifecycle": ["inspect", "reconnect", "cancel", "destroy"]
},
{
"backendId": "docker-sbx-local",
"kind": "microvm",
"implementation": "docker-sandboxes",
"version": "approved-version",
"capabilityDigest": "sha256:...",
"isolationBoundary": "microvm",
"supportsUntrustedCode": true,
"sessionScopes": ["task", "workflow"],
"workspaceModes": ["clone"],
"networkEnforcement": ["deny-by-default", "http-l7"],
"credentialDelivery": ["proxy-injected"],
"containerEngineAccess": "private-daemon",
"lifecycle": ["inspect", "reconnect", "cancel", "destroy"]
},
{
"backendId": "toolbx-local",
"kind": "host-integrated",
"implementation": "toolbx",
"version": "approved-version",
"capabilityDigest": "sha256:...",
"isolationBoundary": "host-integrated",
"supportsUntrustedCode": false,
"sessionScopes": ["none"],
"hostExposure": [
"home",
"dbus",
"devices",
"network",
"ssh-agent",
"system-journal",
"host-sockets"
]
}
],
"imageDigest": "sha256:...",
"configHash": "sha256:...",
"commandAllowlistHash": "sha256:...",
"workspacePolicy": "release-workspace-v1",
"workspaceChangePolicyDigests": ["sha256:..."],
"trustBundleDigests": ["sha256:..."],
"provenanceModes": ["slsa-provenance-v1-signed"],
"localCleanup": {
"watchdog": true,
"durableJournal": true,
"backendResourceScan": true,
"policyDigest": "sha256:..."
},
"networkZone": "tenant-private",
"version": "0.3.0"
}
controller-rs validates the registration against server-side runtime policy.
If accepted, it creates a runner session and issues short-lived credentials for
heartbeat and task claim operations.
Local runner config can request capabilities, but the server decides the
effective capabilities. A runner cannot claim work merely because it sets an
environment variable or local values.yml value.
Registration is an admission request, not attestation by itself. Backend self-report does not establish a security boundary. Server-owned compatibility records and conformance tests decide which capabilities are trusted. Server policy must still constrain every lease, and the runner must prove the selected backend, immutable template or image, command, resource, network, workspace, host-exposure, workspace-change, credential, provenance, trust-bundle, and local cleanup settings for each attempt. A runner without a healthy watchdog and durable cleanup journal cannot claim backend-creating work. Unsupported or unverifiable required controls fail closed.
Execution Lease Model
The execution lease is the enforcement object. The runner should execute an attempt only when it has a valid lease issued by the control plane. Workflow correlation is present for a workflow subject; agent correlation is present for an agent subject.
Lease example:
{
"executionId": "01970f5d-0000-7000-8000-000000000000",
"leaseId": "01970f5d-0000-7000-8000-000000000001",
"fencingToken": 17,
"origin": {
"service": "light-workflow",
"instance": "workflow-main-east"
},
"subject": {
"kind": "workflow-task",
"id": "01970f5d-0000-7000-8000-000000000020",
"attempt": 1
},
"tenantId": "tenant-a",
"hostId": "host-a",
"runnerId": "release-runner-01",
"policySnapshotId": "01970f5d-0000-7000-8000-000000000010",
"policyDigest": "sha256:...",
"workflow": {
"wfInstanceId": "release-2026.06.0",
"taskId": "01970f5d-0000-7000-8000-000000000020",
"wfTaskId": "build-java-products"
},
"operationType": "run.shell",
"runnerPool": "release",
"executionProfile": "release-sandbox",
"profileVersion": 7,
"capabilities": ["git", "maven"],
"commandTemplateId": "light-fabric-release-build",
"commandIdempotencyKey": "release-2026.06.0/build-java-products/1",
"executionRequirements": {
"minimumBoundary": "microvm",
"allowedHostExposure": [],
"workloadTrust": "untrusted"
},
"executionBackend": {
"backendId": "cube-prod-east",
"kind": "microvm",
"implementation": "cubesandbox",
"version": "approved-version",
"capabilityDigest": "sha256:..."
},
"sandbox": {
"templateId": "tpl-immutable-id",
"templateDigest": "sha256:...",
"sessionScope": "workflow",
"workspaceMode": "copy-on-write"
},
"networkPolicy": "release-egress-v3",
"trustBundleRef": "trust-bundle://enterprise-egress-v3",
"trustBundleDigest": "sha256:...",
"resourcePolicy": "release-build-medium-v1",
"artifactPolicy": "release-artifacts-v2",
"inputRefs": [
{
"kind": "skill-package",
"id": "skill-package://coding/rust-review/7",
"digest": "sha256:...",
"size": 18432,
"mountMode": "read-only"
}
],
"provenancePolicy": {
"format": "slsa-provenance-v1",
"mode": "signed",
"policyDigest": "sha256:..."
},
"credentialRefs": [],
"approvalRef": null,
"deadlineAt": "2026-06-08T19:25:00Z",
"environmentExpiresAt": "2026-06-08T19:25:00Z",
"cleanupDeadlineAt": "2026-06-08T19:30:00Z",
"expiresAt": "2026-06-08T19:10:30Z",
"heartbeatIntervalSeconds": 15
}
Server-side validation must check:
- runner session is active,
- runner is not quarantined,
- tenant and host match,
- subject runner pool matches the registered pool,
- subject execution profile is allowed,
- required capabilities are a subset of effective runner capabilities,
- policy snapshot and digest match the active origin-owned subject,
- attempt number and fencing token match the active execution attempt,
- command template is approved,
- selected backend compatibility, workload trust, minimum isolation boundary, immutable template or image, host-exposure, network, resource, artifact, workspace, workspace-change, trust-bundle, provenance, lifecycle, local cleanup, and credential policies are supported,
- required runtime approval is valid and bound to the exact attempt inputs,
- execution and lease deadlines have not expired.
The runner reports execution start, logs, progress, and final result using the lease. The control plane rejects reports that do not match the active attempt, lease, and fencing token.
Attempt And Fencing
Remote runner execution is at-least-once. Every execution uses a durable attempt with a monotonically increasing attempt number and fencing token. The lease is short-lived but renewable while the subject is active. Renewal proves runner liveness; it does not extend the execution wall-clock deadline.
Start, progress, log, artifact, and result messages include the execution origin, subject, attempt, lease ID, and fencing token. Result acceptance uses compare-and-set semantics against the active attempt. An expired runner or late backend result cannot overwrite a newer attempt or transition origin-owned state.
If the runner loses contact after backend dispatch, the attempt becomes
UNKNOWN. The runner or control-plane reconciler inspects the backend
operation before the origin service decides to accept a result, wait, cancel,
retry, or require operator intervention. It must not assume that a transport
failure means the command did not run.
The subject idempotencyKey and lease commandIdempotencyKey are propagated to
the backend and external action where supported. Side-effecting command
templates must define an external idempotency or reconciliation contract before
automatic retry is allowed.
Cancellation And Lease Loss
Cancellation and policy revocation fence the attempt before asking the runner to stop. The runner cancels the backend operation, destroys or quarantines the execution environment as required, revokes execution credential handles, and reports cleanup state. A completion received after fencing is retained only as diagnostic evidence.
If cleanup cannot be confirmed, the attempt enters cleanup-pending and an
orphan reconciler continues inspecting backend resources. Lease loss alone
does not prove that the command stopped.
Local Watchdog And Disconnected Cleanup
The runner must be able to clean tenant-local resources without contacting
controller-rs. A supervisor separate from the execution worker writes a durable
local cleanup journal before backend preparation, stops new work when the
control-plane session is lost, and lets active work continue only until the
locally tracked lease expiry. A connectivity grace period cannot extend the
lease or execution deadline.
At lease expiry or the earlier execution or environment deadline, the supervisor locally fences the attempt, revokes execution credential handles, cancels the backend operation, and destroys or quarantines the environment. It uses a monotonic deadline bounded by the authenticated absolute lease deadline. Backend resources carry owner and expiry tags and use a native TTL where the backend supports one, so cleanup does not depend on the runner host restarting.
Runner startup and periodic sweeps replay incomplete journal records and inspect
tagged resources with bounded cleanup backoff. Reconnection reports outcome and
cleanup evidence, but it cannot make an expired result valid. External actions
that may already have taken effect remain UNKNOWN and are reconciled rather
than blindly retried. The detailed journal and watchdog contract is defined in
the Execution Backends And Sandbox Execution design.
Capacity Scheduling And Claim Backoff
controller-rs keeps execution subjects with no eligible slot in a bounded,
per-tenant fair PENDING_CAPACITY queue. light-workflow retains
authoritative workflow-task state and light-agent retains authoritative
agent-turn/action state. The controller atomically reserves runner and backend
capacity and returns a short-lived reservation token; idempotent, fenced
attempt creation and lease issuance bind that token. Temporary saturation
therefore does not consume an origin retry, create an execution environment, or
cause every runner to race for the same subject.
Claims use long polling or server push. An empty claim or temporary backend
capacity response includes retryAfter; clients apply capped exponential
backoff with jitter, and the controller wakes only a bounded number of eligible
waiters when capacity returns. Hard quota or policy failures are terminal
admission denials until configuration changes. Queue timeout, origin deadline,
cancellation, and policy revocation remove pending work without ever
dispatching it.
Task Routing
light-workflow should execute pure control-plane tasks locally:
ask
assert
set
switch
context merge
workflow branching
workflow persistence
approved internal call tasks
light-workflow-runner should execute effectful or tenant-local tasks:
run.shell
run.script
run.container
call.mcp to tenant-local servers
deployment commands
release build and test commands
AI repair with filesystem access
browser automation
external tool processes
Some call.* tasks can run on either side. The routing decision should come
from effective task policy:
| Task | Default Runtime | Notes |
|---|---|---|
call.http internal SaaS API | light-workflow | Use host-side service credentials. |
call.http tenant-private API | runner | Needs tenant network access. |
call.mcp approved SaaS gateway | light-workflow | Gateway enforces tool access. |
call.mcp tenant-local server | runner | Local sidecar or private MCP server. |
call.agent no tools | light-workflow | Bounded model call. |
call.agent with file/tools | runner | Requires sandbox/tool policy. |
Agent Call Placement
Workflow agent calls need an explicit placement decision. The same workflow can use more than one agent execution mode, but the placement must come from server-side policy and task metadata, not tenant-side local config.
Use three agent execution modes.
Native Workflow Agent
Native call: agent stays in the SaaS-owned light-workflow process. This is
the current bounded agent task model: light-workflow resolves the portal
agent, skill, and tool metadata, builds a constrained prompt from workflow
context, calls the configured model provider, validates structured output, and
continues the workflow.
Use native workflow agents for bounded reasoning:
- classify a request or command result,
- summarize API responses or logs,
- choose a workflow branch,
- draft a customer-facing explanation,
- decide whether human review is required,
- produce JSON output that must match a schema.
Native workflow agents should not receive filesystem access, local network
access, release secrets, or dynamic tool execution. API orchestration should
remain explicit workflow tasks such as call.http, call.mcp, assert,
switch, and ask.
By default, native workflow agents use SaaS-approved model providers and model credentials managed by the Light control plane. Tenant-private repository content, tenant-local logs, local files, and private network data should not be sent to this path unless the tenant policy explicitly allows it.
Runner Agent
Runner agents execute through light-workflow-runner under a server-issued
execution lease. Use this mode when the agent needs access to tenant-local
state or effectful tools:
- checked-out repositories,
- command output plus working directory inspection,
- private tenant network access,
- local MCP servers,
- sandbox tools,
- AI repair of source code,
- test reruns,
- branch or pull-request creation.
The origin service still creates the workflow task or standalone agent action
and records the result. controller-rs issues a lease only to a runner whose
effective capabilities, runner pool, execution profile, command allowlist,
workspace policy, and audit state match the execution requirements.
Runner agent lease example:
{
"executionId": "01970f5d-3333-7000-8000-000000000001",
"origin": {
"service": "light-workflow",
"instance": "workflow-main-east"
},
"subject": {
"kind": "workflow-task",
"id": "01970f5d-3333-7000-8000-000000000020",
"attempt": 1
},
"operationType": "call.agent",
"agentPlacement": "runner",
"runnerPool": "release",
"executionProfile": "release-sandbox",
"profileVersion": 7,
"executionRequirements": {
"minimumBoundary": "microvm",
"allowedHostExposure": [],
"workloadTrust": "untrusted"
},
"executionBackend": {
"backendId": "docker-sbx-local",
"kind": "microvm",
"implementation": "docker-sandboxes",
"version": "approved-version",
"capabilityDigest": "sha256:..."
},
"sandbox": {
"sessionScope": "task",
"isolationClass": "agent-call",
"workspaceMode": "clone"
},
"modelProviderScope": "tenant",
"modelAccessMode": "brokered-proxy",
"modelProxyRef": "tenant-model-proxy-eastus",
"workloadIdentityRef": "attempt://model-access",
"dataBoundary": "tenant-network",
"runtimeToolManifestDigest": "sha256:...",
"allowedTools": [
{
"toolRef": "runner://command/cargo-test",
"modelAlias": "cargo_test",
"schemaDigest": "sha256:...",
"capability": "command.cargo.test"
}
],
"workspaceAccess": "copy-on-write-release-workspace",
"workspaceBaseRevision": "git:...",
"workspaceChangePolicyId": "agent-source-only-v1",
"workspaceChangePolicyDigest": "sha256:...",
"networkPolicy": "release-egress",
"credentialPolicy": "brokered-task-scoped",
"maxRepairAttempts": 2,
"requiresHumanApprovalBefore": ["publish", "sign", "tag"]
}
The runner agent can inspect files and propose or apply bounded patches inside the approved workspace. It must not publish artifacts, sign releases, push final tags, read unrestricted secrets, or expand its own permission scope.
By default, runner agents use tenant-approved model providers and tenant-owned credentials. This keeps private workspace data and private network context inside the tenant boundary and avoids exposing SaaS model credentials to tenant-side runtimes.
Runner-local tools do not appear in light-gateway tools/list. Before worker
startup, the controller intersects catalog entries placed on the runner with
the server-approved runtime compatibility record, execution profile, lease
allowedTools, and immutable runtime-tool manifest. The worker may narrow that
set using live local availability or sandbox-local MCP tools/list, but cannot
add authority. Each model alias remains bound to one stable internal tool
reference, schema digest, placement, and dispatcher; cross-placement name
collisions fail closed. Broker/control sockets and backend lifecycle operations
are never exposed as tools.
Agent Workspace Change Boundary
The lease for a write-capable agent includes the immutable base commit or tree
and a server-owned workspaceChangePolicyId and digest. The policy intersects
allowed paths with protected-path denies and limits file count, bytes, types,
creation, deletion, rename, mode, submodule, nested-repository, and binary
changes. Workflow metadata and agent output cannot weaken it.
The default policy denies CI/CD definitions, reusable automation, workflow
definitions, CODEOWNERS and approval policy, .git internals and hooks, and
release, publish, signing, deployment, credential, runner, and execution-policy
configuration. Repository-specific equivalents are added by operator policy.
Write interception inside the sandbox is defense in depth. After execution, a
trusted runner component computes the authoritative diff from the immutable
base in a fresh trusted checkout without repository-provided hooks or mutable
Git configuration. It normalizes case, Unicode, and separators according to
repository rules, detects link and rename tricks, and creates an immutable
canonical patch whose digest is checked. A violation fails with
workspace_change_denied; no branch, pull request, artifact, publish, or
signing action may consume the patch.
The agent receives no push credential. Branch or pull-request creation is a separate fixed action over the immutable accepted patch and its policy result, never the mutable agent workspace. Even an accepted agent patch remains untrusted: tests run under the same isolation, and a release rebuilds from the reviewed and merged immutable commit rather than publishing artifacts directly from the repair workspace.
Runner Agent Execution Isolation
The runner itself is a tenant-side execution agent. For stronger isolation, the runner can launch the agent task inside a separate environment such as Cube Sandbox, Docker Sandboxes, a dedicated VM, or a Kubernetes Job using an approved runtime class. This should be a tenant-selectable policy because the runner is deployed in the tenant namespace, but the effective choice must still meet the server-owned minimum boundary and be recorded in the execution lease.
Recommended isolation levels:
| Session Scope | Isolation Class | Use Case | Default Policy |
|---|---|---|---|
none | bounded-runner | Model call with no tools, files, or private-network mutation | Allowed only for explicitly approved low-risk profiles |
workflow | release-build | Build, test, or diagnosis sharing one checkout and cache | Useful for release workflows |
agent-session | interactive-agent | Coding workspace reused across authenticated turns | Explicit TTL and identical principal/policy/base required |
task | agent-call | AI repair, generated patches, dynamic tools, or untrusted scripts | Preferred for high-risk agent tasks |
task | publish | Fixed publish or signing action over immutable artifacts | Required for high-value credentials |
Backend purpose and execution session scope are separate. Recommended defaults are:
| Execution need | Minimum boundary | Candidate backend | Important constraint |
|---|---|---|---|
| Trusted local development | host-integrated | Runner operating in Fedora Toolbx | Toolbx is not a sandbox and cannot run untrusted or secret-bearing tasks |
| Trusted CI, tests, or packaging | shared-kernel-container | Rootless Docker or Podman; ordinary Kubernetes Job | No privileged mode, host namespaces, or host container-engine socket |
| Autonomous agent or untrusted code | microvm | Cube Sandbox; Docker Sandboxes; Kubernetes only with an approved stronger runtime | Docker Sandboxes use clone workspace mode; no fallback to a shared-kernel container |
| Privileged or long-running tenant work | dedicated-vm | Approved tenant-dedicated VM | Pin and attest the image, network, identity, limits, and teardown policy |
| Publishing, signing, or deployment | external-service | Fixed typed action or dedicated service | Accept immutable inputs; do not expose a general shell with release credentials |
For a release workflow, the runner should usually orchestrate a separate task
sandbox with the agent-call isolation class for AI repair. The runner provides
only the leased copy-on-write workspace, approved tools, network policy, and
opaque credential handles allowed by the task policy. It collects bounded logs,
artifacts, patches, and structured output, then destroys the sandbox. Freezing
or checkpointing is allowed only when retention policy permits it and no raw
credential entered the sandbox.
For a reused workflow or agent session, effective expiry is the earliest of
the origin session idle/max expiry, execution-session policy, credential or
broker grant expiry, and backend-native TTL. Closing, revoking, or expiring the
origin session creates a durable idempotent common cleanup request in the same
origin transaction. controller-rs fences and cancels active attempts and
dispatches cleanup; the runner destroys the physical session and records
evidence. Cleanup retries across restarts. Backend-native TTL is the final
fail-safe, not the expected way to reclaim an abandoned session.
Action ownership and session retention are separate. Ending an action lease
removes executable authority, broker access, and action credentials. It does
not by itself delete a compatible reused session. An origin may create a
durable IDLE_APPROVAL_HOLD for a non-secret workspace with an explicit hold
ID, reason, policy digest, holdUntil, checkpoint/patch evidence, and cost
policy. The runner pauses or checkpoints where supported. The hold cannot
extend the session idle/fixed maximum or survive origin close, revocation,
policy mismatch, or cleanup request. Controller and runner reconcilers use this
session state rather than treating zero active attempts as abandonment.
This creates a layered boundary:
SaaS light-workflow or light-agent
-> controller-rs fenced execution lease
-> tenant light-workflow-runner
-> ExecutionBackend
-> task execution environment, isolationClass=agent-call
-> model, tools, files, network
Tenants may choose Cube Sandbox, Docker Sandboxes, dedicated VM isolation, an
approved Kubernetes runtime, or no additional environment for explicitly
trusted profiles. Fedora Toolbx and ordinary shared-kernel containers must not
satisfy a microVM requirement. Runner registration advertises supported
execution backends, isolation boundaries, session scopes, workspace modes,
host exposures, and enforcement capabilities. If the approved compatibility
record cannot satisfy the task requirements, controller-rs must not issue the
lease; it must not silently choose a weaker backend.
Local runner config can select among tenant-approved profiles, but it cannot
weaken a task requirement. The lease contains the final effective
executionBackend identity, implementation, version, capability digest,
sandbox.sessionScope, sandbox.isolationClass, immutable template or image,
workspace, host-exposure, network, resource, tool, artifact, lifecycle, and
credential policy. Heartbeat, backend inspection, and audit snapshots should
prove the runner is still operating under that profile.
Execution Backend Boundary
Backend-specific execution lives behind ExecutionBackend in the runner, not
in light-workflow. The core operations are capability discovery,
effective-configuration validation, idempotent environment preparation,
environment and operation inspection, reconnect, idempotent execution,
resumable logs, cancellation, safe artifact copy, and cleanup. Checkpoint and
session operations are optional capabilities rather than assumptions every
backend must emulate. Backends also report measured execution evidence where
supported; the trusted runner or control-plane attestor, not the tenant task,
constructs the final provenance statement.
The runner fails closed when the backend cannot enforce a required isolation, host-exposure, resource, network, credential, lifecycle, or inspection control. A backend timeout or transport error is not automatically a command failure; the runner records an unknown outcome and reconciles it through backend inspection.
Trusted Input And Skill-Package Staging
Immutable context, workspace bases, trust bundles, and skill packages are
resolved into digest-bound input records before dispatch. The trusted runner,
outside the sandbox payload boundary, downloads them with runner authority,
checks kind, size, digest, signature/provenance where required, and archive
safety, and stages them before backend creation. The backend exposes only the
selected bytes as read-only mounts with nodev, nosuid, and noexec unless
an approved entrypoint requires execution.
light-agent-worker may revalidate the mounted manifest, but neither it nor
generated code downloads a package or receives artifact-store credentials.
Input verification or staging failure prevents sandbox start. Staging paths
are attempt/session scoped, journaled without secrets, and removed through the
same idempotent cleanup contract as the backend environment.
The complete policy, lifecycle, Cube production baseline, artifact, secret, and failure contracts are defined in the Execution Backends And Sandbox Execution design.
Agent Service
Containerized light-agent services should be invoked explicitly. They are the
right runtime for interactive or independently scaled agents:
- chat and session memory,
- dynamic
tools/listandtools/callloops, - long-lived specialist agents,
- independently deployed model/tool runtime,
- local catalog caching.
Do not silently change native call: agent to call a containerized
light-agent service. Use an explicit contract such as call: agent-service
or call: agent with mode: service so operators can audit which runtime path
was used.
When an interactive light-agent session needs local execution, it submits an
agent-turn or agent-action execution subject through the same controller
and runner substrate. It does not create a fake workflow task. Session, turn,
tool authorization, memory, and model-provider rules remain governed by the
Light-Agent Execution design.
For a workspace-aware coding or external-agent turn, the runner starts the
small sandbox-side light-agent-worker with a pinned runtime adapter. It does
not start the public light-agent service inside the sandbox. The worker owns
only the leased local loop and normalized event stream; light-agent remains the
agent-domain authority.
Model Provider Boundary
Agent placement and model-provider placement should be decided together.
Recommended defaults:
native call: agent in SaaS light-workflow
-> SaaS-approved model provider
-> SaaS workflow context data boundary
leased runner agent in tenant workflow runner
-> tenant-approved model provider
-> tenant network/workspace data boundary
containerized light-agent service
-> service-owned or tenant-approved model provider
-> explicit service data boundary
The default SaaS model is useful for bounded reasoning over workflow-safe context, such as classification, summaries, branch decisions, and structured JSON output. It should not be the default path for tenant-local source code, private command logs, local files, or private network data.
The default runner model is useful when the task needs tenant-local context. The runner owns an attempt-scoped model broker outside the untrusted payload boundary. It binds a protected local channel to the subject, adapter, approved model, data-boundary and policy digests, token/cost budget, rate, audience, and expiry. Provider keys and reusable proxy bearer tokens never appear in a sandbox environment variable, argument, prompt, workspace, or persistent file. SaaS model credentials must not be sent to tenant runners.
The control plane should still make this policy-driven instead of hard-coding it. Some tenants may require every agent call, including bounded summaries, to use their own provider or regional model endpoint. In that case, the workflow task should be routed to a runner or to an approved tenant model gateway even if the reasoning itself is small.
Lease examples:
{
"agentPlacement": "workflow",
"modelProviderScope": "saas",
"modelProviderRef": "light-managed-default",
"credentialRef": "saas-secret://llm-provider",
"dataBoundary": "saas-workflow-context"
}
{
"agentPlacement": "runner",
"modelProviderScope": "tenant",
"modelAccessMode": "brokered-proxy",
"modelProxyRef": "tenant-model-proxy-eastus",
"workloadIdentityRef": "attempt://model-access",
"dataBoundary": "tenant-network"
}
The sandbox worker receives only a runner-created preconnected descriptor,
peer-credential-checked Unix-domain socket, vsock, or backend-equivalent local
channel. A socket path is not sufficient authority: the broker authenticates
the peer and attempt and independently enforces model and budget policy. The
worker/runtime and generated payload use separate identities and process/mount
namespaces; ptrace and cross-process /proc access are denied, and the worker
does not pass its broker descriptor to child payloads. An adapter that requires
an extractable provider key is ineligible for an untrusted runner profile.
Recommended placement rule:
bounded reasoning over workflow context -> native call: agent in light-workflow
agent needs one isolated local effect -> leased agent-action
workspace-aware coding or external agent loop -> light-agent-worker in leased agent-turn sandbox
interactive session or dynamic tool loop -> containerized light-agent service
For release workflows, use native call: agent to summarize and classify a
failed command. Use a runner agent for repo inspection, patch generation, test
rerun, and pull-request creation. Human approval remains required before
publish, signing, or final tag creation, but approval waiting is a durable
light-workflow orchestration state and never an active runner lease.
Effective Policy
Workflow definitions and tasks can request runner execution through metadata, but the control plane computes the effective policy.
Workflow-level example:
document:
dsl: "1.0.3"
namespace: release
name: java-release
version: "0.1.0"
metadata:
lightWorkflow:
runner:
runnerPool: release
capabilities:
- git
- maven
- rootless-buildkit
security:
schemaVersion: 1
executionProfile: release-sandbox
profileVersion: 7
placement: runner
isolation:
minimumBoundary: microvm
allowedHostExposure: []
workloadTrust: untrusted
sandbox:
sessionScope: workflow
workspace:
mode: copy-on-write
Task-level example:
do:
- build-java:
run:
shell:
command: light-release-build
arguments:
- "${ .release.version }"
metadata:
lightWorkflow:
runner:
runnerPool: release
commandTemplateId: light-fabric-release-build
security:
isolation:
minimumBoundary: microvm
sandbox:
sessionScope: workflow
Runtime policy resolution:
- Operator-approved immutable profile definitions set the base allowed commands, workload trust, minimum isolation, host exposure, execution backends, templates or images, networks, resources, mounts, session scopes, workspace modes and change policies, trust bundles, model-provider scopes, data boundaries, artifact provenance, local cleanup, and credentials.
- SaaS service policy intersects the profiles and approved backend compatibility records available in the deployment.
- Tenant policy further restricts the allowed set.
- The workflow requests one profile and immutable version.
- Task metadata may request stricter isolation or a subset of capabilities; it cannot downgrade operator-derived workload trust.
light-workflowpersists the effective workflow policy snapshot and derives an effective task-policy digest for each attempt.controller-rsvalidates the registered runner and selected backend against the server-owned compatibility record and capability digest.- The fenced execution lease contains the selected backend identity and final allowed execution scope.
For an agent origin, light-agent performs the analogous immutable
agent/turn/action policy resolution described in
Light-Agent Execution. The controller and runner
consume the same final execution-policy fields without taking ownership of how
the origin derived them.
Policy merging is field-specific. Allowlists intersect, explicit denies win, numeric limits use the lowest permitted maximum, and task-scoped isolation may strengthen workflow-scoped isolation. The selected backend must meet the minimum boundary and every required capability; there is no fallback from a microVM to a shared-kernel or host-integrated environment. Backend and template choices must be members of the approved compatibility set. A task cannot weaken the effective policy.
The policy snapshot, execution session, execution attempt, backend operation, and lease must use dedicated runtime tables. They must not be stored in mutable workflow or agent context. Profile versions and backend compatibility records are immutable; emergency revocation fences new and active attempts and records the reason.
Runtime Configuration Audit
Tenant-controlled local configuration cannot be the source of truth. A runner can load local config for its own startup, but the server must verify and audit the effective runtime state.
controller-rs should audit at three points.
Startup Admission
On registration, the runner reports:
- binary version,
- image digest or VM image ID,
- effective config hash,
- command allowlist hash,
- enabled execution profiles,
- runner pools,
- mounted workspace paths,
- supported execution session scopes, isolation boundaries, and isolation classes,
- backend IDs, kinds, implementations, versions, capability digests, and server-approved enforcement capabilities,
- host exposures, workspace modes, workspace-change policy digests, container-engine access, and immutable template or image IDs and digests,
- local watchdog and cleanup-journal health and policy digest,
- trust-bundle digests and supported language-runtime adapters,
- provenance formats, modes, and trusted attestor identity where applicable,
- allowed model provider scopes,
- network zone,
- resource, artifact, and credential-delivery policies,
- host and tenant identity.
controller-rs compares this report with approved server-side policy before
allowing claims.
Heartbeat
Each heartbeat should include:
{
"runnerId": "release-runner-01",
"sessionId": "01970f5d-1111-7000-8000-000000000001",
"status": "ready",
"configHash": "sha256:...",
"commandAllowlistHash": "sha256:...",
"imageDigest": "sha256:...",
"watchdog": {
"status": "healthy",
"lastSweepAt": "2026-06-08T18:59:45Z",
"cleanupPending": 0,
"policyDigest": "sha256:..."
},
"activeAttempts": [
{
"leaseId": "01970f5d-0000-7000-8000-000000000001",
"attempt": 1,
"fencingToken": 17,
"policyDigest": "sha256:...",
"backendId": "cube-prod-east",
"backendOperationId": "backend-op-123",
"leaseExpiresAt": "2026-06-08T19:00:30Z"
}
],
"timestamp": "2026-06-08T19:00:00Z"
}
If a hash changes unexpectedly, the controller marks the runner suspicious and stops issuing new leases.
Periodic Deep Audit
Periodically, controller-rs should request an effective runtime snapshot from
the runner and compare it with the approved policy. For high-risk runners, the
snapshot should include command allowlist, immutable sandbox template, mount
list, host exposure, workspace mode and change policy, resource, network, and
trust-bundle policy, backend effective configuration, artifact and provenance
policy, local cleanup journal health, tagged-resource scan result, and
credential binding names without values. Where the backend permits it, the
control plane should compare runner claims with backend inspection rather than
relying only on runner self-reporting.
On mismatch:
- Mark the runner as
quarantinedand stop issuing new leases. - Fence active task attempts so late results cannot transition workflows.
- Revoke claim and task credentials.
- Request cancellation and backend cleanup for affected operations.
- Emit an append-only runtime audit event.
- Create an operator task when outcome or cleanup remains unknown.
Audit is not the only enforcement mechanism. It detects drift after admission. The fenced task lease is the primary runtime authorization boundary, while the backend must independently enforce its declared isolation, resource, network, workspace, lifecycle, and credential policy.
Release Runner Mode
A release runner is a specialized light-workflow-runner profile.
It can run in:
- an approved dedicated VM,
- a Cube Sandbox or Docker Sandboxes microVM,
- a Kubernetes Job with a recorded runtime class and node policy,
- a rootless shared-kernel container for explicitly trusted tasks,
- a controlled bare-metal or Toolbx environment for trusted local helper tasks.
The last two options are operational environments, not substitutes for a microVM boundary. A runner using them cannot claim untrusted-code, isolated agent, publish, signing, or secret-bearing tasks unless a separate eligible backend performs that task.
Recommended default for release workflows:
- one workflow-scoped sandbox or VM workspace for checkout, build, test, and package steps,
- task-scoped
agent-callisolation for AI repair, source inspection, generated patches, and test reruns driven by an agent, - immutable artifact export through controlled storage with trusted-side hashing and trusted-side in-toto/SLSA provenance,
- server-owned protected-path policy plus trusted post-export diff validation for every agent patch,
- a task-scoped fixed publish action or separate release service for publishing,
- an external signing service or task-scoped fixed signing action,
- human approval bound to the exact artifact digest, release target, version, command template, policy snapshot, and expiry,
- clean checkout inside the runner rather than writable host repository mounts,
- AI repair limited to sandbox workspace changes, with branch or PR creation performed by a separate fixed action,
- a clean release rebuild from the reviewed and merged immutable commit rather than publishing an artifact directly from an agent-repair workspace.
Writable host mounts should be avoided for AI repair and release commands. If host repositories must be mapped, default to read-only mounts and copy the repo into a runner-owned working directory before mutation.
Publish and signing actions must not execute arbitrary scripts from the mutable build workspace. They consume only immutable artifact records and use operator-owned command templates. They verify artifact provenance and approval bindings before use. Prefer brokered short-lived identity or backend-side credential injection so the raw credential never enters arbitrary workflow code. Per-task isolation limits exposure but does not make untrusted code safe to receive a release token.
Do not mount a host Docker socket into tenant-authored runners or sandboxes. Container image builds use an approved rootless builder or remote build service with pinned builder and base-image digests.
Approval Ownership And Lease Handoff
Human approval is owned entirely by the authenticated origin service:
light-workflow for workflow tasks and light-agent for standalone agent
actions. The runner never polls a person. Before an approval wait, the origin
commits any known result and immutable evidence, terminalizes the current
attempt, ends its action lease, closes its model-broker channel, and revokes
task credentials. A task-scoped environment is cleaned.
For a reusable non-secret session workspace, WAITING_APPROVAL may coexist
with the distinct bounded IDLE_APPROVAL_HOLD described above. The hold is not
an action lease, carries no executable authority, is preferably paused or
checkpointed, and consumes observable retained-resource quota. If a safe hold
cannot be established, export an immutable approved patch/checkpoint and clean
the environment; important uncommitted work must not depend only on a live
sandbox.
The origin transaction that enters WAITING_APPROVAL also persists exactly one
session disposition—cleanup or policy-valid bounded hold. If origin and common
session state later use separate databases, an idempotent transactional outbox
provides that handoff. A session reconciler must never observe an ended action
lease without the durable disposition and guess whether to delete the
workspace.
When policy knows approval is required before execution, the origin records
the bound intent but creates no common execution attempt. If a running runtime
discovers an approval boundary, it returns a known approval_required terminal
result and its attempt is fenced and cleaned or explicitly checkpointed under
non-secret retention policy.
After approval, the origin revalidates the exact operation, arguments,
artifacts/provenance where applicable, destination, policy digest, expiry, and
single-use nonce. It consumes the approval into a new numbered domain attempt
and a new common execution_attempt_t; controller-rs issues a fresh lease,
monotonic fencing token, and fresh task-scoped grants. The pre-approval attempt,
lease, backend handle, and grants remain immutable and cannot be reused.
Rejection or expiry changes only origin orchestration state and dispatches no
runner work.
If the held physical workspace still exists after approval, the new action may reuse it only after principal/base/runtime/policy/expiry and cleanup-state revalidation. Otherwise it starts in a fresh environment and restores only a verified policy-permitted checkpoint or patch.
A non-secret workflow session may be checkpointed or retained during approval only under explicit maximum-lifetime, cost, and retention policy; approval must not depend on it. The default release flow cleans the build environment after export. An unknown prior side effect is reconciled before approval can authorize another attempt.
Runner API
The first runner API can be small.
POST /runner/register
POST /runner/heartbeat
POST /runner/claim
POST /runner/execution/{leaseId}/started
POST /runner/execution/{leaseId}/renew
POST /runner/execution/{leaseId}/progress
POST /runner/execution/{leaseId}/log
POST /runner/execution/{leaseId}/complete
POST /runner/execution/{leaseId}/fail
POST /runner/execution/{leaseId}/unknown
POST /runner/execution/{leaseId}/cancelled
POST /runner/execution/{leaseId}/cleanup
POST /runner/execution-session/{executionSessionId}/hold
POST /runner/execution-session/{executionSessionId}/resume
POST /runner/execution-session/{executionSessionId}/cleanup
POST /runner/audit-snapshot
POST /runner/drain
controller-rs can expose these APIs directly or mediate them over its
existing persistent connection model. For private tenant networks, outbound
runner registration and polling is preferable to inbound SaaS calls into the
tenant environment.
The claim response should include only the origin-neutral subject envelope and
payload needed for execution, not a full workflow definition, agent session,
or conversation history. /runner/claim supports long polling and an empty
response includes retryAfter; the runner applies capped exponential backoff
with jitter. A subject waiting for capacity has no lease and is not returned
until the controller has atomically reserved an eligible slot.
Every execution API request includes a unique message ID, origin, subject, attempt number, lease ID, and fencing token. Repeated delivery of the same message is idempotent. Lease renewal extends execution ownership only up to the execution deadline. Cancellation can be delivered through the persistent connection or returned from heartbeat and renewal calls.
There is no runner API for waiting on human approval. Approval is handled by
the origin service, and only a newly created post-approval attempt appears
through /runner/claim. Session hold and resume are idempotent lifecycle
commands from the controller; they contain no human decision and cannot create
or renew an action lease. They carry the session state version/fence, policy
digest, bounded holdUntil, and checkpoint/patch policy.
Command Result Contract
Runner results should use a normalized command result so light-workflow,
light-agent, human tasks, AI diagnosis, and audit do not depend on raw
console parsing.
{
"executionId": "01970f5d-0000-7000-8000-000000000000",
"leaseId": "01970f5d-0000-7000-8000-000000000001",
"fencingToken": 17,
"origin": {
"service": "light-workflow",
"instance": "workflow-main-east"
},
"subject": {
"kind": "workflow-task",
"id": "01970f5d-0000-7000-8000-000000000020",
"attempt": 1
},
"workflow": {
"taskId": "01970f5d-0000-7000-8000-000000000020",
"wfTaskId": "build-java-products"
},
"runnerId": "release-runner-01",
"policyDigest": "sha256:...",
"commandTemplateId": "light-fabric-release-build",
"backendId": "cube-prod-east",
"backendKind": "microvm",
"backendImplementation": "cubesandbox",
"backendVersion": "approved-version",
"backendOperationId": "backend-op-123",
"status": "failed",
"outcome": "known",
"exitCode": 1,
"startedAt": "2026-06-08T19:10:00Z",
"completedAt": "2026-06-08T19:18:30Z",
"summary": "Maven test failure in db-provider",
"stdoutRef": "artifact://release/2026.06.0/build/stdout.log",
"stderrRef": "artifact://release/2026.06.0/build/stderr.log",
"artifacts": [
{
"artifactId": "01970f5d-0000-7000-8000-000000000030",
"name": "surefire-reports.zip",
"sha256": "sha256:...",
"size": 42000,
"storeUri": "artifact://release/2026.06.0/build/surefire-reports.zip",
"provenanceRef": "provenance://release/2026.06.0/build",
"provenanceDigest": "sha256:..."
}
],
"resourceUsage": {
"wallTimeSeconds": 510,
"peakMemoryBytes": 2147483648
},
"cleanupState": "complete",
"approvalRef": null,
"workspaceBaseRevision": null,
"workspaceChangePolicyDigest": null,
"patchDigest": null,
"changedFiles": [],
"aiDiagnosisAllowed": true
}
The runner streams bounded, ordered log chunks with sequence numbers and resumable cursors. Full logs are stored as tenant-scoped artifacts only when policy allows it. Origin domain context keeps summaries and immutable references, not unbounded stdout or stderr.
Artifact names and paths are untrusted. The runner must enforce canonical-root and no-follow extraction, reject traversal and special files, apply count and byte limits, and compute the authoritative digest after bytes cross the sandbox trust boundary.
For an agent task, changedFiles is the canonical manifest produced by the
trusted post-export diff, not a list supplied by the agent. The result is
accepted only when the base commit, patch digest, and workspace-change policy
digest match the lease. For a build requiring provenance, command success is
not sufficient: failure to generate or authenticate the required in-toto/SLSA
statement fails the attempt before any publish action can consume its artifacts.
If command outcome is unknown, the runner sends status: "unknown" with the
backend operation ID and diagnostic reference instead of fabricating a
failure. A later reconciliation report uses the same attempt and fencing token
unless the control plane has already fenced it.
Security Requirements
- Runners authenticate to
controller-rswith tenant-scoped credentials. - Runner and backend control-plane traffic is encrypted and mutually authenticated where it crosses a host boundary.
- Execution leases are short-lived, renewable, scoped to one attempt, and protected by a monotonically increasing fencing token.
- A durable local watchdog stops new work on disconnect, locally fences work at lease expiry, revokes credential handles, and cleans tagged backend resources; backend-native expiry protects against runner-host failure where available.
- Runners never see workflow start events unless they are explicitly deployed as trusted orchestrators in a non-SaaS topology.
- Runners receive bounded execution payloads, not complete workflow definitions or agent sessions.
- Server-side policy decides runner pools, execution profiles, minimum isolation boundaries, approved backend compatibility records, immutable templates or images, capabilities, command templates, resources, networks, host exposures, mounts, workspace modes, session scopes, model provider scopes, data boundaries, artifacts, and credentials.
- Required backend controls and the effective rendered backend configuration are verified before execution; missing controls fail closed. Backend self-report alone cannot upgrade its trusted capability record.
- Temporary capacity shortage remains in a bounded fair queue with
retryAfterand jittered claim backoff; it does not create an execution attempt or consume an origin retry. - SaaS model credentials must not be sent to tenant-side runners.
- Tenant-private source code, local files, and private command logs should use tenant-approved model providers unless tenant policy explicitly allows SaaS model processing.
- Leases contain logical credential references or opaque redemption handles, never raw credential values.
- Immutable skill packages and other external inputs are downloaded and verified by trusted runner code before sandbox creation, then mounted read-only. Sandbox code receives no artifact-store credential or package download authority.
- Prefer backend-side credential injection and short-lived workload identity. Raw secret fallback cannot use a shared or checkpointed execution session.
- Raw tokens are forbidden in environment variables, argv, process titles,
shell history, and persistent files. Use an attempt-bound local credential
broker or an attempt-unique read-only
tmpfsfile when the process must receive a token; environment variables may carry only non-secret endpoint or path references. - Secrets are task-scoped and never included in workflow context, logs, artifacts, snapshots, or AI prompts.
- Sandboxed model access uses a runner-owned, peer/attempt-bound local broker with no reusable bearer visible to the worker or generated payload. Separate process identities/namespaces and descriptor controls prevent generated code from stealing the worker’s model capability; broker-side policy enforces model, budget, rate, cancellation, and expiry.
- TLS interception uses an operator-owned immutable trust-bundle digest and approved runtime adapters. Workflow code cannot add a CA or disable certificate verification.
- AI repair runs only in approved runner profiles and cannot publish, sign, or receive push credentials. A trusted diff enforces the server-owned protected path policy before a fixed branch or pull-request action can consume a patch.
- Publish and signing use fixed actions over immutable artifacts and require
verified provenance, digest-bound human approval, and task-scoped isolation.
Light-workflow dispatches typed
publishandsignrequests to a dedicated release-action service; branch and pull-request requests use a separate repository-action service. These credential-owning services receive exact immutable bindings and an idempotency key, while agents, sandboxes, runners, and workflow context receive no platform or signing credential. - Human approval waiting occurs only in the origin service,
light-workfloworlight-agent; no action lease, model channel, action credential, or secret-bearing task environment remains active. An eligible non-secret session workspace may use a separate bounded hold/checkpoint, and approval creates a fresh common attempt and fencing token. - Origin session close, revocation, or expiry creates a durable common cleanup request and promptly reclaims its physical sandbox; backend TTL is a last-resort bound rather than routine cleanup.
- Required build provenance is constructed and authenticated by trusted runner or control-plane code. Provenance signing material is inaccessible to tenant-controlled build steps.
- Tenant-authored jobs cannot mount host container-engine sockets.
- CPU, memory, disk, process, time, network, output, artifact, and concurrency limits are backend-enforced. A host-integrated backend that cannot prove a required limit cannot claim the task.
- Runtime drift causes quarantine, attempt fencing, credential revocation, cancellation, and backend cleanup.
- Unknown backend outcomes are reconciled before retry.
- All task results include runner identity, attempt, fencing token, effective policy digest, command template ID, backend identity and operation ID, artifact and provenance digests, workspace-change policy and patch digests where applicable, cleanup state, and approval references.
Implementation Plan
Phase 1: Contracts And Persistence
- Create
apps/light-workflow-runner. - Reuse
workflow-coremodels forrun.*task payloads. - Define strict versioned security metadata and immutable policy snapshots.
- Define local-cleanup, workspace-change, trust-bundle, credential-projection, capacity-queue, and provenance policy contracts.
- Define origin-neutral execution IDs, origins, and workflow-task, agent-turn, and agent-action subject types in the first protocol version.
- Add common scheduling request, execution attempt, execution session, immutable input, execution-session cleanup request, artifact, and append-only runtime-audit persistence. Keep workflow approval and agent approval/domain state under their origin services.
- Persist tenant, trigger principal, correlation ID, and policy snapshot at workflow start.
- Define runner registration, heartbeat, claim, renewal, cancellation, reconciliation, cleanup, and result APIs.
- Define the transactional identifiers-only result-ready PostgreSQL wakeup and authoritative startup/periodic origin catch-up query.
- Keep the existing
light-workflowevent consumer as the only workflow start consumer. - Keep unsupported
run.*tasks disabled.
Phase 2: Lease, Attempt, And Fencing
- Add attempt numbers, short-lived renewable leases, fencing tokens, and compare-and-set result acceptance.
- Add execution deadlines, cancellation, unknown-outcome state, backend operation IDs, and reconciliation.
- Add the durable local cleanup journal, disconnected watchdog, startup resource scan, and backend-native expiry handling.
- Add bounded per-tenant fair capacity queues, atomic slot reservation, long-poll claims, and capped jittered backoff.
- Add normalized results and bounded resumable log streaming.
- Emit the result-ready wakeup in the same transaction that stores a newly terminal common attempt; prove lost/duplicate notification recovery through indexed conditional origin acceptance.
- Prove that stale or duplicate reports cannot transition workflow or agent domain state and that a disconnected runner cleans resources without control-plane reachability.
Phase 3: Minimal Per-Task Execution
- Add server-side runner pools, execution profiles, capability matching, and approved command templates.
- Implement one
run.shelltemplate in a task-scoped sandbox. - Start with no credentials, no irreversible external effects, deny-all egress, an ephemeral workspace, and hard resource limits.
- Define the
ExecutionBackendinterface, capability document, server-owned compatibility record, and boundary-specific conformance suite. - Implement Cube Sandbox as the first task-scoped microVM backend, including idempotent prepare and execute, inspection, cancellation, and cleanup.
Phase 4: Artifacts, Sessions, And Additional Backends
- Add safe artifact export, trusted-side hashing, immutable storage, and trusted-side in-toto/SLSA provenance generation and authentication.
- Add workflow-scoped sessions with single-writer enforcement and copy-on-write isolation for parallel branches.
- Support agent-turn task scope and explicitly bounded agent-session reuse without assuming equal lifetimes; origin close/revoke/expiry must still trigger prompt backend cleanup.
- Add an execution-session state/fence and bounded
IDLE_APPROVAL_HOLDlifecycle distinct from action leases. Missing an active attempt must not clean a valid held session; hold expiry must not extend the fixed maximum. - Add idempotent pause/checkpoint/hold/resume commands and retained-resource quota, cost, evidence, and cleanup metrics.
- Compute effective session expiry as the minimum of origin, execution policy, broker/grant, and backend limits. Add durable origin-driven cleanup requests so close/revoke/expiry fences active work and destroys the sandbox promptly.
- Make trusted runner code download, verify, safely extract, stage, and mount immutable skill packages and other input records before sandbox creation; workers only revalidate mounted content.
- Add backend-specific production-baseline validation, including Cube authentication, private control-plane access, restricted inbound traffic, and deny-by-default egress.
- Add Docker Sandboxes for approved local or managed agent execution, requiring clone workspace mode for untrusted code and prohibiting the host Docker socket.
- Add a rootless OCI backend for explicitly trusted tasks and a Kubernetes Job backend only for approved runtime classes and node policies.
- Permit a runner to operate inside Toolbx only as a declared
host-integratedenvironment; do not initially require a separate Toolbx adapter or allow it to claim isolated tasks. - Add immutable TLS trust-bundle projection and approved Java, Node, Python, OpenSSL, and OS-store adapters.
- Add brokered credential delivery, attempt-bound local metadata service,
read-only
tmpfsfallback, and prohibit secret-bearing checkpoints. - Add protected runner-broker transports using a preconnected descriptor, peer-checked Unix-domain socket, vsock, or backend equivalent; separate the trusted worker/runtime from generated payload processes and prevent broker descriptor inheritance.
- Add periodic effective runtime snapshots.
- Compare runner-reported config with server-approved policy.
- Quarantine drifted runners, fence attempts, revoke credentials, and clean up backend resources.
Phase 5: Release And AI Workflows
- Add release-runner profile.
- Execute Java and Rust release build/test tasks through the runner.
- Add ConfigProfile manifest and
event-importerdry-run tasks. - Add AI failure analysis and bounded repair loops.
- Add server-owned runtime-tool manifests and placement-bound tool references.
Gateway tools intersect gateway
tools/list; runner tools intersect execution policy, leaseallowedTools, trusted manifest, and live local enumeration before the independently authorized sets are combined. - Add runner-staged immutable skill packages and runner-owned model inference brokering with model/data-boundary/budget enforcement outside the payload.
- Add protected-path policies, trusted post-export diff validation, and fixed branch or pull-request creation over accepted patches.
- Export immutable artifact sets with signed provenance and rebuild releases from the reviewed immutable commit after AI repair.
Phase 6: Publish And Signing
- Add fixed publish actions or a separate release service.
- Add an external signing service or fixed task-scoped signing action.
- Bind human approval to the artifact set, target, version, command template, policy digest, expiry, and single-use nonce.
- End the build action lease before the origin enters
WAITING_APPROVAL; create a fresh numbered domain/common fixed-action attempt, lease, monotonic fencing token, and grants only after approval. Apply the same contract tolight-agentapprovals. - Reconcile unknown outcomes before any retry or approval reuse.
Open Questions
- Should runner registration and task claim be direct HTTP APIs, WebSocket
messages through
controller-rs, or both? - Where should long-running task logs and artifacts be stored for SaaS deployments?
- How should the control plane attest VM-based runners that do not have a container image digest?
- Which backend-native TTLs and local-watchdog deadlines are required for each compatibility record?
- Which protected runner-local broker transport is supported first on each backend: preconnected descriptor, peer-checked Unix-domain socket, vsock, or backend-native equivalent?
- Which protected repository paths belong in the default agent policy, and how are repository-specific additions approved?
- Which attestor, signing identity, storage convention, and target SLSA Build level should release profiles use?
- Which portal or service owns immutable execution profiles, command templates, backend compatibility records, conformance evidence, and their approval lifecycle?
- Should all publish and signing operations use a separate release service, or should a small set of fixed runner actions be supported?
- How much of the existing
TaskExecutorshould move into shared crates solight-workflowandlight-workflow-runnercan share evaluation and result handling without sharing orchestration responsibilities?
Recommendation
Create light-workflow-runner as a separate executable and keep
light-workflow as the single SaaS-owned orchestrator. The runner should be a
fenced leased execution agent, not a workflow starter or workflow definition
loader. Integrations such as Cube Sandbox, Docker Sandboxes, rootless OCI,
approved Kubernetes runtimes, dedicated VMs, and fixed external actions belong
behind ExecutionBackend in the runner. Toolbx is recorded as a trusted
host-integrated runner environment, not advertised as a sandbox.
This gives tenants a practical way to run workflow tasks near their own APIs, gateways, repositories, clusters, and sandboxes while keeping workflow start events, policy decisions, task visibility, and audit under the SaaS control plane. Publish and signing remain fixed, approval-bound operations over immutable artifacts rather than arbitrary commands with release credentials.
References
- Execution Backends And Sandbox Execution Design
- SLSA v1.2 Build Provenance
- SLSA v1.2 Build Requirements
- in-toto Attestation Framework
Light-Axum Implementation
Implementation plans in this section cover concrete service examples and
runtime integration work built on light-axum.
Insurance Claim MCP Server Example
This plan describes a new light-example-rs MCP server example for the
insurance claim workflow demo. The server should be built on light-axum so it
uses the same runtime startup, config-server bootstrap, service registration,
logging control, TLS, and graceful shutdown pattern as the existing REST demo
APIs.
Source Review
The product workflow doc at
docs/src/product/light-workflow/insurance-claim-agentic-workflow.md defines two
execution variants:
- the REST variant calls demo APIs directly
- the MCP variant calls the same capabilities through
light-gateway
The existing light-example-rs apps provide two REST APIs:
apps/demo-customer-profile-apiapps/demo-offer-decision-api
The current MCP workflow definition,
apps/light-workflow/examples/insurance-claim-mcp-v1.yaml, currently calls MCP
tools for capabilities that already exist as REST endpoints. That workflow
should be replaced. The new version should deliberately mix both integration
styles:
- REST calls remain responsible for existing demo API capabilities.
- MCP calls cover only the functional gaps that are not already implemented by the REST APIs.
Functional Gaps
The workflow doc names a broader insurance tool set than the existing REST demo APIs currently provide. The gaps fall into three groups.
1. No Native Backend MCP Server
light-gateway can expose REST APIs as MCP tools and can proxy backend MCP
servers, but light-example-rs does not yet include a backend MCP server. This
means the demo does not prove the end-to-end path:
light-workflow call:mcp
-> light-gateway /mcp
-> backend MCP server built with light-axum
-> tool implementation
The new example should fill this first.
2. Coverage And Liability Are Still Agent Mock Output
The workflow currently uses a native coverage-liability-agent task with
mockOutput for:
- coverage status
- liability status
- risk level
- estimated loss
- deductible
- adjuster review flag
- SIU review flag
The product doc lists coverage-review tools such as evaluate_coverage,
score_claim_risk, and classify_liability, but the concrete workflow does not
call those as MCP tools yet. A backend MCP server can make this part
deterministic and testable.
3. Settlement Support Is Still Too Coarse
The workflow uses a native settlement-agent with mockOutput, then calls
recommendSettlement. The existing offer decision API returns a settlement
recommendation, but it does not expose separate tools for:
- required documents
- customer-facing summary generation
- repair versus total-loss explanation
- denial-draft explanation
The MCP server should fill those smaller support functions. It should not
reimplement recommendSettlement, because that endpoint already exists in
demo-offer-decision-api.
Decisions
- The backend MCP server is stateful.
- Session state is in memory for the demo.
- The server returns and validates
Mcp-Session-Id. - The first implementation exposes camelCase tool names only.
- The MCP server implements only gap-filling tools.
- Existing REST API capabilities stay in the REST APIs and are not duplicated.
- Small duplicated fixtures or deterministic rule tables are acceptable if they make the example faster to deliver.
- The existing
insurance-claim-mcp-v1.yamlworkflow should be replaced rather than copied into a second MCP workflow version.
Proposed App
Create a new app in light-example-rs:
apps/demo-insurance-claim-mcp-server/
Cargo.toml
src/main.rs
config/
client.yml
portal-registry.yml
server.yml
startup.yml
values.yml
Suggested service identity:
server.serviceId: com.networknt.demo.insurance-claim-mcp-1.0.0
server.environment: demo
server.httpPort: 8087
server.enableHttp: true
server.enableRegistry: true
For local standalone development, server.enableRegistry can be overridden to
false. For the full demo, it should register with controller discovery so
light-gateway can resolve it by serviceId.
Runtime Shape
The app should follow the same pattern as the REST examples:
#![allow(unused)]
fn main() {
#[derive(Clone, Default)]
struct InsuranceClaimMcpApp;
#[async_trait]
impl AxumApp for InsuranceClaimMcpApp {
async fn router(&self, _context: ServerContext) -> Result<Router, RuntimeError> {
Ok(build_router())
}
}
}
The server should expose:
GET /healthPOST /mcpDELETE /mcpfor session cleanup
Use LightRuntimeBuilder::new(AxumTransport::new(InsuranceClaimMcpApp)) and
the same config-dir environment override pattern used by the REST demos.
MCP Protocol Scope
Keep the first server deliberately small:
- support JSON-RPC
initialize - support
notifications/initialized - support
tools/list - support
tools/call - issue an in-memory
Mcp-Session-Idfrominitialize - require later
tools/listandtools/callrequests to send a knownMcp-Session-Id - support
DELETE /mcpto remove the in-memory session - return JSON-RPC errors for unknown methods, unknown tools, invalid arguments, and tool execution failures
Streaming can be deferred. The first version can return normal JSON responses
from POST /mcp.
Tool Catalog
Implement only the tools that fill gaps in the current demo.
| Tool | Purpose |
|---|---|
evaluateCoverage | Determine whether the incident date, policy status, and vehicle coverage allow the claim to continue. |
classifyLiability | Classify liability as clear, unclear, contested, or external-party based on claim facts. |
scoreClaimRisk | Produce risk level and SIU recommendation from prior claims, injury, drivable status, and claim facts. |
listRequiredDocuments | Return required documents for repair, total-loss review, denial draft, or more-information path. |
generateCustomerSummary | Produce a deterministic customer-facing summary from claim, coverage, triage, and settlement context. |
Do not implement these existing REST API capabilities in the MCP server:
getCustomerProfilegetCustomerPreferencesgetCustomerPoliciesgetCoveredVehiclelistPriorClaimstriageClaimrecommendSettlement
The workflow should still show bounded agents. The MCP tools provide deterministic coverage, liability, risk, document, and summary support that the agents can reason over.
Data Strategy
Because the MCP server does not duplicate REST endpoints, it does not need to own the full customer, policy, vehicle, prior-claim, triage, or settlement data sets. The workflow passes the REST API outputs into MCP gap tools as tool arguments.
Small duplicated constants are acceptable for speed, for example:
- coverage rule thresholds
- liability classification labels
- risk scoring thresholds
- document templates
- customer summary text templates
Avoid creating a shared fixture crate unless duplication becomes hard to maintain.
Gateway Configuration
The full demo path should configure light-gateway with an apiType: mcp
backend target. The gateway remains the public MCP endpoint used by
light-workflow; the new server is the backend MCP implementation.
Conceptual target:
mcp-router.enabled: true
mcp-router.path: /mcp
mcp-router.tools:
- name: evaluateCoverage
apiType: mcp
serviceId: com.networknt.demo.insurance-claim-mcp-1.0.0
envTag: demo
path: /mcp
Repeat the tool entries for the gap-filling tools. Access-control rules and
response filtering should remain enforced at light-gateway.
Implementation Phases
Phase 1: App Skeleton
- add
apps/demo-insurance-claim-mcp-server - add workspace membership in
light-example-rs/Cargo.toml - implement
light-axumstartup, config-dir overrides, tracing,/health - add config files and config-registry values
- add release/build wiring consistent with the two existing demo APIs
Phase 2: Minimal MCP Protocol
- define JSON-RPC request, response, error, and MCP content/result structs
- implement in-memory session storage
- implement
initializewithMcp-Session-Id - implement
notifications/initialized - implement
tools/list - implement
tools/call - validate
Mcp-Session-Idon later requests - implement
DELETE /mcpsession cleanup - add request validation and JSON-RPC error mapping
- add unit tests for protocol errors
Phase 3: Gap-Filling Tools
- implement
evaluateCoverage - implement
classifyLiability - implement
scoreClaimRisk - implement
listRequiredDocuments - implement
generateCustomerSummary - keep output fields aligned with the replacement workflow assertions
- add handler tests for one success and one failure path per tool group
- verify with direct
POST /mcptools/listandtools/call
Phase 4: Gateway And Workflow Integration
- add demo
mcp-router.ymlentries that point to the MCP server byserviceId - enable registry for the MCP server in the full demo environment
- verify
light-gatewaycan initialize the backend MCP server - replace
insurance-claim-mcp-v1.yamlso it calls REST APIs for existing demo API capabilities and MCP tools for gap-filling capabilities - run the replaced
insurance-claim-mcp-v1.yamlflow throughlight-gateway - keep the existing REST workflow unchanged for comparison
Tests And Verification
Minimum verification:
cargo check -p demo-insurance-claim-mcp-server
cargo test -p demo-insurance-claim-mcp-server
Direct protocol checks:
curl -sS http://127.0.0.1:8087/health
curl -i -sS -X POST http://127.0.0.1:8087/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","clientInfo":{"name":"demo","version":"1.0.0"},"capabilities":{}}}'
curl -sS -X POST http://127.0.0.1:8087/mcp \
-H 'Content-Type: application/json' \
-H 'Mcp-Session-Id: <session-id-from-initialize>' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
Gateway checks:
curl -k -sS -X POST https://localhost:8443/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
Workflow checks:
- start
insurance-claim-mcp-v1 - confirm customer-context outputs are loaded through REST API calls
- confirm
triageClaimandrecommendSettlementare still REST API calls - confirm
evaluateCoverage,classifyLiability,scoreClaimRisk,listRequiredDocuments, andgenerateCustomerSummaryrun through MCP - complete the adjuster and claimant human tasks
- verify the final workflow output is
CLAIM_APPROVEDfor the happy path
Asymmetric Decryptor
asymmetric-decryptor decrypts RSA encrypted configuration values.
It is used by config-loader when a service loads encrypted values that use
the CRYPT:RSA: prefix. The crate supports RSA private keys in PKCS#8 and
PKCS#1 PEM formats and decrypts payloads with RSA-OAEP using SHA-256.
Main Types
AsymmetricDecryptor: owns the RSA private key and decrypts supported payloads.AsymmetricError: error type for prefix, base64, key, and decrypt failures.CRYPT_RSA_PREFIX: the requiredCRYPT:RSA:payload prefix.
Usage
#![allow(unused)]
fn main() {
use asymmetric_decryptor::AsymmetricDecryptor;
let decryptor = AsymmetricDecryptor::from_pem(private_key_pem)?;
let plaintext = decryptor.decrypt("CRYPT:RSA:...")?;
}
Notes
This crate is intentionally small. It does not fetch keys, rotate keys, or
perform configuration merging. Those concerns belong to config-loader and the
runtime layer.
Config Loader
config-loader loads, merges, resolves, and decrypts service configuration.
It provides the common configuration behavior used by fabric services and runtime modules. Configuration can be loaded from YAML, JSON, or TOML files, merged across layers, expanded from values maps, and decrypted when encrypted values are present.
Main Types
ConfigLoader: loads files and resolves${key:default}style values.ConfigManager<T>: stores hot-swappable typed configuration behind an atomic reference.ConfigError: shared error type for IO, parse, decrypt, and conversion failures.
Resolution Model
The loader supports:
- merging multiple config files in order
- external overlays through
LIGHT_RS_CONFIG_DIR - whole-value variable replacement
- embedded variable expansion inside strings
- typed deserialization through Serde
- symmetric encrypted values through
symmetric-decryptor - asymmetric encrypted values through
asymmetric-decryptor
Usage
#![allow(unused)]
fn main() {
use config_loader::ConfigLoader;
use std::collections::HashMap;
let loader = ConfigLoader::from_values(HashMap::new(), None, None)?;
let config: MyConfig = loader.load_typed(["config/my-service.yml"])?;
}
Consumers
light-runtime uses this crate for service bootstrap and runtime config.
Application crates can also use it for app-specific policy or domain config.
Hindsight Client
hindsight-client provides a small client abstraction for persistent agent
memory.
It stores and recalls memory units from PostgreSQL. The current implementation
uses sqlx and pgvector for vector similarity search.
Main Types
HindsightMemory: trait used by applications that need memory retention and recall without coupling to a specific database implementation.PgHindsightClient: PostgreSQL-backed implementation ofHindsightMemory.MemoryUnit: returned memory record with content, type, metadata, and bank identity.
Usage
#![allow(unused)]
fn main() {
use hindsight_client::{HindsightMemory, PgHindsightClient};
let memory = PgHindsightClient::new(pool);
let unit_id = memory
.retain(host_id, bank_id, "User prefers concise answers", "fact", None, metadata)
.await?;
}
Data Model
The PostgreSQL implementation writes to agent_memory_unit_t and uses
host_id plus bank_id to isolate memory between tenants, users, or sessions.
Consumers
light-agent uses this crate to persist and recall agent conversation memory.
Light Rule
light-rule is the Rust rule engine for evaluating rule definitions and
executing registered actions.
It is designed to align with the rule.yaml specification while remaining
runtime-neutral. Java services can use yaml-rule; Rust services use this
crate.
Main Types
RuleEngine: evaluates rule conditions and determines action execution.MultiThreadRuleExecutor: executes rules with runtime state.RuntimeState: input/output state passed through rule evaluation.ActionRegistry: registry for action plugins.RuleActionPlugin: trait implemented by Rust action handlers.Rule,RuleCondition,RuleAction,RuleConfig,EndpointConfig: rule model types.
Action Model
Rules reference actions by actionRef. In Rust, actionRef resolves to a
registered RuleActionPlugin; it is not a Java class name. This keeps the rule
format portable across Java and Rust executors.
Usage
#![allow(unused)]
fn main() {
use light_rule::{ActionRegistry, RuleEngine};
let registry = ActionRegistry::default();
let engine = RuleEngine::new(registry);
}
Related Design
See Light-Rule for the rule format and its relationship to workflow assertions and portal rule management.
Light Runtime
light-runtime is the shared service runtime for Light Fabric applications.
It owns bootstrap, configuration loading, transport startup, graceful shutdown,
and optional portal registry registration. Apps such as light-agent and
light-deployer should start through this crate instead of binding sockets
directly.
Main Types
LightRuntimeBuilder: builds a runtime from a transport.LightRuntime: configured runtime before start.RunningRuntime: running service handle with shutdown support.Module: lifecycle hook abstraction.RuntimeConfig: resolved runtime configuration.ServerConfig: HTTP/HTTPS bind and service identity settings.BootstrapConfig: remote config bootstrap settings.PortalRegistryConfig: portal registry connection settings.
Startup Pattern
#![allow(unused)]
fn main() {
use light_axum::AxumTransport;
use light_runtime::LightRuntimeBuilder;
let runtime = LightRuntimeBuilder::new(AxumTransport::new(app))
.with_config_dir("config")
.build();
let running = runtime.start().await?;
running.shutdown().await?;
}
Configuration
At minimum, runtime services need server.yml. Optional files include
startup.yml, client.yml, and portal-registry.yml.
Related Frameworks
light-runtime is transport-neutral. light-axum supplies the Axum transport
implementation.
MCP Client
mcp-client is a client for calling MCP-compatible gateway endpoints.
It provides a small API for listing and invoking tools through a configured MCP gateway path. It is intentionally focused on the client side; MCP server implementations live in applications or framework layers.
Main Types
McpGatewayClient: gateway client used by applications.McpTool: tool metadata returned by the gateway.McpContent: content item returned by MCP tool calls.McpToolCallResult: structured result for a tool invocation.
Usage
#![allow(unused)]
fn main() {
use mcp_client::McpGatewayClient;
let client = McpGatewayClient::new(gateway_url, path, timeout_ms);
let result = client.call_tool("tool.name", arguments).await?;
}
Consumers
light-agent uses this crate when an agent session needs to discover or invoke
tools exposed through an MCP gateway.
Model Provider
model-provider defines a common abstraction over LLM providers and implements
multiple provider adapters.
The goal is to let agent and workflow code depend on one Provider trait while
supporting local models, hosted APIs, and provider-specific features.
Main Types
Provider: async trait implemented by model providers.ChatRequest,ChatResponse,ChatMessage: common chat data model.ToolSpec,ToolCall: tool-calling model.ProviderCapabilities: capability metadata.TokenUsage: usage accounting.ReliableProvider: reliability wrapper.RouterProvider: route requests across multiple providers.
Provider Implementations
Current modules include:
- Anthropic
- Azure OpenAI
- Bedrock
- Claude Code
- Codex
- OpenAI-compatible providers
- Copilot
- Gemini
- Gemini CLI
- GLM
- Kilo Code CLI
- Ollama
- OpenAI
- OpenRouter
- Telnyx
Consumers
light-agent uses this crate to send chat requests and tool specs without
hard-coding a single LLM provider.
Portal Registry
portal-registry provides client support for registering services with Light
Portal or Light Controller.
It uses a JSON-RPC style WebSocket protocol for service registration, metadata
updates, discovery, and cache-management control. Runtime services normally use
this through light-runtime, but applications can also use the client directly
when they need custom registry behavior.
Main Types
PortalRegistryClient: WebSocket client for registry communication.RegistryHandler: trait for handling registry callbacks and messages.RegistrationState: client registration state.RegistrationBuilder: helper for constructing registration parameters.ServiceRegistrationParams: service identity and advertised endpoint.ServiceMetadataUpdate: metadata update payload.
Usage
#![allow(unused)]
fn main() {
use portal_registry::RegistrationBuilder;
let registration = RegistrationBuilder::new(
"com.networknt.service-1.0.0",
"1.0.0",
"http",
"127.0.0.1",
8080,
)
.with_env("dev")
.with_jwt(token)
.build();
}
Runtime Integration
light-runtime can register a service automatically when server.yml enables
registry support and portal-registry.yml supplies the portal connection.
Symmetric Decryptor
symmetric-decryptor decrypts legacy symmetric encrypted configuration values.
It supports payloads with the CRYPT prefix and decrypts AES-256-CBC data with
a key derived from the configured password using PBKDF2-HMAC-SHA256.
Main Types
Decryptor: trait implemented by decryptors.SymmetricDecryptor: password-based decryptor.DecryptError: error type for prefix, format, hex, and cipher failures.CRYPT_PREFIX: requiredCRYPTpayload prefix.
Usage
#![allow(unused)]
fn main() {
use symmetric_decryptor::{Decryptor, SymmetricDecryptor};
let decryptor = SymmetricDecryptor::new("password");
let plaintext = decryptor.decrypt("CRYPT:...")?;
}
Consumers
config-loader uses this crate when it encounters symmetric encrypted values
and a config password is available.
Workflow Builder
workflow-builder provides fluent builders for creating Agentic Workflow
definitions programmatically.
It depends on workflow-core for the actual model types and layers a builder
API on top so applications and tests can construct valid workflows without
manually assembling nested maps.
Main Areas
- workflow metadata construction
- authentication definitions
- task definitions
- nested
do,for,fork,try, and other task structures - YAML/JSON serialization through
workflow-coremodel types
Usage
#![allow(unused)]
fn main() {
use workflow_builder::services::workflow::WorkflowBuilder;
let workflow = WorkflowBuilder::new()
.use_dsl("1.0.0")
.with_namespace("lightapi")
.with_name("example")
.with_version("1.0.0")
.build();
}
Relationship To Workflow Core
Use workflow-core when you need direct access to the schema model. Use
workflow-builder when you want an ergonomic construction API.
Workflow Core
workflow-core contains the Rust model for the Agentic Workflow DSL.
The crate is schema-oriented: its structs and enums represent workflow documents, tasks, authentication blocks, durations, timeouts, errors, and supporting map types.
Main Areas
- workflow document metadata
- task definitions
- call task protocol definitions
- ask and assert task definitions
- duration and timeout models
- error definitions
- ordered map support for workflow task lists
Usage
#![allow(unused)]
fn main() {
use workflow_core::models::workflow::{
WorkflowDefinition,
WorkflowDefinitionMetadata,
};
let document = WorkflowDefinitionMetadata::new(
"lightapi",
"example",
"1.0.0",
Some("Example".to_string()),
None,
None,
None,
);
let workflow = WorkflowDefinition::new(document);
}
Consumers
workflow-builder builds on this crate. light-workflow and workflow-related
services use the model for loading, validating, and executing workflow
documents.
Light-Axum
light-axum adapts Axum applications to light-runtime.
Applications implement AxumApp and return an axum::Router. The framework
owns binding, optional TLS, runtime metadata resolution, and graceful shutdown
through the runtime transport contract.
Main Types
AxumApp: trait implemented by an application.AxumTransport: transport passed toLightRuntimeBuilder.ServerContext: runtime context passed into the app when building routes.AxumBoundHandle: running Axum server handle.
Pattern
#![allow(unused)]
fn main() {
use light_axum::{AxumApp, AxumTransport, ServerContext};
use light_runtime::LightRuntimeBuilder;
#[derive(Clone)]
struct App;
impl AxumApp for App {
fn router(&self, _context: ServerContext) -> axum::Router {
axum::Router::new()
}
}
let runtime = LightRuntimeBuilder::new(AxumTransport::new(App))
.with_config_dir("config")
.build();
}
Consumers
light-agent and light-deployer use this framework.
Building REST APIs With Light-Axum
light-axum lets a service use normal Axum routing while delegating listener
binding, TLS, config loading, runtime metadata, logging control, and graceful
shutdown to light-runtime.
The application owns the HTTP API shape. The framework owns how the service is started and managed.
Working Examples
The light-example-rs repository contains two REST API demos built with
light-axum:
| Demo | Purpose | Local port | OpenAPI |
|---|---|---|---|
apps/demo-customer-profile-api | Customer profile, preferences, policies, vehicles, and prior claims | 8085 | apps/demo-customer-profile-api/openapi.yaml |
apps/demo-offer-decision-api | Offer search, offer decisions, claim triage, and settlement recommendations | 8086 | apps/demo-offer-decision-api/openapi.yaml |
Both demos follow the same service pattern:
- Define request and response models with
serde. - Build a standard Axum
Router. - Implement
AxumAppand return that router. - Start the app through
LightRuntimeBuilder::new(AxumTransport::new(app)). - Keep runtime config in the app
config/directory. - Publish an OpenAPI document for endpoint import and API management.
Dependencies
A minimal REST API needs these crates:
[dependencies]
anyhow = { workspace = true }
async-trait = { workspace = true }
axum = { workspace = true }
light-axum = { workspace = true }
light-runtime = { workspace = true }
serde = { workspace = true }
tokio = { workspace = true }
tracing = { workspace = true }
Add serde_json when handlers accept or return dynamic JSON values.
Application Shape
Create a service type and implement AxumApp. The runtime passes a
ServerContext into router. Most simple REST APIs do not need it, but it is
available when routes need runtime metadata.
#![allow(unused)]
fn main() {
use async_trait::async_trait;
use axum::{Json, Router, routing::get};
use light_axum::{AxumApp, ServerContext};
use light_runtime::RuntimeError;
use serde::Serialize;
#[derive(Clone, Default)]
struct CustomerProfileApp;
#[async_trait]
impl AxumApp for CustomerProfileApp {
async fn router(&self, _context: ServerContext) -> Result<Router, RuntimeError> {
Ok(build_router())
}
}
#[derive(Debug, Serialize)]
#[serde(rename_all = "camelCase")]
struct HealthResponse {
status: &'static str,
service: &'static str,
}
fn build_router() -> Router {
Router::new().route("/health", get(health))
}
async fn health() -> Json<HealthResponse> {
Json(HealthResponse {
status: "UP",
service: "demo-customer-profile-api",
})
}
}
Everything inside build_router is standard Axum. Use Path, Query,
State, Json, HeaderMap, middleware, extractors, and response types the
same way you would in a standalone Axum service.
Runtime Startup
Start the service through LightRuntimeBuilder instead of binding a
TcpListener directly.
use anyhow::{Context, Result};
use light_axum::AxumTransport;
use light_runtime::{LightRuntimeBuilder, TracingOptions, init_tracing};
use tracing::info;
const CONFIG_DIR_ENV: &str = "CUSTOMER_PROFILE_CONFIG_DIR";
const EXTERNAL_CONFIG_DIR_ENV: &str = "CUSTOMER_PROFILE_EXTERNAL_CONFIG_DIR";
const LOG_ANSI_ENV: &str = "CUSTOMER_PROFILE_LOG_ANSI";
const DEFAULT_CONFIG_DIR: &str = "apps/demo-customer-profile-api/config";
const DEFAULT_EXTERNAL_CONFIG_DIR: &str =
"apps/demo-customer-profile-api/config-cache";
#[tokio::main]
async fn main() -> Result<()> {
let tracing_guard = init_tracing(
TracingOptions::new("demo-customer-profile-api")
.with_legacy_ansi_env(LOG_ANSI_ENV),
)
.context("failed to initialize tracing")?;
let config_dir =
std::env::var(CONFIG_DIR_ENV).unwrap_or_else(|_| DEFAULT_CONFIG_DIR.to_string());
let external_config_dir = std::env::var(EXTERNAL_CONFIG_DIR_ENV)
.unwrap_or_else(|_| DEFAULT_EXTERNAL_CONFIG_DIR.to_string());
let runtime = LightRuntimeBuilder::new(AxumTransport::new(CustomerProfileApp))
.with_config_dir(config_dir)
.with_external_config_dir(external_config_dir)
.with_logging_control(tracing_guard.logging_control())
.build();
let running = runtime
.start()
.await
.context("failed to start demo customer profile API")?;
info!("demo customer profile API started");
tokio::signal::ctrl_c()
.await
.context("failed to listen for shutdown signal")?;
running
.shutdown()
.await
.context("failed to shut down demo customer profile API")?;
Ok(())
}
This startup path gives the application the same runtime behavior as other Light services:
- listener configuration comes from
server.yml - local and external config directories are resolved by
light-runtime - TLS is controlled by runtime config, not by route code
- graceful shutdown goes through the runtime handle
- logging can be controlled by the runtime logging control object
Routing Patterns
The customer profile demo shows read-only REST endpoints:
#![allow(unused)]
fn main() {
fn build_router() -> Router {
Router::new()
.route("/health", get(health))
.route("/customers/{customer_id}", get(get_customer))
.route(
"/customers/{customer_id}/preferences",
get(get_customer_preferences),
)
.route(
"/customers/{customer_id}/policies",
get(get_customer_policies),
)
.route(
"/customers/{customer_id}/vehicles/{vehicle_id}",
get(get_covered_vehicle),
)
.route(
"/customers/{customer_id}/prior-claims",
get(get_prior_claims),
)
.with_state(AppState::seeded())
}
}
The offer decision demo shows query parameters, request bodies, headers, and shared mutable state:
#![allow(unused)]
fn main() {
fn build_router() -> Router {
Router::new()
.route("/health", get(health))
.route("/offers", get(search_offers))
.route("/offer-decisions", post(record_offer_decision))
.route("/claim-triage", post(triage_claim))
.route("/settlement-recommendations", post(recommend_settlement))
.with_state(AppState::seeded())
}
}
Use typed handlers for predictable API behavior:
#![allow(unused)]
fn main() {
async fn search_offers(
State(state): State<AppState>,
Query(query): Query<OfferQuery>,
) -> Json<Vec<Offer>> {
Json(state.search_offers(&query))
}
}
For request bodies:
#![allow(unused)]
fn main() {
async fn record_offer_decision(
State(state): State<AppState>,
headers: HeaderMap,
Json(request): Json<OfferDecisionRequest>,
) -> Result<Json<OfferDecisionResponse>, ApiError> {
// validate request and return a typed API response
}
}
Errors
Define one service error type and implement IntoResponse. This keeps handlers
small and ensures failures return stable JSON.
#![allow(unused)]
fn main() {
use axum::{
Json,
http::StatusCode,
response::{IntoResponse, Response},
};
use serde::Serialize;
#[derive(Debug, Serialize)]
#[serde(rename_all = "camelCase")]
struct ErrorResponse {
code: &'static str,
message: String,
}
#[derive(Debug)]
struct ApiError {
status: StatusCode,
code: &'static str,
message: String,
}
impl IntoResponse for ApiError {
fn into_response(self) -> Response {
(
self.status,
Json(ErrorResponse {
code: self.code,
message: self.message,
}),
)
.into_response()
}
}
}
The customer profile API returns 404 with CUSTOMER_NOT_FOUND. The offer
decision API returns 400 with INVALID_DECISION_REQUEST when request content
is invalid.
Configuration
Each application should keep its runtime configuration under its own config/
directory:
apps/<service-name>/
Cargo.toml
openapi.yaml
src/main.rs
config/
client.yml
portal-registry.yml
server.yml
startup.yml
values.yml
server.yml controls the listener and service identity:
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8085}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
tlsCertPath: ${server.tlsCertPath:}
tlsKeyPath: ${server.tlsKeyPath:}
serviceId: ${server.serviceId:com.networknt.demo.customer-profile-1.0.0}
enableRegistry: ${server.enableRegistry:false}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:demo}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
Use a unique server.serviceId per service. The demos use:
com.networknt.demo.customer-profile-1.0.0com.networknt.demo.offer-decision-1.0.0
values.yml supplies defaults for template variables:
server.serviceId: com.networknt.demo.customer-profile-1.0.0
server.environment: demo
server.ip: 0.0.0.0
server.advertisedAddress: 127.0.0.1
server.httpPort: 8085
server.enableHttp: true
server.enableHttps: false
server.enableRegistry: false
server.startOnRegistryFailure: true
Enable registry integration only when the service should register with Portal
or Controller discovery. For local standalone development, keep
server.enableRegistry: false.
OpenAPI
Keep the OpenAPI document beside the service source. The OpenAPI file is the
contract used by API management and workflow tooling, while src/main.rs is the
runtime implementation.
The demo specs are:
light-example-rs/apps/demo-customer-profile-api/openapi.yamllight-example-rs/apps/demo-offer-decision-api/openapi.yaml
When adding or changing a route, update both the Axum router and openapi.yaml.
Use operation IDs that match the business action, such as
getCustomerProfile, searchOffers, or recordOfferDecision.
Running Locally
From the light-example-rs repository:
cargo run -p demo-customer-profile-api
Then verify the health endpoint:
curl http://127.0.0.1:8085/health
Run the offer decision API the same way:
cargo run -p demo-offer-decision-api
curl http://127.0.0.1:8086/health
Override config locations with environment variables when running from a different working directory:
CUSTOMER_PROFILE_CONFIG_DIR=/path/to/config \
CUSTOMER_PROFILE_EXTERNAL_CONFIG_DIR=/path/to/config-cache \
cargo run -p demo-customer-profile-api
Checklist
Use this checklist when creating a new REST API with light-axum:
- create an app crate under
apps/ - add
axum,light-axum,light-runtime,tokio,serde, andasync-trait - define typed request, response, and error models
- build routes with a standard Axum
Router - implement
AxumAppfor the service type - start the service with
LightRuntimeBuilderandAxumTransport - add
config/server.yml,startup.yml,portal-registry.yml,client.yml, andvalues.yml - assign a stable
server.serviceId - add or update
openapi.yaml - verify
/healthand one representative business endpoint locally
Light-Axum IPv6 Support
Problem
light-axum binds application listeners from server.yml through the shared
light-runtime server configuration. The bind IP is configured with
server.ip, and most product templates default it to 0.0.0.0.
IPv4 wildcard binding works for IPv4-only networks, but dual-stack container and Kubernetes networks can publish both IPv4 and IPv6 service addresses. If a client resolves an Axum service name to IPv6 first, the service must either listen on IPv6 or the client must retry an IPv4 address. Relying on client fallback is not enough for gateway and service-to-service traffic.
Before IPv6 support, the transport built bind addresses with string concatenation:
#![allow(unused)]
fn main() {
format!("{}:{port}", config.server.ip).parse()
}
That works for 0.0.0.0:8080, but fails for IPv6 wildcard binding because
:: plus port becomes :::8080 instead of [::]:8080.
Goals
- Support IPv4 and IPv6 bind addresses for all applications using
AxumTransport. - Keep the existing default of
server.ip: 0.0.0.0. - Preserve current runtime configuration names and deployment templates.
- Fail early with a clear error when
server.ipis not a valid IP address.
Non-Goals
- Do not enable IPv6 by default.
- Do not change advertised address resolution or portal-registry registration.
- Do not change TLS behavior or application routing.
- Do not add client-side IPv4 fallback in this change.
Configuration
The bind address remains the existing server.ip property.
IPv4 wildcard:
server.ip: 0.0.0.0
IPv6 wildcard:
server.ip: "::"
Specific IPv4 address:
server.ip: 172.16.1.9
Specific IPv6 address:
server.ip: "fdd0:0:0:1::9"
External templates should continue projecting this value into server.yml:
ip: ${server.ip:0.0.0.0}
Implementation
light-axum parses server.ip as an IpAddr and constructs the listener with
SocketAddr::new(ip, port).
This keeps IPv4 output unchanged and produces bracketed IPv6 socket addresses where required by Rust networking APIs:
0.0.0.0 + 8080 -> 0.0.0.0:8080
:: + 8080 -> [::]:8080
The change lives in the framework transport, so it applies to products using
AxumTransport, including portal-service, light-agent, and
light-deployer.
Deployment Guidance
Only set server.ip: "::" when the host or container network is intended to
serve IPv6 traffic. In dual-stack deployments, verify both sides:
- the runtime receives an IPv6 address;
- DNS or service discovery returns reachable addresses;
- dependent clients or gateways can connect to the IPv6 endpoint;
- health checks cover the selected address family.
If the environment is IPv4-only, keep server.ip: 0.0.0.0.
Verification
For an Axum service configured with IPv6 wildcard binding:
server.ip: "::"
verify from a peer in the same network:
getent ahosts <service-name>
curl -k -g https://[<service-ipv6>]:<port>/health
If access goes through service DNS:
curl -k -v https://<service-name>:<port>/health
The first resolved address family must be reachable, or the caller must have a retry/fallback strategy.
Light-Pingora
light-pingora adapts Pingora proxy services to light-runtime.
It is the framework layer for high-performance gateway and proxy products. The crate keeps runtime concerns such as configuration and service lifecycle separate from Pingora-specific proxy behavior.
Role
- bridge Pingora services into the common runtime lifecycle
- expose transport metadata to
light-runtime - support gateway products without duplicating bootstrap code
Consumers
light-gateway uses this framework.
MSAL Exchange
The msal-exchange handler is a BFF security handler for SPA applications
that authenticate with Microsoft Authentication Library, MSAL, and need an
internal light-oauth security profile for gateway authorization.
The SPA obtains Azure MSAL tokens in the browser. It sends the MSAL ID token to the gateway for light-oauth token exchange. In the Azure authorization placement pattern, it also sends the MSAL access token during the exchange so the gateway can store it in a secure BFF cookie. The internal light-oauth token set is stored in secure BFF cookies and is used on later requests together with CSRF protection.
This page documents the current behavior and the token placement extension for
deployments that must keep the Azure MSAL access token in the downstream
Authorization header while forwarding the light-oauth token in a separate
header.
Use Cases
Use msal-exchange when:
- The UI is a browser SPA using MSAL.js.
- Azure Entra ID is the identity provider for the browser login.
- The gateway must exchange the Azure token for a light-oauth token containing the enterprise security profile and custom claims.
- The gateway must protect browser requests with HttpOnly cookies and CSRF.
- Downstream routing needs either the light-oauth token or the Azure MSAL token
in the
Authorizationheader.
Handler Placement
Enable the handler in the gateway handler chain before downstream routing and before handlers that depend on the authenticated principal.
Example:
handlers:
- exception
- cors
- msal-exchange
- header
- prefix
- router
chains:
bff:
- exception
- cors
- msal-exchange
- header
- prefix
- router
paths:
- path: /auth/ms/exchange
method: POST
exec:
- bff
- path: /auth/ms/exchange
method: OPTIONS
exec:
- bff
- path: /auth/ms/logout
method: POST
exec:
- bff
- path: /auth/ms/logout
method: OPTIONS
exec:
- bff
Exchange and logout are POST-only. Keep the OPTIONS routes permanently, with
cors before msal-exchange, so preflight is handled before the auth method
guard.
When the handler is active, the gateway needs these resolved config files:
msal-exchange.ymlsecurity-msal.ymlsecurity.ymlclient.yml
security-msal.yml validates Azure MSAL tokens. security.yml validates the
light-oauth tokens stored in BFF cookies. client.yml provides the
light-oauth token-exchange client configuration.
Exchange Flow
The exchange endpoint receives the Azure MSAL ID token from the SPA and creates the BFF session.
POST /auth/ms/exchange
Authorization: Bearer <azure-msal-id-token>
-> read the Azure MSAL ID token
-> verify the ID token with security-msal.yml
-> generate a CSRF value
-> call light-oauth with the token-exchange grant
-> verify the returned light-oauth access token with security.yml
-> set BFF cookies
-> return { "scopes": [...] }
The token-exchange request uses client.yml oauth.token.token_exchange.
The outgoing form body contains:
grant_type=urn:ietf:params:oauth:grant-type:token-exchange
subject_token=<azure-msal-id-token>
subject_token_type=urn:ietf:params:oauth:token-type:jwt
csrf=<generated-csrf>
subjectTokenType can be set in msal-exchange.yml. When it is blank, the
shared token client default from client.yml is used.
On success, the response body contains the scopes from the light-oauth token:
{
"scopes": ["scope1", "scope2"]
}
The exchange body is optional. A zero-length request is valid even when a
shared client declares Content-Type: application/json.
Session Cookies
The handler uses the same cookie contract as the stateless SPA auth handler.
| Cookie | HttpOnly | Description |
|---|---|---|
accessToken | true | light-oauth access token |
refreshToken | true | light-oauth refresh token, when returned |
msalAccessToken | true | Azure MSAL access token when authorizationToken is azure-msal |
csrf | false | Generated CSRF value |
userId | false | User id from uid, user_id, or sub |
userType | false | User type from userType |
roles | false | Base64 encoded role value, default user |
host | false | Host claim |
email | false | Email claim from eml |
eid | false | Enterprise id claim |
accessToken and refreshToken are HttpOnly so browser JavaScript cannot read
the light-oauth tokens. The SPA reads the non-HttpOnly csrf cookie and sends
it back with protected requests.
CSRF Validation
For normal protected requests, the handler validates the request CSRF value
against the csrf claim in the light-oauth access token.
CSRF source order:
X-CSRF-TOKENrequest header.Sec-WebSocket-Protocolvalue starting withcsrf.for WebSocket requests.csrfquery parameter.
If the CSRF value is missing or does not match the JWT claim, the request is rejected.
Token Placement
authorizationToken selects which token owns the downstream Authorization
header after the BFF session has been established.
Supported values:
| Value | Authorization header | Light-oauth token location | Use case |
|---|---|---|---|
light-oauth | Bearer <light-oauth-token> | Authorization | Existing enterprise BFF pattern |
azure-msal | Bearer <azure-msal-access-token> | lightTokenHeader, default X-Light-Token | Azure-whitelisted downstream systems, such as AWS Agent Core |
authorizationToken: light-oauth
This is the current default behavior.
After the exchange, the SPA calls the gateway with cookies and CSRF:
GET /api/orders
Cookie: accessToken=...; csrf=...
X-CSRF-TOKEN: <csrf>
The handler:
-> reads the light-oauth accessToken cookie
-> verifies it with security.yml
-> validates CSRF
-> refreshes the token if it is close to expiry
-> injects Authorization: Bearer <light-oauth-token>
-> continues the handler chain
Downstream services receive:
Authorization: Bearer <light-oauth-token>
This mode is appropriate when downstream services and MCP tools trust
light-oauth directly and expect fine-grained security claims in the normal
Authorization header.
authorizationToken: azure-msal
This token placement pattern uses both Azure and light-oauth tokens downstream.
At exchange time, the SPA sends the MSAL ID token in Authorization and the
MSAL access token in msalAccessTokenHeader, which defaults to
X-MSAL-Access-Token:
POST /auth/ms/exchange
Authorization: Bearer <azure-msal-id-token>
X-MSAL-Access-Token: Bearer <azure-msal-access-token>
-> verify the MSAL ID token with security-msal.yml
-> verify the MSAL access token with security-msal.yml
-> exchange the ID token for a light-oauth token
-> store the light-oauth token in accessToken
-> store the MSAL access token in msalAccessToken
For later protected requests, the SPA sends cookies and CSRF. The SPA does not
need to put the Azure access token in the browser request Authorization
header because the gateway reads it from the HttpOnly msalAccessToken cookie:
GET /agent/chat
Cookie: accessToken=...; msalAccessToken=...; csrf=...
X-CSRF-TOKEN: <csrf>
The handler:
-> read the MSAL access token from the msalAccessToken cookie
-> verify the MSAL access token with security-msal.yml
-> read the light-oauth accessToken cookie
-> verify the light-oauth token with security.yml
-> validate CSRF
-> refresh the light-oauth token if it is close to expiry
-> inject Authorization: Bearer <azure-msal-access-token>
-> inject X-Light-Token: Bearer <light-oauth-token>
-> continue the handler chain
Downstream systems receive both tokens:
Authorization: Bearer <azure-msal-access-token>
X-Light-Token: Bearer <light-oauth-token>
This mode is intended for systems that only allow Azure as the OAuth provider
for the normal Authorization header, while still needing the light-oauth
security profile for API and MCP authorization decisions.
The SPA should not read or send X-Light-Token itself. The gateway should
derive that header from the HttpOnly light-oauth cookie after CSRF validation.
That keeps the light-oauth token out of browser JavaScript.
If a downstream light-gateway is responsible for fine-grained authorization,
it must be configured to verify X-Light-Token as the light-oauth token or to
promote X-Light-Token to Authorization at a trusted boundary before the
normal security/access-control handlers run.
Configuration
Example default configuration:
enabled: ${msal-exchange.enabled:true}
exchangePath: ${msal-exchange.exchangePath:/auth/ms/exchange}
logoutPath: ${msal-exchange.logoutPath:/auth/ms/logout}
logoutCsrfEnforced: ${msal-exchange.logoutCsrfEnforced:false}
cookieDomain: ${msal-exchange.cookieDomain:localhost}
cookiePath: ${msal-exchange.cookiePath:/}
cookieSecure: ${msal-exchange.cookieSecure:false}
sessionTimeout: ${msal-exchange.sessionTimeout:3600}
rememberMeTimeout: ${msal-exchange.rememberMeTimeout:604800}
renewBeforeSeconds: ${msal-exchange.renewBeforeSeconds:90}
refreshSingleFlightWaitMs: ${msal-exchange.refreshSingleFlightWaitMs:5000}
refreshSingleFlightCacheMs: ${msal-exchange.refreshSingleFlightCacheMs:3000}
refreshSingleFlightMaxEntries: ${msal-exchange.refreshSingleFlightMaxEntries:10000}
cookieSameSite: ${msal-exchange.cookieSameSite:None}
cookieTimeoutUri: ${msal-exchange.cookieTimeoutUri:/}
subjectTokenType: ${msal-exchange.subjectTokenType:}
authorizationToken: ${msal-exchange.authorizationToken:light-oauth}
lightTokenHeader: ${msal-exchange.lightTokenHeader:X-Light-Token}
msalAccessTokenHeader: ${msal-exchange.msalAccessTokenHeader:X-MSAL-Access-Token}
msalAccessTokenCookie: ${msal-exchange.msalAccessTokenCookie:msalAccessToken}
Fields:
| Field | Default | Description |
|---|---|---|
enabled | true | Enables or disables the handler once it is active in the chain. |
exchangePath | /auth/ms/exchange | Endpoint that receives the Azure MSAL ID token and creates the BFF session. |
logoutPath | /auth/ms/logout | Endpoint that clears BFF cookies. |
logoutCsrfEnforced | false | Enforces readable csrf cookie versus X-CSRF-TOKEN validation on logout after environment-specific observe-only qualification. |
cookieDomain | localhost | Cookie domain for session cookies. |
cookiePath | / | Cookie path for session cookies. |
cookieSecure | false | Adds the Secure cookie attribute. Use true for HTTPS deployments. |
sessionTimeout | 3600 | Default max age in seconds for session cookies. |
rememberMeTimeout | 604800 | Max age in seconds for long-lived refresh-token cookies when light-oauth returns remember-me behavior. |
renewBeforeSeconds | 90 | Refresh the light-oauth access token when it expires within this window. |
refreshSingleFlightWaitMs | 5000 | Maximum wait time for concurrent refresh requests sharing the same refresh token. |
refreshSingleFlightCacheMs | 3000 | Short cache window for a successful refresh result. |
refreshSingleFlightMaxEntries | 10000 | Maximum refresh single-flight cache entries. |
cookieSameSite | None | Cookie SameSite attribute. Supported values are None, Lax, and Strict. |
cookieTimeoutUri | / | URI returned when the session expires and cannot be refreshed. |
subjectTokenType | blank | Optional token-exchange subject token type override. |
authorizationToken | light-oauth | Token to place in downstream Authorization: light-oauth or azure-msal. |
lightTokenHeader | X-Light-Token | Header used for the light-oauth token when authorizationToken is azure-msal. |
msalAccessTokenHeader | X-MSAL-Access-Token | Header that carries the Azure MSAL access token on the exchange request when authorizationToken is azure-msal. |
msalAccessTokenCookie | msalAccessToken | HttpOnly cookie used to store the Azure MSAL access token after exchange when authorizationToken is azure-msal. |
Invalid authorizationToken values should fail startup. lightTokenHeader
should not be Authorization; use authorizationToken: light-oauth for that
case. In azure-msal mode, msalAccessTokenHeader must not be
Authorization because Authorization carries the MSAL ID token on the
exchange endpoint. msalAccessTokenHeader must also be different from
lightTokenHeader.
Security Configuration
security-msal.yml validates Azure MSAL tokens. It is required when the handler
is active.
Example:
enableVerifyJwt: ${security-msal.enableVerifyJwt:true}
ignoreJwtExpiry: ${security-msal.ignoreJwtExpiry:false}
enableRelaxedKeyValidation: ${security-msal.enableRelaxedKeyValidation:false}
issuer: ${security-msal.issuer:}
audience: ${security-msal.audience:}
jwt:
clockSkewInSeconds: ${security-msal.jwt.clockSkewInSeconds:60}
Recommended settings:
- Set
issuerto the Azure tenant issuer when the tenant is known. - Set
audienceto the SPA client id or the expected Azure access-token audience. - Keep
ignoreJwtExpiry: falsein production. - Use the configured Microsoft JWK supported by the gateway security runtime.
security.yml remains the normal light-oauth verifier. It validates the
light-oauth access token stored in the accessToken cookie and provides the
principal used by gateway authorization logic.
SPA Integration
Initial exchange:
await fetch("/auth/ms/exchange", {
method: "POST",
credentials: "include",
headers: {
Authorization: `Bearer ${azureMsalIdToken}`
}
});
Initial exchange with authorizationToken: azure-msal:
await fetch("/auth/ms/exchange", {
method: "POST",
credentials: "include",
headers: {
Authorization: `Bearer ${azureMsalIdToken}`,
"X-MSAL-Access-Token": `Bearer ${azureMsalAccessToken}`
}
});
Subsequent requests with the existing light-oauth authorization pattern:
await fetch("/api/orders", {
credentials: "include",
headers: {
"X-CSRF-TOKEN": csrf
}
});
Subsequent requests with the Azure MSAL authorization pattern:
await fetch("/agent/chat", {
credentials: "include",
headers: {
"X-CSRF-TOKEN": csrf
}
});
In both patterns, the SPA must send cookies with credentials: "include".
In the Azure MSAL authorization pattern, MSAL.js is responsible for obtaining
the Azure access token before calling /auth/ms/exchange. The gateway stores
that access token in the HttpOnly msalAccessToken cookie, validates it on
later BFF requests, injects it into Authorization, and injects the
light-oauth token into lightTokenHeader.
Logout
Logout clears all BFF cookies managed by the handler:
POST /auth/ms/logout
Cookie: accessToken=...; csrf=...
X-CSRF-TOKEN: <csrf>
Send credentials and no request body. A zero-length request is also accepted
when a shared client sets Content-Type: application/json. On success the
handler returns 204 No Content, no response content type or body, and
deletion cookies for every cookie name the runtime can set.
A legacy GET or any other unsupported exchange/logout method returns 405,
ERR10008, and Allow: POST before token-server, cookie, or proxy side
effects. OPTIONS remains routed to CORS.
Error Handling
Important error codes:
| Code | Meaning |
|---|---|
ERR11647 | Required Azure MSAL bearer token is missing on the exchange endpoint or in the MSAL access-token cookie. |
ERR11648 | light-oauth token exchange failed. |
ERR10000 | Azure MSAL token or light-oauth token verification failed. |
ERR10036 | CSRF token is missing from the request. |
ERR10038 | CSRF claim is missing from the light-oauth token. |
ERR10039 | Request CSRF and token CSRF do not match. |
ERR10052 | Token response does not contain expires_in and the JWT has no usable exp. |
ERR10008 | Method is not allowed for the exchange or logout endpoint. |
ERR11649 | Logout CSRF cookie/header validation failed without exposing either value. |
Implementation Notes
Rust light-pingora and Java light-spa-4j use the same token placement
contract:
authorizationToken: light-oauthpreserves the existing behavior and injects the light-oauth token intoAuthorization.authorizationToken: azure-msalverifies the exchange request’s MSAL ID token and MSAL access token withsecurity-msal.yml, stores the MSAL access token inmsalAccessToken, injects it into downstreamAuthorization, and injects the light-oauth token intolightTokenHeader.lightTokenHeaderdefaults toX-Light-Tokenand must not beAuthorizationwhenauthorizationTokenisazure-msal.msalAccessTokenHeaderdefaults toX-MSAL-Access-Tokenand is used only on the exchange endpoint.msalAccessTokenCookiedefaults tomsalAccessTokenand is HttpOnly.
In azure-msal placement, the gateway requires the MSAL access-token cookie
only when a BFF session cookie is present. Requests without accessToken or
refreshToken cookies keep the existing pass-through behavior so public
endpoints are not forced to authenticate at this handler.
Light-Agent
light-agent is the interactive agent service in Light Fabric.
It provides a WebSocket chat interface, integrates with model providers,
invokes MCP tools through mcp-client, and stores conversation memory through
hindsight-client. The current executable implements the enterprise
API/MCP-oriented service path. Coding and personal-assistant support extend the
same durable agent domain through additional runtime profiles rather than
forking separate agent engines.
Execution Model
light-agent is a long-lived interactive session service. A logical agent does
not automatically receive its own container or VM.
Remote model calls and gateway-only API/MCP tools can remain in the service. Turns that need a local CLI model, shell, browser, filesystem, repository, private local MCP server, or other effectful tenant execution use a runner-managed backend selected from server-owned policy. High-value publish, signing, deployment, branch, and pull-request operations use fixed structured actions.
Tool availability is placement-specific: gateway catalog entries intersect
live gateway tools/list, while runner-local shell/filesystem/browser/local-MCP
entries intersect the execution profile, lease allowlist, approved runtime
manifest, and live local availability. The independently authorized sets can
be combined for the model, but a tool remains bound to one server-owned
placement and dispatcher.
Human approval ends the current action lease and credentials. A task sandbox is cleaned; an eligible non-secret coding-session workspace may instead use a separate bounded pause/checkpoint hold. The hold is not executable authority and cannot extend the session maximum lifetime.
The target profiles are:
- enterprise business agents: long-lived light-agent reasoning plus typed light-gateway API/MCP tools;
- coding agents: a bounded
light-agent-workerinside a runner-managed workspace sandbox, using a native or external agent runtime adapter; - personal assistants: light-agent reasoning plus a separately deployed
light-agent-channelfor messaging and proactive triggers, with an optional personal edge runner for local-device effects.
Codex, Pi, Claude Code, Gemini CLI, Kilo, Hermes, OpenClaw, and similar harnesses are integration candidates behind an agent-runtime adapter. They are not launched directly by the shared light-agent service or by light-workflow. Centralized skills are materialized for the selected profile, but never grant execution authority by themselves.
See Light-Agent Execution for session and turn durability, tool authorization, sandbox placement, deployment profiles, runtime adapters, channel ingress, workflow handoffs, and the origin-neutral runner contract shared with workflow execution. See Centralized Skills for profile-specific skill materialization.
Key Dependencies
light-runtimelight-axummodel-providermcp-clienthindsight-clientportal-registry
Runtime
The app follows the standard runtime pattern:
- load config from
config/ - implement an Axum app
- start through
LightRuntimeBuilder - optionally register through portal registry
LLM User and Agent Authorization
Policy lifetime
Published Agent configuration and derived gateway assignments remain valid until
replaced or explicitly revoked through an applied configuration update. They do
not expire with elapsed time. PERSISTENT is the publication default; legacy
BOUNDED and LOCAL_DEMO inputs remain accepted for compatibility, but new
publications do not emit runtimePolicy.refreshAfter or runtimePolicy.expiresAt.
Updated runtimes accept old snapshots containing these fields without enforcing
them. Gateway assignment expiresAt is likewise a legacy field only.
For disconnected operation, download the complete validated values.yml and
mount it read-only under an operator-controlled configuration directory. Disable
remote bootstrap/controller registration for a deliberately disconnected
deployment. Keep the file immutable until an intentional update; an embedded
content digest checks consistency but is not a trusted signature against an actor
who can rewrite both the content and its digest. File ownership and read-only
mounts provide the deployment boundary.
During Portal or Config Server downtime, keep the accepted configuration. Remote bootstrap already falls back to available local/cached values on connection or HTTP failure; persist that cache if restart recovery is required. A restart without any complete local configuration still cannot bootstrap during an outage. Identity, digest, schema, valid-from, and rollback checks continue to apply. Explicit revocation takes effect when delivered and applied; disconnected Agents cannot discover undelivered revocations.
This lifetime applies only to configuration. User and workload JWT expiry, temporary delegation grants, execution deadlines, and session limits remain enforced. OAuth token issuance, model providers and operational stores remain runtime dependencies; removing policy expiry does not make those services offline. Upgrade both runtimes and the Portal publisher before emitting expiry-free snapshots: older binaries require/enforce the legacy fields.
Status: Phases 1–3 implemented; deployment qualification remains Phase 4. The observations under “Current implementation and gaps” describe the original September 7, 2026 baseline. Linked phase reports document the implemented changes.
Problem and decision
An interactive chat connection can succeed while its first model request fails
with 403 permission_denied. The browser authenticates to light-agent as a
user, but the current agent model client sends the agent service token as
Authorization to llm-gateway. An endpoint protected by
req-access-light-portal.lightapi.net can require admin or host-admin, which
the service token does not contain.
Preserve the original user access token in Authorization and send the agent’s
own service token in X-Scope-Token. The gateway must independently validate
both tokens, authorize the user, enforce applicable agent-to-model assignments,
and audit both identities. Do not grant administrator roles to service tokens
or disable endpoint access control to make interactive inference work.
Scope
This design covers interactive user-delegated model calls from light-agent
to llm-gateway, including subsequent model calls within a tool loop, retries,
and streaming admission. The same contract should apply to other generation
endpoints when they support agent-mediated requests.
Existing direct-user inference and separately authorized workloads, such as Knowledge embeddings, retain explicitly defined admission profiles. Background jobs and A2A invocations without an original user access token use the proposed User, Application, And Workflow Authorization design. They must not fabricate a user identity or fall back to presenting the service token as the user; unattended grants require separate qualification.
Current implementation and gaps
| Area | Current behavior | Required change |
|---|---|---|
apps/light-agent/src/main.rs | build_model_provider passes the service token to CompatibleProvider; authenticated user authorization is retained for MCP calls | Carry current-turn user authorization into model calls and attach the service token separately |
crates/model-provider/src/compatible.rs | send_request places its configured API key in Authorization | Provide a typed, request-scoped gateway authentication path that can send both headers |
frameworks/light-pingora/src/security.rs | JWT verification selects Authorization, falling back to X-Scope-Token | Explicitly verify both tokens for dual-token requests; the fallback is not dual validation |
frameworks/light-pingora/src/access_control.rs | Endpoint rules evaluate the primary principal | Keep the verified user as the primary principal and expose verified workload context separately |
apps/light-gateway/src/main.rs | llm_billing_context derives principal, billing subject, and optional routeAlias from one principal | Separate user, workload, authorization, and accounting identities |
crates/llm-gateway/src/http.rs | bound_model_alias constrains a signed route alias | Add authoritative agent assignment enforcement; alias equality alone is insufficient |
These observations do not establish that the existing control-plane binding projection contains all required workload identity and assignment fields. Inventory that projection and its canonical schema before implementing it.
Existing delegation and selected tradeoff
crates/agent-delegation already defines signed lad1 credentials with caller
claims, agent actor/definition, audience, expiry, policy digests, and replay IDs.
The gateway’s authenticate_agent_delegation verifies them and consumes their
replay identity through shared storage. Its constructed principal deliberately
retains both client_id = agent_actor and user_id = caller_subject.
The current kinds are tools-list, tool-call, Knowledge-retrieve, and Knowledge-upload; there is no LLM inference kind. Knowledge currently mints 60-second tokens; the verifier accepts a maximum lifetime of 300 seconds. The current agent LLM path uses its registry credential, not a 60-second LLM delegation. Therefore this proposal does not replace an already qualified LLM delegation path, but it must account for the existing delegation infrastructure.
The selected user-requested contract forwards the original user JWT. This
exposes a reusable, potentially broad-authority bearer credential to another
service and is weaker in containment than an audience-bound, single-use
credential. TLS, redaction, and destination restrictions do not make those
properties equivalent. Confirm that the user’s token issuer authorizes the LLM
gateway as a recipient; do not disable audience checking to permit forwarding.
If that contract cannot be met, block this profile and design an exchange using
the existing delegation infrastructure, rather than silently changing token
semantics. Extending DelegationKind for LLM operations would require explicit
operation/alias binding and replay/retry semantics; it is not a second parallel
mechanism to add in this implementation.
Existing tools/Knowledge delegation remains unchanged. For the new LLM profile,
reject lad1 credentials before generic delegation authentication consumes
replay state: neither existing tool nor Knowledge authority authorizes LLM
inference. Test this route/profile separation explicitly.
agentPolicy.gatewayDelegation already exists as an untyped Value and has no
runtime consumer in the agent. Type and validate this existing projection slot
for per-agent gateway policy, trusted destination, credential source, and claim
requirements. Do not introduce a parallel per-agent configuration surface or
put bearer tokens in published policy. Schema changes must participate in the
existing immutable snapshot/digest and publisher compatibility workflow.
Request and trust contract
UI -> light-agent
Authorization: Bearer <original user access token>
light-agent -> llm-gateway
Authorization: Bearer <original user access token>
X-Scope-Token: Bearer <agent service token>
llm-gateway -> inference provider
Provider-specific credentials owned by llm-gateway
Neither inbound token is forwarded to the inference provider. Request bodies,
prompts, tools, and caller-supplied identity headers cannot override either
verified principal. The agent ignores any UI-supplied X-Scope-Token and injects
its configured service credential.
Maintain separate typed user and workload principals. The user principal
contains the verified user identifier, including the deployment’s established
uid mapping, roles, issuer, and host. The workload principal contains the
verified service subject/client identity, issuer, host, and applicable audience,
environment, and scopes. Never merge workload claims into user roles.
Reuse the existing control-plane agent/service identity and alias binding
projection to resolve the workload to its agent definition within the
host/environment. Inventory the exact canonical identity format before coding;
do not introduce a new workload mapping store. An
arbitrary request agentDefId is not evidence of identity. If one credential
maps to multiple logical agents, fail ambiguous resolution until a trusted,
verifiable discriminator is defined; a caller-selected agent ID is insufficient.
Agent request lifecycle
Capture the original access token from the authenticated invocation and pass it through the current turn’s model request context. Keep it out of shared client default headers and globally cached provider objects. A provider created for one turn may retain that turn’s credentials only for its bounded lifetime.
Every tool-loop continuation and retry carries the same explicit user/workload
pair. Today authenticate_request runs at WebSocket upgrade, and handle_socket
retains that AuthenticatedRequest for subsequent turns. A new turn on the same
socket does not refresh authentication. Implement UI token renewal followed by
reauthenticated reconnect before accepting another turn after expiry. Preserve
session ownership and queued-turn identifiers across reconnect, reject identity
changes for an existing session, and prevent duplicate turn submission. Check
expiry before every outbound model call; renewal cannot extend an already
captured token. An in-band replacement protocol is outside this initial scope. Tokens are
not persisted in conversation history, prompts, checkpoints, logs, or audit
records. Gateway destinations remain trusted configuration; redirects must not
forward either credential to another origin.
If the user token expires before a later model call, reject that call and let the UI renew authentication through its established flow. An open chat connection does not extend the token’s authority. Do not automatically repeat a potentially billable inference after ambiguous completion.
Workload credential issuance and renewal
The current agent credential comes from registry_token and is retained in
AppState.llm_gateway_token at startup. It is not evidence of a gateway-audience
credential, and a static expired token cannot recover through retries.
Before enabling this profile, implement an issuer-supported service credential acquisition and refresh path. Obtain a token with the gateway’s accepted audience, agent identity, host/environment, and required scopes; cache it until a safety margin before expiry, refresh with bounded concurrency, and rotate the credential used for acquisition. Existing valid cached credentials may serve only until expiry. After expiry or failed initial acquisition, stop new inference with an explicit credential-unavailable outcome; do not retry the stale startup token or turn off expiry checks. Keep registry authentication separate unless its issued token is explicitly qualified for both purposes.
The issuer endpoint, supported grant, audience value, and source of refresh credentials are Phase 1 deliverables, not assumptions that current registry configuration already satisfies. Prove renewal across an actual expiry boundary before rollout.
Gateway admission sequence
Retain X-Scope-Token to preserve the requested header contract and avoid a new
header allowlist/configuration surface. A distinct header would remove the need
to distinguish its two wire forms, but this design accepts that compatibility
cost explicitly. Select parsing behavior from trusted route configuration, not
from caller input or a fallback after failed verification.
- Select a route admission profile and parse headers before authentication.
The current
request_headeraccessor returns one value and cannot establish uniqueness. Inspect all values from the underlying header map and reject duplicates, comma-combined credentials, malformed values, and conflicts. Add a shared tested Bearer parser for this dual-token path: the current scope fallback passes its header directly to JWT verification without removingBearer. Do not route the new wire form through that fallback. Explicitly preserve and regression-test legacy raw scope-token profiles separately. - Validate the user token and, when supplied or required, the workload token independently: signature, trusted issuer, expiry/not-before, applicable audience, and required identity claims. A supplied invalid scope token is never ignored because the user token is valid.
- Resolve and validate the workload identity. Check its required scope, environment, and host against the user and resource host. Audience and environment requirements must come from an explicit issuer/profile contract.
- Evaluate
req-access-light-portal.lightapi.netagainst the user principal. A valid agent credential cannot compensate for a denied user. - Parse and resolve the requested model under the active LLM policy snapshot.
Apply any existing signed
routeAliasrestriction as an additional bound. - Build the explicit execution identity described below and evaluate existing internal-alias binding with it; do not derive it from the user JWT.
- Record the admission decision and its policy revision, then enter the existing quota, budget, routing, and inference path.
Protected routes must not permit mock identity or expiry bypass in qualification
or production. In particular, deployments currently using
security.ignoreJwtExpiry: true need current credentials and enforced expiry
before this contract can be qualified.
Keep the existing primary-principal API compatible where possible, adding a
separate verified workload context for LLM admission rather than changing the
meaning of AuthPrincipal for unrelated handlers.
Execution identity and existing alias bindings
The runtime already implements assignment enforcement through alias internal
and boundPrincipal. The compiler requires a principal for an internal alias;
buffered generation, streaming, embeddings, probes, and model listing use that
binding. Reuse these fields and the existing control-plane publication path.
Do not add a second assignment database or a competing runtime permission map.
Phase 1 verifies the publisher’s exact agent identity representation and missing
wiring; it does not recreate the already implemented binding mechanism.
Changing only Authorization is unsafe: llm_billing_context currently prefers
client_id, and its principal_id drives internal alias access, per-principal
concurrency/rate buckets, and default billing. Use an explicit adapter:
| Consumer | Dual-token source |
|---|---|
| Endpoint RBAC/CEL and user audit | Verified user principal; preserve established uid/role mapping |
Internal alias matching and runtime principal_id | Verified agent identity in the exact existing boundPrincipal representation |
| Per-principal limits | Same canonical agent execution principal as binding; retain configured limits |
| Billing subject | Existing trusted agent billing attribution and fallback to execution principal; no automatic switch to UI client or human |
| Workload audit | Verified service identity plus authoritatively resolved agent definition |
bound_model_alias | Effective verified route restriction, independent of primary RBAC principal |
Never let untrusted body fields or the user’s client_id select the execution
principal. Any desired future user-level budget is additive and needs its own
policy; it must not replace the agent bucket accidentally.
Move route restriction extraction out of the single-principal helper. Preserve
routeAlias from the verified workload credential, and intersect it with any
applicable verified user restriction and existing control-plane alias binding.
Conflicting restrictions deny. A missing user routeAlias cannot erase the
workload restriction. A deployment that previously relied on a signed alias
claim must issue it in the replacement workload credential or have an explicitly
qualified equivalent binding; absence must not broaden access during migration.
For direct-user profiles, preserve existing identity and listing behavior. An
internal alias queried without an authorized workload remains indistinguishable
from an unknown alias, using the existing AliasNotFound behavior. Header
omission must not turn an internal alias public. Intentionally public aliases
retain their current policy; invalid supplied workload credentials still fail.
agentDefault selects a default and is not a separate authorization grant. Local
immutable agent model policy may further restrict selection but cannot expand
gateway access. Fallback deployments stay under the authorized alias policy.
Every request uses one snapshot generation; later tool-loop requests see newly
activated bindings. Immediate cancellation of admitted streams is outside scope.
Errors and audit
Use 401 for invalid supplied credentials or credentials required by a route
profile independent of the requested model, 403 for authenticated callers
denied by endpoint or workload-context policy, and 503 for unavailable required
authorization state. For inaccessible internal aliases, including absent workload
identity on a direct-user profile, preserve AliasNotFound just as for unknown
aliases. Do not choose a missing-token 401 based on model lookup: that would
reveal that the model exists and requires an agent. Preserve the established public
error envelope and avoid disclosing restricted model existence. Detailed denial
reasons belong in internal audit and diagnostics with a correlation ID.
Audit admission denials as well as successful and failed inference completion. Record:
- Correlation/request ID and available agent turn/session reference.
- Verified user identity and issuer; verified workload service identity and issuer; resolved agent definition and host/environment.
- Requested alias, effective model policy, selected deployment when available, and assignment/snapshot revision.
- Separate user-access and agent-assignment decisions and stable reason codes.
- Completion status, token usage, cost, and explicitly resolved billing subject.
Do not store raw tokens or unnecessary claims. A failed token check must not label decoded claims as verified identity. A client-supplied turn identifier is correlation metadata, not authorization evidence.
Apply the execution-identity table above to accounting and limits, and record both actors independently. Regression tests must assert exact bucket and billing keys, not merely that a request succeeds.
Extend the audit schema and WAL serialization together, with a compatibility and migration strategy for persisted events. An observed local audit WAL decoding failure must be diagnosed and repaired before live audit qualification; it is a deployment observation, not proof of the cause of the original 403. Honor each route’s configured audit delivery guarantees and prove eventual delivery for retained WAL records after recovery.
Implementation phases
Phase 1: Contracts and control-plane projection
The Phase 1 contract records the implemented typed projection, compatibility checks, issuer constraints, and reauthentication contract. Runtime activation remains gated on later phases.
Inventory existing delegation, canonical alias publication, and identity formats.
Type the existing agentPolicy.gatewayDelegation slot, identify actual projection
gaps, and specify issuer-supported token renewal and UI reconnect behavior.
Validate audience compatibility for both credentials and preserve execution,
route-binding, quota, and billing identity contracts. Add only demonstrated
schema/projection gaps with publication and digest compatibility tests.
Exit: existing alias bindings round-trip with the intended agent principal; credential acquisition/renewal and user reauthentication contracts are concrete. No new assignment store is introduced.
Phase 2: Gateway verification and enforcement
Implemented: see Phase 2 gateway implementation for the activation prerequisites, compatibility contract, and qualification gate.
Implement independent token validation, separate principal context, user rule evaluation, assignment admission, and structured audit. Preserve unrelated single-token admission profiles and provider credential isolation.
Exit: gateway integration tests prove allowed and denied cases before any provider invocation, including real audit persistence.
Phase 3: Agent forwarding
Implemented: see Phase 3 Agent forwarding for lifecycle behavior, the repeatable gate, and remaining issuer/deployment prerequisites.
Add request-scoped dual-token support to the gateway model client and thread the user context through interactive turns, retries, and tool loops. Give noninteractive paths an explicit unsupported or separately authorized outcome.
Exit: concurrent requests cannot exchange credentials, the mock gateway sees the correct pair, and reconnect and workload renewal pass expiry-boundary tests.
Phase 4: Deployment and live qualification
See Phase 4 rollout and qualification for the issuer change, executable checks, ordered rollout, rollback, and outstanding live evidence.
Deploy compatible gateway verification and projection support first, then agent forwarding, then activate the required restricted-model policy. Prepare valid service/user tokens and enforced expiry before activation. During transition, any supplied second token is validated; no compatibility mode may ignore an invalid credential or bypass an active assignment restriction.
Rollback must preserve enforcement: a gateway version that cannot enforce the active policy must not serve it. Revert compatible policy and application versions together through the normal publication/deployment process, or stop affected traffic until enforcement is restored.
Exit: /app/genai/chat returns a model response for an authorized pair and
persists audit evidence identifying the user, agent, model, and policy revision.
Verification matrix
| Test | Expected result |
|---|---|
| Valid user role, valid agent, matching assignment | Inference succeeds; both identities audited |
| Denied user role, valid assigned agent | 403; provider not called |
| Valid user, expired/invalid service token | 401; provider not called |
| Expired user, valid assigned agent | 401, including later tool-loop calls |
| Cross-host or wrong-environment agent | Denied before dispatch |
| Another agent’s internal alias or removed binding | Same AliasNotFound as an unknown alias |
| Internal alias with omitted scope token on direct-user profile | Same AliasNotFound as unknown alias |
| Route profile requires both credentials, workload absent | 401 independent of model existence |
| Invalid supplied scope token on an unbound model | 401 |
| Intentionally unbound model with valid direct-user request | Existing policy preserved |
| Unknown or ambiguous workload mapping | Denied before dispatch |
| Assignment projection missing or invalid | 503; no unrestricted fallback |
| Binding revoked between two model requests | Next admission denies; revisions audited |
User JWT has UI client_id, agent has bound principal | Binding, limiter, and billing keys remain the agent keys |
User JWT lacks routeAlias, workload restricts alias | Workload restriction remains enforced |
| Conflicting signed alias restrictions | Deny before dispatch |
| Duplicate or combined credential headers | Reject before single-value access |
| Bearer scope header on new profile; raw scope on legacy profile | Correct independent parsing and compatibility |
Tool/Knowledge lad1 presented for LLM | Reject before replay consumption or dispatch |
| Static registry token lacks gateway audience | Profile cannot activate using that credential |
| User token lacks the gateway audience required by its issuer/profile contract | Profile cannot activate with that token contract; runtime rejects the token with 401 before dispatch, without disabling audience validation |
| Workload acquisition, expiry, refresh failure, and recovery | No stale-token loop; valid renewal resumes service |
| User expires on open socket, renews and reconnects | Same-owner session resumes without duplicate turns |
| Concurrent users sharing HTTP connection pool | Each request retains its own user token |
| Retry, tool continuation, and streaming admission | Same authorization contract enforced |
| Provider fallback or redirect | No policy escape or inbound credential disclosure |
| Audit sink outage and recovery | Configured delivery policy honored; retained events recover |
Use focused unit tests, a mock inference provider, and a real PostgreSQL audit integration gate. Database tests skipped for lack of configuration do not count as qualification. Complete the live UI-to-agent-to-gateway test with both positive and negative assignment cases and persisted audit inspection.
LLM Dual-Token Phase 1 Contract
Phase 1 introduced the prepared configuration and publication contract from
LLM User and Agent Authorization. Its original
activation guard was removed by Phase 3, after gateway
enforcement and Agent forwarding were implemented. Empty {} publications retain
their existing behavior and canonical digests. The issuer and deployment gates
below still apply before activating a nonempty policy.
Implemented contract
agent-runtime-protocol::gateway_delegation::GatewayDelegationPolicy replaces
the untyped agentPolicy.gatewayDelegation value. The shape is either exactly
{} or the following prepared contract; unknown fields, null dualToken, and
incomplete contracts are rejected. The values below are test fixtures, not
qualified deployment credentials or issuer configuration.
{
"dualToken": {
"schemaVersion": 1,
"profile": "user-agent-dual-token-v1",
"gatewayUrl": "https://llm-gateway:8443/v1",
"userIssuer": "https://oauth.example.test",
"userAudience": "urn:com.networknt",
"workloadIssuer": "https://oauth.example.test",
"workloadAudience": "urn:com.networknt",
"tokenEndpoint": "https://oauth.example.test/oauth2/provider/token",
"clientId": "019d8349-41c6-72ef-95c2-4428a40d0e49",
"clientSecretFile": "/run/secrets/agent-llm-client-secret",
"scopes": ["portal.r"],
"refreshBeforeSeconds": 60,
"routeAlias": "assistant-dev"
}
}
Both URLs require HTTPS and exclude credentials, query, and fragment.
gatewayUrl and routeAlias must equal the enclosing immutable model policy.
Both issuer/audience pairs are explicit nonempty exact strings. clientId is a
nonnil UUID. clientSecretFile is a normalized absolute path beneath
/run/secrets/; it names deployment material, never a secret in Config Server.
Scopes are distinct OAuth scope tokens. The refresh margin is 1–599 seconds,
strictly below the current issuer’s 600-second client-credentials lifetime.
These checks validate publication structure, not the truth of issuer claims or
live credentials. Phase 2 must validate actual tokens against these expectations.
The contract deliberately has no ignoreExpiry, audience bypass, parser fallback,
identity override, or enabled switch.
Publication and digest compatibility
The Java publisher in light-portal/db-provider reads the optional instance map
property agent-policy-authoring.gatewayDelegation, assigned to the agent’s
product version through the existing configuration authoring model. Register
that optional map property and assignment using the existing Portal configuration
commands before authoring a prepared candidate; no new SQL assignment table is
needed. Absent authoring produces {}.
AgentPolicyPublicationPersistence.loadGatewayDelegation tracks the property,
assignment, instance-value versions, and authored value in
sourceAggregateVersions.gatewayDelegationAuthoring. This source is separate
from the generated runtime property, avoiding a publication invalidating its
own inputs. Ambiguous property matches fail. Publication also requires the configured
clientId to have an active auth_client_t row in the same host with
api_version_id equal to the Agent definition UUID. Registration version and
identity enter sourceAggregateVersions.gatewayClientRegistration; missing or
mismatched registration denies candidate compilation. AgentGatewayDelegation validates
the authored contract before compilation.
AgentPolicyProjectionCompiler emits the complete value as one map property,
agentPolicy.gatewayDelegation, consumed by the existing agent.yml placeholder
agent.agentPolicy.gatewayDelegation. It does not flatten child properties.
For nonempty contracts, gateway delegation enters the product-profile digest
material and consequently the policy digest, as well as the complete content
digest. Empty legacy policy material remains byte-compatible. The existing
publication source/content digests detect changes without a parallel version or
assignment store.
Java exports a complete prepared Agent policy fixture, and the Rust consumer
recomputes its canonical content digest. Both runtimes preserve the contract
without adding default fields to {}. This is a schema extension requiring the
new consumer; it is not safe to publish to older consumers that ignore the slot.
Existing alias and identity contract
LlmModelPersistenceImpl.appendV4Routes already projects:
| Portal alias visibility | Gateway alias fields |
|---|---|
PUBLIC | No internal binding |
INTERNAL_AGENT | internal: true, boundPrincipal: <bound_agent_def_id UUID> |
INTERNAL_WORKLOAD | internal: true, boundPrincipal: <bound_workload_principal> |
instancePropertyCandidateV4 preserves those fields in the aliases map. The
Rust AliasConfig, compiler, and runtime already consume them. Phase 1 tests the
real Java route publisher with controlled JDBC rows and feeds its exported
agent-bound alias to the Rust config deserializer. Existing internal-alias
admission and model-listing tests cover rejection/concealment behavior.
For the interactive agent profile, the execution principal must be the existing
agent-definition UUID, not the user JWT’s UI client ID and not necessarily the
service-token client ID. Resolve the latter through the existing host-scoped
auth_client_t.api_version_id registration and active Agent definition. Do not accept an authored or caller-supplied UUID
as identity proof. That verification and adapter are Phase 2 work.
The binding, limiter, billing, and routeAlias rules in the parent design remain
normative. The prepared routeAlias is an additional immutable restriction and
cannot replace or widen signed token restrictions. This phase adds no new alias
or permission store and changes no live execution principal.
Concrete issuer acquisition and renewal contract
The checked portal-service/apps/light-oauth/src/main.rs implements
POST /oauth2/{providerId}/token with grant_type=client_credentials and supports
client_secret_basic and client_secret_post. The selected client contract uses
HTTP Basic client authentication with clientId and the secret read from
clientSecretFile, and a form body containing grant_type=client_credentials
and the space-joined requested scopes. Use the configured HTTPS endpoint and
trusted CA; disallow redirects. Never request the long_lived grant.
The current normal access token lasts 600 seconds and has no refresh token.
Renewal means repeating the client-credentials grant, not using refresh_token.
The Phase 3 provider must read the current mounted secret on acquisition, refresh
before the configured margin, single-flight concurrent refreshes, validate the
returned token before cache replacement, and retain the previous verified token
only until its expiry. Bound backoff to the remaining valid lifetime; after
expiry fail new inference explicitly, then recover when acquisition succeeds.
A static registry_token is not a renewable inference credential.
Issuer limitations are concrete deployment gates:
generate_jwt_with_keyuses the server’s configuredjwt_audience; the client-credentials form does not choose a per-request audience. Qualify that issuer-wide audience as an accepted gateway recipient for both user and workload profiles.urn:com.networkntis not proof of that deployment decision. If a distinct gateway-only audience is required, issuer changes or a separate qualified issuer configuration are required before activation.- At the Phase 1 baseline,
handle_client_credentialspassed empty extra claims. It did not emit agent definition, host, environment, orrouteAliasclaims from custom registration claims on this grant. The target workload contract requiring host/environment therefore needed an issuer change that derives them from trusted registration; Phase 4 adds that registered workload projection. Do not let client-supplied form claims assert them. Existing signed alias bounds must likewise survive migration, or have a separately qualified equivalent. subandclient_idrepresent the OAuth client. They are not the Portal agent definition UUID. Verify the canonical mapping before using the alias binding.
Phase 1 specifies these limits; it does not claim the current issuer can already issue a deployable token satisfying every gate. Issuer changes and live token qualification must precede Phase 4 activation, without relaxing the contract.
Concrete user reauthentication contract
portal-view/src/pages/genai/Chat.tsx opens the WebSocket using the browser’s
access-token cookie and CSRF subprotocol. light-agent authenticates at upgrade
and captures the principal/token in handle_socket; subsequent turns do not
refresh them. UserContext.tsx has an HTTP-based cookie-renewal path, but no chat
reauthentication protocol exists yet.
Phase 3 must check captured-token expiry before each model call and stop new
admission when it expires. Use a structured authentication_required chat event
and close code 4401 as the new protocol contract. The UI renews authentication
through the existing HTTP login/refresh-cookie flow, then reconnects with the
same session ID and fresh cookies/CSRF context. If renewal fails, require sign-in.
Never carry an old captured bearer into the reconnected socket.
The agent revalidates host, user, and agent ownership on reconnect. Preserve accepted turn IDs and query/reconcile their durable status; do not automatically resubmit an accepted or ambiguously completed turn. Unaccepted drafts may remain in the UI for explicit submission. A tool-loop call encountering expiry ends that turn with an explicit authentication outcome; reconnect does not replay its side effects. An already admitted stream may finish under its admission context, but a later model call requires valid credentials.
These requirements are implemented and covered by the Phase 3 qualification gate. Full deployed-path qualification remains Phase 4.
Verification
Run scripts/run-agent-llm-phase1-gates.sh from light-fabric. It runs the Java
contract/publication tests, compares regenerated Java fixtures with checked-in
Rust fixtures, tests Rust contract and agent validation, tests alias parsing and
existing binding/concealment behavior, and builds the book.
JDBC publication tests use controlled mocks; they prove compiler/source tracking and serialization contracts, not a live database migration or deployment. Phase 2 still requires real audit persistence, and Phase 4 requires live issuer, database, and UI qualification. Phase 1 changes neither live configuration nor credentials.
Agent LLM dual-token Phase 2 implementation
Phase 2 implements gateway admission and durable identity audit. Phase 3 now supplies Agent forwarding, credential renewal, and interactive reauthentication. Deployment activation remains Phase 4.
Admission and identity
Trusted llm-router.agentDelegation.endpoints selects the dual-token parser for
POST chat completions, Responses, and Anthropic messages. A true value requires
both tokens; false permits user-only calls and validates any supplied workload
token. Model listing and routes outside this map retain their existing profile.
An overlapping HMAC profile fails closed. There is no parser retry or legacy
lad1 fallback on these routes.
Authorization: Bearer <user JWT> and X-Scope-Token: Bearer <workload JWT>
are parsed from all header values; duplicates and malformed forms are rejected.
Both signatures use the shared JWT verifier. Issuer, audience, expiry, and
not-before checks are enforced independently of legacy expiry bypass settings.
The workload must match a current projected registration, host, environment,
scopes, and route alias. Expired publication evidence fails closed.
User access-control rules receive the user principal. Model binding and limiter
identity receive the registered agent definition ID; billing uses the verified
workload billing subject with the agent ID as fallback. Signed route-alias
restrictions intersect the published alias. Existing internal aliases and
boundPrincipal remain the model assignment authority. Provider requests strip
both incoming credential headers and use provider credentials.
Publication and activation prerequisites
The Portal candidate compiler materializes llm-router.agentDelegation from the
existing published agentPolicy.gatewayDelegation, current Agent snapshots,
active agent definitions, and active OAuth registrations for the same host and
environment. It carries registration version, policy digest, and publication
expiry. This is a derived projection, not another assignment store; no client
secret or secret-file reference is exported.
Register the optional agentDelegation property in Portal’s llm-router config
metadata before publishing the gateway candidate. Older metadata remains
compatible. Initial generated inference endpoints permit direct user calls;
subsequent publication preserves trusted endpoint requirements. Removing the
last binding retains the profile with an empty binding list instead of restoring
legacy parsing. Registration revocation takes effect after gateway publication
and reload, bounded by the projected publication expiry.
Before activation, apply audit PostgreSQL migrations through
0006_authorization_context.sql, configure the audit database environment
reference and writable WAL, and require local_durable audit for generation
aliases. Binding host IDs must match the audit host. Gateway loading rejects
missing aliases and conflicting internal alias bindings. Prepare issuers that
actually issue the specified audiences; audience checking must remain enabled.
Requests pin both the verified identity and the matching routing snapshot across reload. Runtime snapshots share concurrency permits so reload cannot reset limits.
Audit and compatibility
Admission denials and authorized inference records persist structured
authorizationContext, including verified user/workload identities, resolved
agent and registration evidence, correlation ID, and access/assignment decisions.
Unverified claim values and bearer credentials are not recorded. Audit admission
failure returns 503 before provider dispatch. Missing or inaccessible internal
models retain the existing indistinguishable 404 response.
The JSON field is additive and absent for legacy calls. PostgreSQL migrations are idempotent. Deploy the updated audit consumer before activating the profile; do not replay a WAL containing dual-token records through an older consumer, which does not preserve the added identity field. Drain the WAL with the updated consumer before rolling back. Existing legacy admission profiles remain unchanged.
Verification
Run scripts/run-agent-llm-phase2-gates.sh with
LLM_AUDIT_TEST_DATABASE_URL pointing to a disposable, dedicated PostgreSQL
database. The gate refuses to silently skip database qualification.
The gate checks Java publication contracts against the Rust fixture, applies migrations twice, runs gateway unit/data-plane/alias regressions, explicitly runs the PostgreSQL idempotency test, and starts an actual Pingora gateway with mock JWKS and provider servers. The live test verifies successful dispatch, user RBAC, audience/expiry checks, host/scope/alias/registration failures, missing or duplicate credentials, legacy token rejection, provider credential isolation, and persisted user/workload/agent attribution. Java publication tests use mocked JDBC; the audit sink test uses real PostgreSQL. This is gateway qualification, not a live Portal chat or credential-renewal qualification.
Agent LLM dual-token Phase 3 implementation
Phase 3 implements request-scoped Agent forwarding, renewable workload credentials,
and interactive reauthentication. A validated nonempty
agentPolicy.gatewayDelegation.dualToken now selects the implementation; {}
retains the legacy single-token path. Deployment and live issuer qualification
remain Phase 4. No live configuration has been activated by this implementation.
Request-scoped forwarding
CompatibleProvider accepts a typed GatewayAuthorization for one turn. The
Agent constructs it only from the authenticated invocation and its published
workload credential source. Each model call, including tool-loop continuations,
requests the credentials again and checks both expiry times immediately before
sending. It sends the original user JWT in Authorization and the acquired
workload JWT in X-Scope-Token, both with the Bearer scheme.
The shared HTTP client never holds user credentials in default headers. Concurrent turns have separate user authority; only the workload cache is shared. The provider pins its configured destination and rejects a changed base URL. Agent outbound clients disable redirects, and activating this profile requires TLS hostname verification and real JWT verification. Incoming duplicate or malformed user credential headers are rejected. UI-supplied scope credentials are ignored.
Credentials are not serialized into conversation history, turn records, snapshots, or audit. Credential-bearing request headers are marked sensitive, authenticated request objects have no derived Debug output, and dual-token gateway errors do not include response bodies that could echo credentials. The registry credential remains separate. Native A2A inference explicitly rejects this profile because it has no original interactive user token; there is no service-token fallback.
The current compatible client uses buffered chat completions. This phase preserves that transport and its tool-loop behavior; it does not add an Agent streaming API. No automatic retry is added after a potentially billable request.
Workload issuance and renewal
gateway_credentials.rs implements the Phase 1 client-credentials contract:
HTTP Basic client authentication, form grant_type=client_credentials, and the
published space-separated scopes. It uses the configured endpoint, trusted CA,
a 15-second timeout, no redirects, and a bounded response body. Every acquisition
reopens clientSecretFile, allowing atomic mounted-secret rotation.
The returned JWT passes the shared signature verifier and independent checks for issuer, audience, expiry, not-before/issued-at, registered client ID, host, environment, required scopes, and any signed route-alias restriction. Issued lifetime must exceed the refresh margin. Legacy expiry bypass cannot extend it.
A mutex makes refresh single-flight. Calls reuse the verified cache until the
refresh margin. A failed refresh may use the previous verified token only until
its original expiry; failures have a two-second retry backoff. Initial acquisition
failure or expiry returns workload_credential_unavailable without dispatch.
A later successful grant recovers the cache. An expired captured user token fails
before workload acquisition, and expiry is checked again after acquisition so
waiting for refresh cannot extend user authority.
Interactive expiry and reconnect
WebSocket upgrade verifies the current user JWT, its explicit gateway audience,
and host/user/Agent ownership. Under this profile the Agent advertises an
authentication_context event containing only expiresAt. It checks captured
expiry before durable turn admission and before every outbound model call.
Expiry emits authentication_required with the client message ID and whether
the turn was already admitted, followed by close code 4401. An admitted turn that
expires during execution terminates with that authentication outcome. Reconnection
does not replay earlier tool side effects or model requests.
The UI uses the existing same-origin HTTP refresh-cookie probe, then reconnects with fresh cookies and CSRF context under the same session ID. Renewal is bounded to 15 seconds and cancelled on disconnect, unmount, or identity/deployment selection change. Cookie owner/host changes block reconnect; the Agent independently checks durable ownership. Renewal failure requires sign-in. An expired unsent draft remains for explicit submission after reconnect.
Ordinary chat now supplies clientMessageId as coding already did. The Agent
acknowledges admission with turnAccepted. The UI retains submitted message IDs
and accepted turn IDs, without credentials, in the existing per-user/deployment
session-storage namespace. Reconnect reports the latest 100 durable turn states
through turn_status; older or unavailable status stays explicitly unresolved.
Accepted or ambiguous turns are never automatically resubmitted. An explicitly
unadmitted message retains its ID when restored as a draft. Existing database
idempotency and session ownership remain authoritative.
Verification and remaining deployment gates
Run scripts/run-agent-llm-phase3-gates.sh with
LIGHT_AGENT_TEST_DATABASE_URL pointing to a disposable PostgreSQL database with
pgvector, operational metadata migrations, and Agent store migrations applied.
The script requires the database and selects agent_ops; it does not silently
skip the database tests. PORTAL_VIEW_SOURCE_DIR can override the sibling UI path.
The passing gate covers:
- Signed JWT issuance through a mock HTTP issuer/JWKS server, concurrent refresh, secret rotation, actual expiry, failed-refresh fallback, expiry denial, recovery, and isolation of users sharing the workload cache.
- Real HTTP model requests with both headers, credential rechecks, registry-token exclusion, destination changes, redirect rejection, and legacy compatible-client regressions.
- Agent configuration, ownership and authentication-event tests; real PostgreSQL same-owner resume, different-owner rejection, duplicate admission, FIFO ordering, and unavailable-execution reconciliation.
- UI reconnect, retained drafts and accepted IDs, durable status reconciliation, account-change rejection, cancellation on Host change, and existing chat tests.
- Existing gateway data-plane regressions, UI lint, documentation build, and diff checks.
The issuer fixtures use loopback HTTP only in tests; production publication still requires HTTPS. These checks are not a live browser-to-Portal-to-Agent-to-gateway qualification. Phase 4 must qualify the deployed issuer, CA, audience acceptance, registration claims, mounted secret rotation, gateway audit sink, and reconnect behavior together. The issuer limitations recorded in Phase 1 remain hard gates: missing workload host/environment or an unsuitable user audience must be fixed at issuance, without disabling validation.
Agent LLM dual-token Phase 4 rollout and qualification
Policy lifetime update: Agent policies and derived gateway assignments now remain effective until replaced or revoked through an applied update. Mandatory policy lease renewal is removed; authentication tokens still expire normally. See Policy lifetime for offline files, cached startup, compatibility, and deployment requirements. This source change is not proof that an existing container image includes it.
Status: implementation and qualification tooling prepared. The default deployment
target is portal-config-loc/all-in-lt; the live /app/genai/chat exit check
remains pending. A passing
implementation gate or direct gateway probe is not Phase 4 completion.
Local deployment target
Use /home/steve/workspace/portal-config-loc/all-in-lt by default. Preserve the
running Compose project’s personal-runner and credentials overlays and its release
and private environment files when updating individual services.
The UI entry point is https://localhost:3000; use
https://localhost:3000/app/genai/chat for browser qualification.
Preflight on 2026-09-07 found the three deployed Agent definitions using the
public assistant-dev alias, with no published dual-token policy. Qualification
requires an internal alias bound to the selected Agent through normal publication.
The mounted OAuth and LLM gateway certificates cover localhost and selected IP
addresses, but not their internal service DNS names; use validated endpoint names
or issue suitable certificates before enabling hostname verification. The earlier
browser certificate failure at https://local.lightapi.net was at the wrong UI
entry point and does not establish a blocker for https://localhost:3000.
Login at the correct URL succeeded. The baseline Tech Support chat reproduced
the reported gateway 403 before policy activation. This is baseline evidence,
not a passing dual-token qualification.
The local llm_audit database has now received migration
0006_authorization_context.sql in one transaction; the JSONB column and its
constraint were verified afterward. This additive migration does not enable the
profile or constitute persisted dual-token audit evidence.
Local deployment now uses portal-config-loc/all-in-lt/docker-compose.yml;
the temporary dual-token overlay has been retired. Public certificates reside in
each service’s config/cert.pem. The separate renewable OAuth workload credential
remains in the external Tech Support credential volume.
Event artifacts and historical qualification notes are retained in
light-portal-event/genai/20260908-agent-llm-dual-token/. The historical notes are
not current deployment instructions. No bearer credentials belong in that record.
Remaining implementation
The normal light-oauth client-credentials grant now derives host from the
authenticated OAuth registration and reads only env/environment, routeAlias,
and billingSubject from its registered custom_claim JSON. Conflicting host or
environment values and invalid workload field types fail closed. The grant does
not accept form fields as workload identity, and does not project custom user IDs,
roles, audiences, subjects, expiry, or scope overrides. Existing scope validation,
issuer-wide audience, RSA signature, and 600-second lifetime remain authoritative.
Registrations without custom claims now include their authoritative host; no additional environment or alias is invented. Existing deployments must author those claims through the normal OAuth-client administration path before enabling the Agent profile. Malformed registered custom JSON, previously ignored by this grant, now causes issuance to fail rather than producing a partial identity.
The dual-token generation profile now returns model_not_found (404) for a request
outside its assigned route alias, matching the unknown-model response. The legacy
single-token signed-route restriction retains its existing 403 response.
The Portal gateway candidate compiler now promotes every generation alias to
local_durable when it emits a non-null delegation profile, including a retained
profile with no remaining bindings. This matches the gateway’s startup invariant.
Embedding-only aliases and publications without the profile retain their existing
audit mode. The source route resources are not mutated. The publisher regression
suite passed 43 tests with no failures or skips on 2026-09-08.
Live publication exposed a second omission: the event consumer’s managed-property
allowlist rejected agentDelegation. The projection now accepts that property
while retaining its metadata checks. Fresh-publication and replay tests include
the ninth property and its ownership write; all 43 targeted tests pass. The local
Portal images tagged 2.3.5-local.dualtoken.20260908.3 contain this correction.
The original failed event must be recovered through the independently approved
replay workflow; deployment of the fix alone does not resolve its DLQ entry.
Ordered rollout
- Record exact image digests and current published snapshot IDs for the issuer, gateway, Agent, UI, and Portal publisher. Retain the previous compatible policy revisions for rollback. Use the selected environment’s normal deployment and publication process; do not activate by editing generated snapshot values.
- Deploy the compatible issuer and Portal publisher. Register the optional
agent-policy-authoring.gatewayDelegationauthoring map and gatewayllm-router.agentDelegationmetadata if absent. Assign the OAuth client to the Agent definition in the same host, with an active provider-client registration. Author the environment and any signed alias/billing claims on that registration. - Apply audit migrations through
0006_authorization_context.sql. Provision a persistent writable WAL and PostgreSQL sink. Deploy the compatible gateway before Agent forwarding. Keep generation aliases onlocal_durableaudit. - Obtain normal client-credentials tokens through the configured TLS endpoint and mounted secret. Verify the declared issuer/audience for both user and workload tokens. The issuer’s audience remains issuer-wide; the request form cannot select it. Missing or unsuitable claims block activation. Never enable expiry bypass or disable CA/hostname checking to make the test succeed.
- Deploy the compatible Agent and UI. Mount the acquisition secret at the
published
/run/secrets/...reference with restrictive permissions. Publish the immutable dual-token Agent contract and its derived gateway projection through the existing candidate/validation/apply workflow. The projection must resolve the client to the Agent definition and agree with the existing internal aliasboundPrincipal. Configure the selected inference route to require both credentials for the restricted-path qualification. - Run gateway admission probes, followed by the browser checks below. Capture audit and publication identifiers; do not capture bearer values, cookies, client secrets, or a raw authenticated network trace in the report.
Repeatable implementation gate
Set LIGHT_OAUTH_TEST_DATABASE_URL to a disposable PostgreSQL database and run:
./scripts/run-agent-llm-phase4-gates.sh implementation
The issuer integration test uses connection-local temporary tables, signs a real
RSA JWT through the production grant handler, verifies its signature and claims,
and proves that injected form host/environment/role/alias values do not override
registered authority. It explicitly runs with --ignored; absence of the database
is an error. The gate also runs issuer unit tests, gateway data-plane regressions,
qualification-runner tests, and documentation/diff checks. These do not deploy
services and produce only IMPLEMENTATION_CHECKS_PASSED.
Live gateway probe
scripts/agent-llm/qualify_gateway.py requires a trusted CA and an explicit HTTPS
chat-completions endpoint. It refuses redirects and reads tokens only from
owner-only files. Configure libpq PGHOST, PGDATABASE, PGUSER, and
PGPASSFILE for read access to the gateway audit database. Credentials must not
appear in command arguments or checked-in config.
Supply a public expectations JSON object with:
endpoint: the deployed HTTPS/v1/chat/completionsURL.alias: the internal alias assigned to this Agent.unassignedAlias: an existing alias assigned to another Agent.userId,workloadClientId,agentDefId,hostId: expected verified UUIDs.agentPolicyDigest: the active Agent content/policy evidence digest from the gateway’s derived binding projection, not an assumed local source revision.requireWorkload:true, matching trusted route configuration.
Run directly with --config, --user-token-file, --workload-token-file,
--ca-file, and --report, or use run-agent-llm-phase4-gates.sh live-gateway
with the corresponding AGENT_LLM_* variables documented in that script.
The probe performs one small inference and four denials: missing workload,
invalid workload, another Agent’s alias, and an unknown alias. It requires
persisted terminal audit records for all five, checks successful user/Agent/model/
policy attribution, and rejects any denied case with a provider-attempt record.
Reports contain request/correlation identifiers and status only. Starting a run
invalidates an older pass at the report path. The result is explicitly
GATEWAY_ADMISSION_PASSED with browserQualification: NOT_RUN.
Browser exit evidence
Session admission must permit a forward publication transition for the same
active runtime scope. The Agent uses existing accepted AGENT_POLICY reference
evidence to establish the previous policy version and requires a strictly greater
incoming version before replacing the scope’s publication and content digest.
Host, instance, service, environment and audience remain bound; missing,
inconsistent or revoked baseline evidence fails closed. Scope changes, new policy
evidence and session creation commit in the same database transaction. Existing
sessions retain their original pinned authority and are not silently migrated.
The local OAuth leaf certificate must include light-oauth because the Agent
fetches JWT keys through that internal DNS name. A healthy Agent listener alone
does not prove that authenticated WebSocket admission works. The normal
all-in-lt/docker-compose.yml mounts the DNS-valid OAuth certificate as well as
the Config Server and LLM gateway certificates; no TLS bypass is required.
Use the deployed /app/genai/chat page under an authorized user account. Select
the intended deployed Agent, submit a small uniquely identifiable turn, and
confirm a model response. Correlate that time window and verified user/Agent with
the gateway’s PostgreSQL audit and the Agent’s durable turn record. Record the
actual user, client, Agent, alias, registration version, Agent policy digest,
gateway snapshot revision, and durable request/turn IDs.
Repeat with a denied user or an incompatible assignment; no provider attempt may occur. Restore the compatible assignment through normal publication and verify the next authorized admission. Exercise user expiry, HTTP cookie renewal and same-owner reconnect without resubmitting accepted turns. Exercise workload renewal across expiry, temporary acquisition failure, recovery and mounted-secret rotation. Inspect retained audit delivery after a sink outage. Use an isolated qualification deployment for destructive outage/rotation exercises.
Only mark Phase 4 complete after this deployed UI-to-Agent-to-gateway response and persisted audit evidence exist. Do not substitute mock-provider unit results, a direct curl response, decoded but unverified JWT claims, or a manually edited report for that exit evidence.
Rollback
Keep the compatible gateway enforcing the active profile while reverting an
Agent or UI change. Disable affected traffic before restoring a gateway version
that cannot enforce the active policy. Revert compatible publication and binary
versions together; never remove agentDelegation merely to bypass a denial.
Drain the audit WAL using the new consumer before any rollback to an older audit
consumer that would discard dual-token identity fields. Preserve database audit
history and additive migrations. Requalify both allowed and denied admissions
before restoring traffic.
Coding Harness Integration
Status: proposed target design. The implementation inventory in this document was verified against the repository on September 4, 2026. Vendor authentication, subscription, and protocol rules must be rechecked before each qualified release.
This design specializes the coding-agent portion of
Light-Agent Execution. It defines how
light-agent can use Codex and, when justified, Claude Code without making an
external coding harness the enterprise policy authority.
The durable requirement, design, implementation-plan, review, multi-repository, and GitHub issue lifecycle built on these workers is defined in Personal Development Workflow Orchestration and Enterprise Development Workflow Orchestration.
For the opt-in locally qualified Claude subscription adapter, integration alternatives,
and qualification plan, see Claude Personal Worker.
Phase 2 supplies light-claude-worker and Agent/runner integration; production
distribution and deployment qualification remain separate gates.
Decision
For the proposed persistent multi-repository input mode, see Shared Task Workspaces. It defines runner-managed worktrees and shared implementation/review access without changing the existing bundle contract or making the adapter the workspace authority.
Use light-agent as the durable enterprise agent authority and run each
workspace-aware coding loop through light-agent-worker in a runner-managed
sandbox.
The first Codex integration candidate is a pinned Codex App Server process over local standard input/output. It provides a typed process and failure boundary that is natural for the Rust worker to drive directly; no Python adapter is required. Because the App Server protocol is currently experimental, this is a qualification decision rather than a claim of production support.
Maintain one trusted worker core and allow separately built adapter variants:
codex-app-server-v1: starts a pinned Codex App Server and translates its JSON-RPC messages into the Light agent runtime protocol;codex-embedded-v1: optionally links pinned Codex Rust crates directly after their library boundary, licensing, upgrade cost, and failure isolation have passed qualification;claude-code-v1: an optional later adapter for workloads that require Claude Code behavior or independent harness diversity;- other versioned native-harness adapters, such as a future Grok-oriented coding worker, when their protocol, isolation, authentication, licensing, and release compatibility pass the same qualification contract.
Remove Pi from the supported target architecture. Preserve its existing scheduling, sandbox, immutable-input, and canonical-patch tests only as a migration baseline until Codex passes the replacement gate; then delete the Pi runtime, profile, image, template, and Node/npm dependency rather than carrying Pi as an optional adapter.
Do not call the Rust-linked option the “Codex SDK” in contracts. The public Codex SDKs are currently documented for TypeScript and Python. Direct Rust linkage is an embedded integration against pinned crates and may depend on APIs that are not maintained as a stable external SDK.
Use logical model aliases such as coding-implementer and coding-reviewer.
The same Codex worker variant may use different gateway-backed models for
implementation and review. A Claude worker is needed only for Claude Code
harness semantics or deliberate harness diversity, not merely to use an
Anthropic model as a reviewer.
Terminology is deliberately distinct: an adapter ID selects a harness
integration, a role execution profile selects workspace/tool authority, an
authentication profile selects personal-subscription or enterprise-api,
and a logical model alias is resolved independently by the configured model
route. None of these identifiers may be reused as another kind.
Goals
- Reuse capable coding harnesses without duplicating their model/tool loop.
- Keep session, policy, approval, quota, audit, and artifact authority in Light.
- Support provider and model diversity through
llm-gatewaywhere the selected harness protocol is compatible. - Isolate repository mutation, shell execution, local MCP, and coding-process credentials in a bounded runner sandbox, subject to the explicitly weaker credential-isolation claim in Local Native Isolation Profile.
- Support a fresh, read-only review turn after an implementation turn.
- Permit personal subscription use only in a dedicated user context that the vendor permits, while keeping pooled enterprise execution API-backed.
Non-Goals
- Treating Codex, Claude Code, or another harness as an ordinary model provider.
- Launching a coding CLI from the long-lived
light-agentservice. - Sending personal subscription credentials through
llm-gateway. - Letting a prompt select a binary, adapter, provider, credentials, or approval mode.
- Using a reviewer model in the same thread as proof of independent review.
- Replacing durable workflow branching, retries, timers, or approvals with a coding harness.
Authority And Trust Boundaries
| Component | Authority |
|---|---|
light-agent | Agent/session policy, turn admission, model alias, budgets, approvals, durable result, and review requirements |
light-workflow | Durable business process, branching, retry, wait, and workflow approval |
controller-rs and runner | Placement, reservation, lease, sandbox lifecycle, resource enforcement, and cleanup |
light-agent-worker core | Lease validation, materialization, adapter launch, event normalization, cancellation, and artifact proposal |
| Runtime adapter | One bounded model/tool loop; it may narrow but never widen the lease |
llm-gateway | Workload authentication, logical-to-physical model routing, provider credentials, rate/cost policy, and protocol mediation |
| Skills | Approved instructions and supporting resources; never execution authority |
| Fixed action | Push, pull-request creation, signing, publishing, and deployment over an accepted immutable artifact |
light-gateway and llm-gateway are distinct logical roles. light-gateway
is the Portal/API ingress and policy-enforcement edge; llm-gateway is the
model-routing, provider-credential, and usage-accounting subsystem. A deployment
may package both roles in one process, but their authorities are not
interchangeable.
The effective coding authority is an intersection:
caller grants
intersect immutable agent and execution profile
intersect runner lease and sandbox policy
intersect approved skill and tool manifests
intersect adapter compatibility record
intersect current approval, quota, and revocation state
An adapter capability announcement can remove a capability. It cannot add one that is absent from this intersection.
For enterprise calls, workload authentication also carries a verified end-user identity and a distinct actor identity. The user is the attribution and default billing subject; the workflow, agent, or worker is the workload actor. The gateway validates both and never treats a pooled service credential as the user.
Architecture
flowchart TB
U[Portal, API, or workflow] --> A[light-agent<br/>session and policy authority]
A --> C[controller-rs and runner<br/>lease and sandbox]
C --> W[light-agent-worker<br/>trusted common core]
W --> AS[Codex App Server adapter]
W -. optional .-> ER[Embedded Codex adapter]
W -. optional .-> CC[Claude Code adapter]
W -. future qualified adapters .-> OA[Other coding harness]
AS --> B
ER --> B
CC --> B
OA --> B
B --> G[llm-gateway<br/>logical model alias]
G --> O[OpenAI API]
G --> AN[Anthropic API or Bedrock]
G --> X[Other qualified provider]
AS -. dedicated personal profile .-> CS[Codex subscription login]
CC -. dedicated personal profile .-> CLS[Claude subscription login]
W --> P[Canonical patch and evidence]
P --> F[Fixed push, PR, publish, or deploy action]
The subscription edges and the gateway edge are mutually exclusive for a
given model call. A subscription credential is consumed only by its native
vendor harness in a dedicated user context. llm-gateway accepts workload/API
credentials and never acts as a subscription proxy.
Worker Packaging
Multiple adapter variants are reasonable, but they should not become separate independent security implementations.
flowchart LR
CORE[Trusted worker core<br/>lease, policy, journal, artifacts] --> APPA[App Server adapter]
CORE --> EMBA[Embedded adapter]
CORE --> CLA[Claude Code adapter]
CORE --> OTHER[Other qualified adapter]
APPA --> I1[App Server worker image]
EMBA --> I2[Embedded worker image]
CLA --> I3[Claude worker image]
OTHER --> I4[Other worker image]
The common core owns:
agent-runtime-protocolframing, identity, fencing, and sequence validation;- immutable execution-spec and capability-digest verification;
- approved context and skill materialization;
- writable-root, protected-path, network, resource, and deadline enforcement;
- broker attachment without exposing reusable provider credentials;
- process-tree cancellation and bounded output;
- canonical diff calculation, artifact limits, and cleanup evidence.
Each adapter owns only launch, protocol translation, feature negotiation, and adapter-specific error normalization. Images pin the adapter implementation and version. Server-owned compatibility policy binds adapter ID, image digest, capability digest, and allowed execution profiles. A request or prompt cannot override that selection.
Local Native Isolation Profile
The enterprise API profile retains the full runner isolation boundary and does not expose a user’s native credential store. The personal-subscription profile is deliberately weaker because the vendor harness must discover credentials in its dedicated local user context.
The local runner still enforces the lease, admitted repository and writable roots, protected paths, tool and approval policy, network policy, resource and deadline limits, process-tree termination, bounded artifacts, and scratch cleanup. It also starts a fresh harness context for independent review. It may, however, use a same-user host process or container with narrowly mounted native configuration and credential-store paths instead of the enterprise MicroVM. Those paths are readable only by the harness process and are never materialized in the repository or general tool workspace.
This profile protects repository integrity and bounds process lifetime; it does not claim that untrusted repository code is isolated from credentials owned by the same operating-system user. Subscription-backed execution is therefore for a trusted single-user local environment. Repositories requiring hostile-code isolation must use the enterprise API profile or a separately qualified host-mediated credential design.
The host-process path is admitted only on an exclusive
maximumConcurrency: 1 local runner and advertises
local-single-user-native-v1, not restricted-model-egress. Enterprise pools
must configure a digest-pinned per-attempt sandbox launcher. Its reviewed
profile owns filesystem, process, and resource separation and the deployment
egress allowlist to llm-gateway; enterprise startup fails closed without it.
Retired pi-rpc-v1 Baseline
Phase 1 used the Phase 0 Pi contract as the migration baseline and then removed its scheduling path, adapter application and image, Cube admission, capability, enum value, policy profile, and Node/npm runtime. The retained deterministic coding fixture exercises immutable input and bounded canonical-patch behavior; it is test infrastructure, not a selectable product worker.
Retain these fixtures during Codex development because this is the only
concrete external RPC coding harness currently wired from light-agent through
Cube to a canonical patch. Treat its security checks and failure cases as the
replacement baseline. Do not spend work moving Pi behind the shared worker
core: after Codex passes the replacement gate, remove the Pi-specific policy,
scheduler path, adapter, image, template, capability advertisement, dependency,
and published profile.
codex-app-server-v1
Run one pinned App Server process inside the leased sandbox and communicate over local stdio using JSON-RPC 2.0 JSON Lines. Generate protocol schemas from the same pinned Codex version used in the image and compile or validate the Rust types in CI.
Prefer local stdio to a shared remote App Server:
- the OS process is a clear lifetime, cancellation, and resource boundary;
- credentials and repository access stay scoped to one sandbox;
- a server crash is isolated from
light-agentand other tenants; - version skew can be rejected using the image and schema digests;
- no independently exposed WebSocket control surface is required.
The adapter maps App Server thread, turn, item, approval, usage, and error notifications into ordered Light runtime events. Light remains authoritative: an App Server notification is evidence, not a durable Light state transition.
The App Server protocol is documented as experimental. Production enablement therefore requires an exact-version conformance suite and a fail-closed upgrade process. Do not advertise compatibility with an untested Codex release.
codex-embedded-v1
An embedded Rust variant can remove the child-process and JSON-RPC translation overhead. It does not need Python, but it creates a tighter coupling to Codex’s crate graph and runtime assumptions.
Qualify it independently for:
- stable or acceptably pinned Rust APIs and compatible licensing;
- Tokio/runtime, tracing, TLS, filesystem, and dependency compatibility;
- panic containment and cancellation behavior;
- equivalent approvals, tool events, usage, patches, and resumability;
- binary size, build time, security updates, and release cadence;
- absence of ambient credential or configuration discovery.
Do not fall back silently between embedded and App Server modes. They are distinct adapter IDs with distinct capability and image digests.
Public SDK Wrappers
The documented TypeScript and Python Codex SDKs are useful for applications in
those languages, but a Rust light-agent-worker does not need a Python bridge.
Driving App Server directly preserves the same explicit protocol boundary
without adding another language runtime. A wrapper should be introduced only
if it supplies a tested semantic layer that is expensive to reproduce and its
operational cost is justified.
Model Routing Through llm-gateway
In the enterprise API profile, the trusted worker launcher obtains an attempt-scoped credential through a command-backed helper before it starts Codex. This is Light launcher policy, not repository configuration. The Phase 2 worker obtains the token over its runner-owned Unix broker socket; it does not launch an external Python or shell adapter:
credential:
source: attempt-broker
target: llm-gateway-attempt
audience: llm-gateway
envelopeDirectory: /run/secrets/llm-gateway-attempts
exportForChildAs: LIGHT_LLM_ATTEMPT_TOKEN
The launcher then configures Codex with a custom Responses provider that exposes only a logical model alias:
model = "coding-implementer"
model_provider = "light_gateway"
[model_providers.light_gateway]
name = "Light LLM Gateway"
base_url = "https://llm-gateway.example/v1"
wire_api = "responses"
env_key = "LIGHT_LLM_ATTEMPT_TOKEN"
The envelope filename is the canonical attempt-binding SHA-256, so concurrent
turns never share a credential slot. If the pinned Codex version later supports
a directly configured credential helper, use that form and remove the launcher
export. The env_key form is a compatibility fallback: it names a short-lived
token present only in the Codex process environment, not a reusable gateway
bearer token in the worker’s ambient or tool-subprocess environment. Both the
private Codex home and command-line configuration pin the provider, gateway
URL, Responses wire protocol, environment key, and shell exclusion so a
repository-local .codex/config.toml cannot replace them. The worker must not
place a reusable gateway token or provider key in the repository, inherited
shell environment, process arguments, logs, runtime events, or artifacts.
The adapter must remove the attempt token from every shell, tool, MCP, and repository-command environment created by the harness. Qualification must prove that this filtering survives both normal tool execution and error logs.
The helper obtains a short-lived, llm-gateway-audience delegation token bound
to the verified end user, workload actor, host, workflow, agent session/turn,
logical model route, policy, and budget. It does not copy the original portal
JWT into the harness. llm-gateway alone holds provider API keys and returns a
trusted usage receipt so light-agent can reconcile its turn budget with the
gateway’s per-user token and cost ledger.
llm-gateway resolves coding-implementer or coding-reviewer to an eligible
physical provider/model deployment. This permits a Codex harness to use a
qualified OpenAI, Anthropic, Bedrock, xAI, or other backend without teaching
light-agent physical model names or credentials.
Protocol compatibility is a release gate, not an assumption. For every harness/provider route, test:
- buffered and streaming Responses ordering;
- function/tool call identifiers, arguments, and results;
- reasoning continuation and encrypted/opaque reasoning fields when used;
- usage and cost accounting;
- context limits, truncation, cancellation, timeouts, and rate limits;
- refusal, safety, retryable, terminal, and malformed upstream errors;
- provider fallback only before observable output, unless a documented continuation contract pins the deployment.
If an Anthropic-backed deployment cannot faithfully implement the Responses features required by the pinned Codex harness, that deployment is ineligible for the alias. Model availability alone is not compatibility.
Implementation And Independent Review
The workflow owns thread creation, resumption, and closure. See the implemented lifecycle contract. Implementer and reviewer threads remain separate, and both may persist through one stage’s remediation rounds. New stages allocate new threads.
A separate Claude worker is not required merely to review with a different model. The first design uses the same qualified Codex App Server worker variant with two immutable role execution profiles:
| Role execution profile | Logical model alias | Workspace | Purpose |
|---|---|---|---|
coding-implement-v1 | coding-implementer | Bounded write access | Diagnose, edit, and run allowed verification |
coding-review-v1 | coding-reviewer | Read-only repository plus writable ephemeral build scratch | Review the accepted patch and evidence |
The gateway may route the reviewer alias to a different model family, including an Anthropic model when the Responses compatibility gate passes. The exact physical name is gateway configuration, not an agent contract.
sequenceDiagram
participant A as light-agent
participant I as Implementer worker
participant R as Reviewer worker
participant F as Fixed action
A->>I: New stage thread: base, requirements, write lease
I-->>A: Canonical patch plus test evidence
A->>A: Validate protected paths and artifact digest
A->>R: New thread: immutable base, patch, requirements, evidence
R-->>A: Structured findings and verdict
alt Blocking findings
A->>I: Resume implementer thread with accepted findings
else Review accepted and approval satisfied
A->>F: Accepted immutable artifact
F-->>A: Push, PR, or publish result
end
The reviewer must receive:
- a new Light turn; create a separate reviewer harness thread at stage entry, then resume that reviewer thread for subsequent rounds in the same stage;
- a clean read-only reconstruction of the immutable base plus candidate patch;
- a writable ephemeral scratch/build directory, with language build outputs and caches redirected there and excluded from canonical patch calculation;
- requirements, relevant policies, and actual test evidence;
- no implementer chain-of-thought, hidden scratch state, writable repository tree, or repository mutation tools;
- a structured result with severity, file/location, evidence, remediation, and verdict.
The semantic reviewer may consume the implementer’s immutable test evidence and may reproduce an allowed build or test when its outputs fit in scratch. It is not the authoritative independent test executor. The fixed pre-publication gate in the development workflow re-executes required checks in a clean test/CI environment before publication.
Using a different model provides model diversity, not harness diversity. Add a
claude-code-v1 worker when Claude Code-specific repository behavior, tool
semantics, or a genuinely independent harness implementation is a requirement.
For high-risk changes, policy may require both a different model family and a
different worker image.
Review Assurance Policy
Model diversity, harness diversity, and human approval address different failure modes:
| Control | What changes | Primary purpose | What it does not prove |
|---|---|---|---|
| Model diversity | coding-reviewer resolves to a different model family in a fresh thread | Reduce correlated reasoning and interpretation errors | Independence from the coding harness, tools, or protocol adapter |
| Harness diversity | Review runs through a separately qualified adapter and worker image | Detect harness-specific prompting, tool, patch, sandbox, or protocol behavior | Authorization to accept business risk or perform an irreversible action |
| Human approval | An authorized person accepts a precisely digested artifact or action | Confirm intent, accountability, residual risk, timing, and business authority | Technical correctness without model review, tests, and fixed gates |
The default policy starts an independent reviewer thread at stage entry, resumes it across that stage’s review rounds, and may require model diversity without deploying another harness. Security/authentication changes, destructive data or schema migrations, signing/release controls, and broad multi-repository contract changes should require a different model family plus human approval. Harness diversity is an additional high-assurance control when the risk includes the primary harness itself or policy demands an independently implemented tool loop. Irreversible production, publication, signing, credential, or data effects always remain fixed actions behind human approval, regardless of model or harness diversity.
This is a policy matrix, not a claim that every environment must deploy every adapter. If no independent harness is qualified, a policy requiring harness diversity must pause rather than silently downgrade to model diversity.
Session Workspace Isolation
Use one runner-owned WorkspaceSet per mutable coding session. A workspace set
is a directory/volume namespace and immutable manifest bound to tenant, end
user, agent session, work package, repository base revisions, and lease. It may
contain many repositories, so a cross-repository feature still presents one
coherent workspace to the harness:
/workspaces/<workspace-set-id>/
manifest.json
repos/
light-fabric/
portal-view/
light-portal-doc/
scratch/
evidence/
Repositories are materialized lazily from the approved work-package manifest; a session does not need to copy dozens of unrelated repositories. Adding a repository requires a new admitted manifest version. The runner serializes mutating turns for the workspace set, and no second user or agent session may attach to it with write authority. A reviewer receives a different read-only reconstruction and scratch directory, never the implementer’s workspace set.
Git worktrees can reduce checkout time and disk use for one repository, but they are not the isolation boundary. Linked worktrees share repository-level state including most refs and, by default, configuration; Git also documents incomplete multiple-worktree support for submodules. A collection of per-repo worktrees under a session directory can be a local optimization, but only when the runner owns their creation, uses detached revisions or session-unique branches, enables worktree-specific configuration where needed, and brokers all shared-metadata operations.
For pooled or multi-user enterprise execution, prefer a separate clone/Git metadata directory for every repository in each workspace set, inside the session sandbox or volume. A read-only content-addressed object cache may be shared for efficiency; indexes, refs, configuration, hooks, credentials, working files, build outputs, and scratch must not be shared writable state. The sandbox remains the security boundary around the complete workspace set.
For a trusted single-user local profile, runner-managed per-repository worktrees are acceptable as a storage optimization, but separate workspace-set roots, exclusive leases, and process isolation still apply. This prevents two sessions from editing the same paths while retaining the weaker same-user security claim described earlier.
On completion or cancellation, the runner either destroys the workspace set or checkpoints its manifest, canonical patches, and evidence under retention policy. Resume creates or revalidates an exclusive lease and every base digest; it never reconnects a different user to an ambient existing directory.
Authentication And Subscription Profiles
The authentication profile is a separate immutable input to each role execution profile.
| Profile | Intended placement | Codex | Claude | Gateway |
|---|---|---|---|---|
personal-subscription | Local Portal via portal-config-loc/all-in-lt or light-portal-install | Native harness discovers its existing local Codex login; Light passes no vendor credential | Native harness discovers its existing local Claude Code login; Light passes no vendor credential | Not used for model calls |
enterprise-api | Pooled or dedicated enterprise runner | Attempt-scoped Light credential to custom Responses provider | Attempt-scoped Light credential to qualified Anthropic facade, if enabled | Required |
The local profile has two equivalent distributions:
portal-config-loc/all-in-lt for source-oriented platform development and
light-portal-install for packaged local use. They expose the same Portal,
workflow, agent, Worklist, Chat, approval, and optional CLI contracts.
Light components pass work, artifact, correlation, and platform-identity data, but do not pass an OAuth token, API key, or copied subscription credential to a native harness. The Codex or Claude Code process loads its own prior login from its normal local credential store. Light may query authentication status and ask the user to log in through the native client, but it does not manage that login. The local Portal JWT and service delegation tokens are a separate identity plane used for API authorization, human-task assignment, approval, and audit; they are not model-provider credentials.
Because local subscription calls bypass llm-gateway, Light has no normalized,
trusted per-user token or cost ledger for that profile. Local enforcement is
limited to round, turn/model-call, wall-clock, and process/resource limits;
provider-reported usage may be retained as advisory evidence but cannot drive a
cost-exhaustion transition. Enterprise API execution additionally enforces
gateway token and cost reservations from signed usage receipts.
Codex
Codex App Server supports API-key authentication and ChatGPT browser/device login. OpenAI also documents Codex access tokens for trusted unattended local automation in eligible Business and Enterprise workspaces.
Therefore the local Codex process may use the user’s eligible subscription identity directly, subject to the current OpenAI plan and controls. For this local-native profile, Light does not mint, copy, pass, or store a Codex vendor credential. The native process owns credential discovery and use.
Once Codex is configured to call llm-gateway as a custom provider, that path
uses Light workload credentials and API/cloud billing. A ChatGPT subscription
cannot be forwarded through the gateway to pay for arbitrary upstream model
calls.
Claude
Anthropic currently permits paid Claude users to authenticate the official Claude Code client, and documents long-lived Claude Code OAuth tokens for scripts, CI, and the Agent SDK. Anthropic separately states that Claude subscriptions and the Claude API are distinct products, and directs developers building third-party tools for others to the API or supported cloud providers. It also prohibits misrepresenting a client or routing third-party traffic against subscription limits.
Consequently:
- a Codex worker cannot reuse Claude subscription OAuth to call an Anthropic model;
- an Anthropic reviewer behind
llm-gatewayrequires Anthropic API, Bedrock, or another supported workload credential; - a dedicated
claude-code-v1worker may use a subscription only through the official Claude Code/Agent SDK path and under the then-current plan terms; - in the local-native profile, Claude Code discovers its existing local login itself; Light does not pass or store that login;
- personal Claude credentials must never become a shared enterprise provider.
Anthropic’s subscription-backed Agent SDK and non-interactive usage has its own credit and policy rules. Treat those rules as time-varying release inputs, not as a permanent platform entitlement.
Skills, Tools, And Workflows
Skills guide the harness; they do not authorize it. At turn admission,
light-agent resolves approved skill package versions and the runner
materializes only their verified contents. The worker intersects any tools a
skill mentions with the immutable execution profile, lease allowlist, adapter
manifest, and live availability.
Tool placement remains explicit:
- remote enterprise API/MCP calls execute through
light-gateway; - shell, filesystem, browser, and local MCP calls execute inside the leased sandbox through their bound dispatcher;
- push, PR creation, signing, publish, and deploy execute as fixed actions over an accepted artifact.
Durable multi-step work remains a light-workflow responsibility. A coding
harness may request a typed workflow handoff, but it does not own workflow
state, retries, timers, or approval transitions.
Approval, Cancellation, And Recovery
- Permission-bypass or blanket auto-approval flags are prohibited.
- An adapter approval request is normalized and checked against Light policy.
- Waiting for human approval ends executable authority unless a separately bounded checkpoint hold is allowed.
- Cancellation terminates the complete process tree. Enterprise execution also revokes the broker grant before cleanup. Local subscription execution has no broker grant to revoke, so cancellation invalidates the Light lease and stops further calls without claiming to revoke the user’s vendor login.
- Runtime events are sequence-checked and journaled; duplicates are idempotent.
- A worker crash, protocol gap, digest mismatch, expired lease, or ambiguous
side effect produces
unknownor failure for reconciliation, never assumed success. - Resume requires the same principal, repository base, policy, adapter/image, capability digest, model route constraints, and unexpired sandbox session.
Current Implementation Inventory
The following distinction prevents this target design from being read as a shipping claim.
Present In The Repository
agent-runtime-protocolversion1.4defines versioned worker commands/events, runtime identity and fencing, strict contiguous event sequencing, capability digests, bounded frames, and broker-grant admission. Its capability document declares adapter protocol version, approvals, streaming, session reuse, thread/turn identity, checkpoint, and usage support.coding-agent-runtimedefines a closed, digest-boundCodingAdapterContract,CodingTurnSpec, structured implementation and review artifacts, immutablecoding-implement-v1andcoding-review-v1profiles, remediation-chain validation, the review closure gate, canonical patch validation, and explicit migration dispositions for the shipped adapter identifiers.light-agentschedules the digest-boundcodex-app-server-v1shared-worker contract with an immutable Git bundle.light-workflow-runnerstages that bundle into an execution-specific private directory and independently validates the worker’s bounded canonical patch.light-agent-workerhosts pinned Codex0.153.4over local stdio, maps App Server lifecycle, streaming, approval, usage, error, and cancellation events, exports the validated implementation artifact, and runs review in a separate workflow-controlled thread over a reconstructed candidate with only an external build scratch directory writable. Reviewer output is constrained to the structuredCodingReviewResultschema and candidate mutation fails the turn.light-github-action-providerrequires an approved review bound to the exact implementation patch before its fixed create-branch or open-PR action can materialize or publish the patch.- Exact generated JSON Schema and TypeScript artifacts plus binary and schema
provenance are stored under
contracts/codex-app-server/v0.153.4. light-agenthas an immutablecodingProfileprojection point.llm-gatewayimplements the/v1/responsesclient surface, and its product design documents Codex custom-provider configuration and logical aliases.llm-gatewayaccepts a principal context and audits a principal digest, charged cost, and usage completeness with per-principal admission controls.- The enterprise coding profile binds the authenticated user, workload actor, optional workflow, session, turn, action attempt, logical route, billing subject, budget policy, policy, data boundary, and correlation ID into one canonical attempt digest. The runner releases only a short-lived credential envelope matching that exact digest and audience.
- The coding turn carries an immutable
personal-subscriptionorenterprise-apiauthentication profile. Runner admission keeps native Codex homes and enterprise brokers in separate pools, validates owner-only local credential-store placement, and rejects user, billing-subject, or host substitution. - Authentication audit evidence is a closed metadata-only record. It contains the profile, credential source, optional broker generation, and usage authority, but no token, account, email, or subscription-plan material.
- Attempt credential envelopes carry schema, identity, generation, issue, expiry, and revocation state. The broker validates and re-reads the envelope at delivery so rotation and revocation before process launch fail closed.
- The worker creates a private ephemeral Codex home containing only the trusted
light_gatewayResponses provider. The attempt token is placed only in the App Server environment and is excluded from Codex-created shell and tool environments. codingWorkerEligiblemakes provider compatibility explicit: an alias is rejected unless it is generation-only, requires streaming and tools, and every deployment has current passing conformance evidence.llm-gatewayprovides a signed usage-receipt contract whose verification covers the complete attempt binding and normalized token/cost result.
Not Yet Implemented Or Qualified
- An embedded Codex Rust adapter.
- A production Claude Code worker adapter.
- The Development Workflow Orchestration state machine that automatically schedules repeated implementer/reviewer rounds and persists the feature-wide finding ledger. The coding harness now enforces each immutable review and remediation handoff and blocks fixed publication until closure.
- Live packaging/browser qualification of the same local contract through both
portal-config-loc/all-in-ltandlight-portal-install. - End-to-end Codex-to-
llm-gatewaycompatibility across every proposed physical provider. - Durable normalized per-user token accounting, budget-window reservation, receipt emission/storage, and reauthorization for workflows that outlive the initiating JWT remain owned by Enterprise Workflow Phase E1. Existing gateway audit persistence records cost and usage completeness but not normalized token counts.
The existing CodexJsonl enum value validates a deprecated generic structured
CLI launch; it is not the App Server integration described here.
Delivery Plan
Phase 0: Freeze The Contract
Give the currently shipped CodingAdapter values an explicit migration
disposition:
| Current enum value | Target disposition |
|---|---|
CodexJsonl | Keep only as a deprecated generic CLI compatibility identifier during migration. It must not alias codex-app-server-v1; remove it after callers migrate unless it receives its own qualified adapter contract. |
ClaudeStreamJson | Keep only as a deprecated, non-production compatibility identifier. Remove it unless a separately versioned claude-code-v1 contract and qualification suite are delivered. |
GeminiJson | Keep only as a deprecated, non-production compatibility identifier. A future Gemini worker requires its own versioned adapter contract and qualification suite. |
KiloJson | Keep only as a deprecated, non-production compatibility identifier. A future Kilo worker requires its own versioned adapter contract and qualification suite. |
- Capture the existing Pi scheduling, sandbox, digest, RPC, and canonical-patch behavior as the initial cross-adapter conformance suite.
- Replace Pi-specific coding-policy and scheduling names with a runtime-neutral
adapter selection bound to immutable adapter, image, capability, and template
digests. Admit
pi-rpc-v1only as a temporary legacy selection. - Extend the runtime capability document for approvals, streaming, session reuse, thread/turn identity, usage, and adapter protocol version.
- Define adapter compatibility and image-digest records in the immutable coding profile.
- Define structured implementation artifacts and review findings.
- Add negative fixtures for unknown fields, oversized frames, wrong fencing, sequence gaps, expired grants, and permission-bypass options.
Exit gate: protocol and policy fixtures pass without starting a vendor harness, and the generic contract captures every Pi behavior required for Codex replacement without making Pi part of the target adapter set.
Implementation status: complete. The protocol, policy, scheduler, immutable adapter binding, artifact schemas, and negative fixtures are implemented and covered by unit/conformance tests. The temporary Pi migration selection has been removed by Phase 1.
Phase 1: Codex App Server Worker
- Build
codex-app-server-v1on the shared worker core. - Pin the Codex binary and generate JSON/TypeScript schemas from that exact release; use the JSON schema to validate Rust protocol types.
- Implement initialization, authentication status, thread/turn lifecycle, streaming items, approvals, cancellation, usage, errors, and shutdown.
- Qualify local stdio only; do not expose a shared App Server socket.
- Run the Pi baseline and Codex cases through the same worker-contract test matrix; adapter-specific protocol details may differ, but lease, artifact, cancellation, and authority outcomes must match.
- After that matrix and migration gate pass, remove the Pi scheduling path, adapter, image, template, capability advertisement, enum value, profile, and Node/npm dependency.
Exit gate: one sandboxed Codex turn can inspect, edit, test, emit a bounded canonical patch, cancel cleanly, and fail closed on version/schema mismatch; the legacy Pi product/runtime artifacts listed above are removed.
Implementation status: complete. The shared worker uses only the pinned local
stdio App Server, verifies both binary and generated-schema digests, performs
the initialize/account/thread/turn sequence, maps streaming and usage, denies
unbrokered approval requests, translates cancellation to turn/interrupt, and
shuts down the process tree. The runner supplies staged immutable input and
revalidates the canonical patch before publishing it. Pi product/runtime
artifacts are removed. App Server frames are bounded before JSON parsing,
inline patches are limited to 128 KiB so their complete event fits the 1 MiB
runtime frame, and the phase gate launches the pinned App Server for a live
initialize/account lifecycle smoke test.
Phase 2: Enterprise Gateway Routing
- Treat Enterprise Workflow Phase E1 as the prerequisite and
single delivery owner for audience-bound delegation, reservations, normalized
usage storage, reconciliation, and signed receipts in
llm-gateway. - Add trusted custom-provider configuration and an attempt-scoped credential helper.
- Qualify Codex against
/v1/responsesfor every eligible logical alias. - Integrate the worker with that gateway contract and verify receipt binding to the exact user, actor, workflow, session, turn, route, and attempt.
- Prove provider and gateway credentials are absent from workspace, child shell, process arguments, logs, runtime events, and artifacts.
Exit gate: the pinned Codex worker completes buffered, streaming, tool-use,
cancellation, usage, and error scenarios through llm-gateway without learning
a physical provider or credential, and every call is charged to the verified
billing subject with its distinct workload actor and correlation IDs.
Implementation status: complete at the coding-harness boundary. The scheduler,
runner, worker, and gateway use the versioned attempt binding; mismatched user,
turn, route, audience, or billing bindings fail closed. Codex receives a
trusted ephemeral custom-provider configuration, while its shell environment
excludes the attempt token. Gateway tests cover buffered Responses, streaming,
tool events, cancellation, usage, errors, route pinning, provider eligibility,
and signed receipt tamper detection. scripts/run-coding-harness-phase2-gates.sh
composes these checks with the complete Phase 1 gate.
Production enablement remains conditional on the separately owned Enterprise Workflow Phase E1 ledger, receipt persistence/emission, token exchange, and reauthorization services. A physical provider/model is suitable for a coding worker only after its current conformance evidence satisfies the alias; unsupported Responses transformations remain ineligible rather than being routed optimistically.
Phase 3: Implement And Review
- Add immutable
coding-implement-v1andcoding-review-v1role execution profiles, routed through thecoding-implementerandcoding-reviewerlogical model aliases. - Reconstruct review input from the accepted base and patch in a fresh read-only repository tree and thread, with only an ephemeral build scratch directory writable and excluded from the canonical patch.
- Enforce structured findings and remediation loops.
Exit gate: tests prove the reviewer cannot mutate the candidate workspace and cannot observe implementer-private thread state, while a blocking finding prevents the fixed publish action.
Implementation status: complete at the coding-harness and fixed-publication
boundary. light-agent selects the two pinned profiles and aliases from trusted
policy; the runner canonicalizes implementation artifacts and validates review
results; the worker reconstructs the accepted patch in the reviewer thread with
scratch-only writes; remediation inputs must carry the complete prior finding
set; and the GitHub fixed action rejects missing, mismatched, or blocking review
evidence. scripts/run-coding-harness-phase3-gates.sh composes the Phase 0-2
qualification with the role, isolation, structured-output, remediation, and
publication-blocking tests.
Automatic multi-round feature orchestration and its durable finding ledger remain owned by Development Workflow Orchestration; they consume these Phase 3 contracts rather than weakening or duplicating them.
Phase 4: Authentication Profiles
- For the local Portal profile, launch the native harness in its dedicated user context under the documented local native isolation profile, check only authentication status, and never add a Light-managed vendor-credential store or token pass-through.
- Run the same harness and interaction qualification through both
portal-config-loc/all-in-ltandlight-portal-install. - For the enterprise profile, add workload token issuance, rotation, and revocation through the trusted credential broker.
- Keep subscription routes physically and logically separate from pooled
enterprise-apiworkers. - Record authentication class, never secret material, in audit evidence.
Exit gate: cross-user, cross-tenant, gateway-proxy, and expired/revoked credential tests fail closed.
Implementation status: complete at the coding-harness and runner-pool boundary.
The immutable coding turn now carries exactly one authentication profile.
personal-subscription requires an owner-only, runner-projected native Codex
home, rejects any broker or enterprise gateway, accepts only an authenticated
ChatGPT account status, and records advisory-usage metadata without account or
secret fields. enterprise-api requires the exact user/host-bound gateway and
attempt broker, rejects native credential-store visibility, and records only
the authentication class, broker source, credential generation, and
authoritative-usage flag.
Local native pools are single-concurrency. Enterprise pools require the pinned per-attempt sandbox-launch contract and restricted-egress profile; the runner advertises those features only when that contract is configured.
Attempt credential envelopes are versioned and bound to a unique credential ID
and positive generation. The broker re-reads the owner-only envelope at
delivery, enabling pre-delivery rotation, and rejects zero-generation,
future-issued, expired, overlong, mismatched, or revoked credentials. Process
cancellation still terminates credential use; durable token minting and gateway
revocation remain owned by the enterprise token-exchange service described in
Enterprise Workflow Phase E1.
Local distribution conformance is
expressed against the same Portal/workflow/agent contract for both
portal-config-loc/all-in-lt and light-portal-install; their packaging and
browser qualification remain distribution release gates rather than a second
coding-worker implementation.
The runner accepts an explicit agentWorker.codexExecutable for host-native
personal pools and projects it as LIGHT_CODEX_EXECUTABLE. This permits an
unprivileged local installation to use the exact qualified native binary
without a /usr/local/bin shim. The personal App Server smoke is enabled with
LIGHT_RUN_CODEX_PERSONAL_SMOKE=1; it requires an authenticated ChatGPT
CODEX_HOME, completes a real ephemeral turn with the native default model,
and can assert that the LLM audit row count is unchanged. Logical Light aliases
are sent as App Server model names only in the enterprise gateway profile;
personal subscription turns use the native Codex default model.
scripts/run-coding-harness-phase4-gates.sh composes every earlier coding
harness gate with the profile-separation, account-status, credential lifecycle,
user/tenant binding, audit-metadata, and documentation checks.
Phase 5: Optional Adapters
Phase 5 is implemented as a fail-closed optional-adapter qualification layer:
CodingAdapterQualificationseparates launch contracts from promotion evidence. A selectable adapter must bind the exact launch-contract digest and have evidence for all 13 lifecycle, isolation, authentication, dependency, and output dimensions.- Immutable agent policy carries the separately issued qualification record.
Both
light-agentand the worker require itscontractDigestto equal the exact admitted launch contract and itsevidenceDigestto equal the reviewed manifest. Runtime code cannot manufacture a qualified record from an incoming candidate contract. codex-app-server-v1is the only qualified production adapter. Its evidence manifest is digest-bound to the worker and composes the Phase 1 through Phase 4 gates.prototypes/codex-embedded-v1pins the official Codex0.153.2source revision and upstream Cargo patches, compile-probes the exportedThreadManagerandStartThreadOptionstypes, and benchmarks only direct typed-call overhead against JSON boundary translation. Its isolated lock currently contains 1,122 packages, which confirms that dependency size and release coupling are material costs rather than theoretical risks.codex-embedded-v1is recorded asprototype-only: only dependency and license dimensions have evidence, it has no launch-contract digest, it is absent from worker capabilities, and selection fails closed.- No
claude-code-v1worker is shipped. The proposed local subscription and independent-harness use case is now described in Claude Personal Worker; implementation and qualification remain pending. Anthropic models remain routable behind logical aliases without implying Claude Code harness semantics. - Future native harnesses, including Grok-oriented workers, require a new versioned adapter, an exact evidence manifest, and the same complete matrix. Model parity never implies adapter parity and adapters never silently fall back to one another.
scripts/run-coding-harness-phase5-gates.sh composes all earlier gates,
validates evidence digests and the fail-closed selection contract, and ensures
unqualified adapter IDs do not enter production selection paths. Setting
LIGHT_RUN_CODEX_EMBEDDED_PROBE=1 additionally compiles and runs the pinned,
network-dependent embedded probe; it does not promote the adapter.
Acceptance Criteria
light-agentselects a pinned adapter and logical model alias from immutable policy; prompt input cannot change either.- No Pi runtime, profile, template, capability, enum, image, or Node/npm dependency remains in the supported product.
- The long-lived service never starts a coding harness or receives its reusable provider/subscription credential.
- App Server runs locally inside one leased sandbox and communicates over bounded structured messages.
- The local native sandbox retains lease, root/path, network, resource, cancellation, and artifact controls while explicitly making the weaker same-user credential-isolation claim documented above.
- A model call uses exactly one authentication path: native vendor subscription
or Light workload/API credentials through
llm-gateway. - In the local Portal profile, Light passes no OAuth token, API key, or subscription credential to the native harness; the harness uses its existing local login while Portal identity separately protects API and approval operations.
- An enterprise model call uses a short-lived delegated token that binds the initiating end user and workload actor; the original portal JWT is not stored in the worker or harness.
llm-gatewaykeeps provider API keys, enforces the per-user or cost-center budget, and records normalized token usage and charged cost for every provider attempt.- The implementer emits a canonical patch relative to an immutable base; the trusted runner enforces protected paths and artifact limits.
- Review runs in its own stage-scoped thread over a clean read-only repository reconstruction, uses only excluded ephemeral build scratch, and emits schema-valid findings.
- Model diversity can be enabled without deploying a second harness; harness diversity requires a separately qualified adapter/image.
- Cancellation kills the process tree and produces cleanup evidence within the lease deadline; enterprise execution also revokes model access, while local execution invalidates the lease without claiming to revoke the native login.
- Unsupported protocol features, unqualified provider routes, unknown adapter versions, and stale capability digests fail before repository mutation.
- Fixed high-value actions accept only the approved immutable artifact and never the live coding workspace.
- Every mutable coding session has an exclusively leased workspace set; no two users or sessions share writable repository, Git metadata, build, or scratch state.
Settled Decisions And Residual Risks
- App Server is the primary integration. Light pins its protocol and binary, upgrades them with each admitted Codex release, and blocks rollout until the versioned conformance suite passes; protocol churn is accepted release work.
- Embedded Codex crate linkage, if delivered, follows the same upstream-release cadence. Light will not fork Codex crates; an upstream incompatibility blocks that adapter upgrade or causes the embedded adapter to be withdrawn.
- Pi has been removed after Codex replacement qualification, including its runtime, Node/npm dependency, and published profile.
llm-gatewaywill preserve Responses semantics where possible. A provider or model that cannot faithfully support required reasoning, tool, streaming, usage, or cancellation behavior is marked ineligible for coding-worker aliases rather than receiving a lossy transformation.- Subscription terms, entitlements, and quotas remain release inputs. Light may change authentication behavior or disable an affected route when terms change, and the adapter contract remains open to other qualified coding workers.
- Review assurance follows the policy matrix above: model diversity, harness diversity, and human approval are independent controls selected by risk.
- Session reuse is allowed only inside one exclusively leased workspace set. Cross-user or cross-session writable sharing is prohibited; fresh reviewer workspaces remain the default independent-review boundary.
References
Internal:
- Light-Agent
- Light-Agent Execution
- Personal Development Workflow Orchestration
- Enterprise Development Workflow Orchestration
- Centralized Agent Skills
- LLM Gateway API
Vendor documentation, verified September 4, 2026:
- OpenAI Codex App Server
- OpenAI Codex SDK
- OpenAI Codex configuration reference
- OpenAI Codex authentication
- OpenAI Codex code review
- Anthropic Claude Code authentication
- Anthropic third-party subscription guidance
- Anthropic subscription and API billing separation
Workspace isolation reference:
Claude Personal Worker
Status: Phases 0, 1, 2, and 3 implemented September 12, 2026 for the pinned Linux x86_64 local candidate. Phase 2 includes Agent admission, dedicated runner selection, fenced worker events, canonical artifacts, independent review, and workflow-controlled sessions. The opt-in worker is locally technically qualified; production/distribution eligibility remains open. Phase 3 supplies repeatable local configuration and published-Agent-to-runner verification in both service layouts. See the Phase 2 record for the exact implemented boundary. Other sections also describe later target capabilities.
This specializes Coding Harness Integration for a user operating their own local Claude subscription. The concrete use case is implementing or independently reviewing a coding task with the native Claude Code harness alongside the tested Codex personal worker.
Recommendation
Build claude-code-v1 as a separately qualified adapter in the existing Rust
light-agent-worker. Initially launch one official Claude CLI process per
attempt, consume bounded JSON Lines, and retain the existing Light lease,
workspace, cancellation, and canonical-artifact boundaries.
An App Server equivalent is unnecessary for this first milestone. Anthropic provides Python and TypeScript Agent SDKs and explicitly recommends a CLI subprocess for other languages. There is no Rust SDK in the documented SDK language set. This is a supported integration surface to investigate, rather than a reason to reverse-engineer a private server protocol. Agent SDK overview.
The recommended first release supports workflow-controlled multi-turn sessions with a policy-selected constrained or trusted-personal automation profile. Each job remains a bounded turn in its own process. Session reuse is required for qualification; interactive approvals and automatic recovery of uncertain turns are not advertised. If Portal-mediated approvals are required for the initial use case, qualify the SDK bridge described below before enabling that profile; do not emulate approvals by parsing terminal text.
Current Repository Boundary
The current production worker is explicitly Codex-specific:
apps/light-agent-worker/src/lib.rsadvertises Codex capabilities and usescodex_app_server.rs,coding_session.rs, andworkspace.rs.crates/coding-agent-runtime/src/lib.rspins the Codex binary/protocol and defines adapter qualification, authentication, roles, and artifact contracts.contracts/coding-adapters/codex-app-server-v1-qualification.jsonrecords evidence across 13 qualification dimensions.crates/model-provider/src/claude_code.rsalready invokes Claude CLI as a model provider. It buffers output and uses--dangerously-skip-permissionsin its agent path. It is not the leased coding-worker integration and must not be reused as the worker’s launch or permission policy.
The existing ClaudeStreamJson compatibility identifier does not establish a
qualified adapter. A new launch contract and evidence record are required.
Keep adapter ID, role profile, authentication profile, and logical model alias
separate, as in the parent design.
Integration Options
| Option | Benefit | Cost or limit | Decision |
|---|---|---|---|
Rust launches claude -p and parses JSON Lines | Small dependency footprint; clear process boundary | Narrower lifecycle and approval integration | First prototype and bounded-turn candidate |
| Rust launches a pinned Python Agent SDK bridge | Typed SDK messages, permission callbacks, interruption and session APIs | Additional interpreter, SDK, and bridge qualification | Preferred escalation path for interactive features |
| TypeScript Agent SDK bridge | Similar programmatic control | Adds Node/runtime packaging | Alternative if deployment ownership favors TypeScript |
| Rust implements SDK-internal bidirectional messages | No interpreter | Assumes responsibility for an internal protocol | Reject unless upstream publishes a suitable supported contract |
| PTY automation of interactive Claude | Looks like local terminal usage | Fragile prompts, escape sequences, approval ambiguity | Reject |
| Direct Anthropic API calls using native login credentials | None within this architecture | Bypasses native-client and credential boundaries | Reject |
The Agent SDK is distinct from the Anthropic API client SDK: the latter would require Light to implement the coding loop and does not turn a subscription into API credits. The SDK bridge remains a local child of the worker, never a shared network service. SDK comparison.
Architecture And Ownership
The diagram shows two configured deployments of the same light-agent
application. Codex is production-integrated; Claude has an opt-in locally qualified adapter and dedicated worker. The workflow never
calls light-agent-worker directly. It selects an agent definition, and that
agent’s configured coding policy selects the worker and adapter.
flowchart TB
WF[light-workflow<br/>Select target agent definition<br/>Own session new / resume / close]
JOBS[Durable agent jobs<br/>agent_job_t<br/>Target agent_def_id and typed coding input]
AC[light-agent deployment: codex-personal<br/>Codex agent definition and pinned coding policy]
AL[light-agent deployment: claude-personal<br/>Claude agent definition and pinned coding policy]
CTRL[controller-rs<br/>Execution placement and leases]
RC[Codex personal runner]
RL[Claude personal runner]
WC[light-agent-worker: Codex variant<br/>Shared core plus codex-app-server-v1 adapter]
WL[light-agent-worker: Claude variant<br/>Shared core plus claude-code-v1 adapter]
CX[Local Codex App Server<br/>Native persisted threads and login]
CL[Local Claude CLI<br/>Native persisted sessions and login]
WF -->|Resolve selected agent and enqueue job| JOBS
JOBS -->|Jobs targeting Codex definition| AC
JOBS -.->|Jobs targeting Claude definition| AL
AC -->|Admitted Codex execution| CTRL
AL -.->|Admitted Claude execution| CTRL
CTRL --> RC
CTRL -.-> RL
RC -->|Launch leased attempt| WC
RL -.->|Launch leased attempt| WL
WC -->|Local stdio JSON-RPC| CX
WL -.->|Local subprocess stdio| CL
How the workflow selects Codex or Claude
Use two logical agent definitions, illustratively named codex-personal and
claude-personal, with different agent definition IDs and separately pinned
coding policies. The workflow’s agent-call target resolves through the agent
catalog to the selected definition. In the existing service-mode path,
light-workflow writes that target into agent_job_t.agent_def_id, together with
the typed coding input and workflow correlation. The matching light-agent
deployment reconciles the job, authorizes the turn, and submits its execution to
the controller/runner path. The runner starts the worker variant fixed by that
agent’s policy. A target name alone is not authority: catalog, policy, ownership,
and execution admission still apply.
For example, the implementation task targets the Codex definition and receives
an implementation result plus checkpoint. The review task targets the Claude
definition with the candidate patch and its own session reference. Follow-up
review tasks keep targeting that same Claude definition with mode: resume and
the expected checkpoint. The workflow receives durable results through the agent
job path; it does not address either native CLI or worker process.
| Layer | Concrete local arrangement |
|---|---|
| Application code | One light-agent application/binary, reused by both deployments |
| Logical agents | Two definitions: Codex personal and Claude personal |
| Agent service processes | Two configured deployments, initially one process each |
| Coding policy | One pinned adapter/worker contract per deployment |
| Personal runners | Separate configured Codex and Claude targets, each single-concurrency |
| Worker processes | Launched per admitted attempt; each contains the common core and its selected adapter |
| Native conversations | Separate persistent sessions per workflow stage and role; process exit does not close them |
This corrects the earlier statement that one current process could host both
coding policies. apps/light-agent/src/main.rs loads
agentPolicy.execution.codingProfile at startup through
coding_profile_from_policy; the current design must not assume a per-job
multi-profile dispatcher. Hosting both definitions in one process would require
an explicit future change to policy loading, job dispatch, isolation, and
qualification. It is not required to implement the Claude worker. Two
deployments may run on the same machine with distinct service configuration;
they do not require two implementations of light-agent.
Both workers reuse the trusted worker core, but the diagram represents separate execution paths, not a single worker running both adapters. Agent admission owns policy; controller/runner owns placement and leases; worker core owns runtime validation and artifact proposals; adapters translate native protocols. Results return through runner and agent into the durable workflow job. Claude’s path becomes selectable only after implementation and qualification.
A new worker or CLI process may be launched for every turn while the same native
conversation is resumed. Implementer and reviewer sessions remain separate, and
neither adapter can resume the other vendor’s history. Both use the common
coding.thread lifecycle; the workflow closes a session explicitly rather than
interpreting process exit as conversation closure.
The worker launches only the server-selected binary in the admitted working directory. The prompt cannot set a binary, environment, provider, credentials, permission mode, MCP server, or writable root. The controller still owns leases; Claude session identifiers never replace Light identity or fencing.
Use a dedicated runner bound to one Portal host and user, with
maximumConcurrency: 1 and local-single-user-native-v1. Codex and Claude must
also share an exclusive workspace lease when targeting the same task; separate
vendor accounts do not make concurrent writes safe. Shared task workspaces obey
Shared Task Workspaces.
This retains the parent design’s weaker trusted-local-user security model. A same-user shell can potentially access that user’s native credentials. CLI permission rules and post-run diff validation do not establish OS isolation. Require proven filesystem/process controls for each advertised role; where they cannot be established, reject the role. Hostile repositories require a stronger separately qualified sandbox and authentication arrangement.
Subscription Authentication And Eligibility
The intended setup is a user installing the official CLI and logging in through Anthropic’s native flow outside Portal. Light launches the CLI in that dedicated user context. Light does not receive, copy, export, broker, or persist OAuth credentials, and does not implement a Claude login button.
There is a material distinction between technical feasibility and permission to ship this integration:
- Anthropic’s support article has a June 15 update stating that proposed billing
changes are paused and SDK,
claude -p, and third-party app usage still draw from subscription limits. The older credit announcement below that update is explicitly superseded. Current support update. - Its compliance page still directs product developers to API authentication and restricts offering Claude login or routing subscription credentials on users’ behalf. Authentication and credential use.
These pages do not conclusively authorize this particular Portal integration. Keep personal deployment conditional on recorded eligibility for the intended user-owned local use and current plan terms; resolve the discrepancy with Anthropic before distributing it as a supported subscription integration. This does not block designing or testing the adapter with synthetic fixtures. An SDK bridge does not change the eligibility question.
At startup, qualify a documented native authentication-status check for the exact
CLI release, parse only the minimum authentication class, and discard account
identifiers. Reject API-key/cloud-provider mode, ambiguous status, or missing
login before workspace mutation. Do not inspect credential-file contents as an
alternative status API. Native reauthentication is a user action outside the
attempt; report an actionable authentication-required result.
Construct an allowlisted environment and reject incompatible inherited settings,
including API keys, OAuth-token overrides, custom API base URLs, API-key helpers,
and Bedrock/Vertex/Foundry selection. Audit native settings as well as environment
precedence. Do not silently fall back to API billing, another account, Codex, or
llm-gateway. The workflow may request a native model through the proposed typed
model-selection field below. If omitted on a new session, use the configured
agent default. Do not send Light role aliases such as coding-reviewer as literal
Claude model names.
Subscription usage is advisory telemetry. It is neither an authoritative invoice nor proof of remaining quota. Exhaustion yields a bounded retryable outcome for the workflow, without a tight retry loop or an automatic API purchase.
Workflow-selected Native Model
Yes: the target design permits the workflow to choose a Claude model for a coding
job, independently of its choice of agent definition and implement/review role.
The worker maps the admitted choice to the CLI’s documented --model option,
which accepts native aliases or full model IDs.
CLI model selection.
Implemented fragment inside the typed coding input:
{
"nativeModel": "opus"
}
A workflow may populate this field with an expression. Agent admission checks it
against its configured allowed native models and account eligibility, resolves
any configured mapping, and binds the effective choice into the immutable turn
spec. The worker then launches Claude with --model <admitted-model>. Unknown or
unavailable choices fail explicitly; no silent model or paid-provider fallback.
Precedence: on new, explicit workflow choice overrides the agent default;
on resume, an omitted choice preserves the session’s admitted model. Initially
require an explicit resume choice to match that model. To change models, the
workflow closes the session and starts a new one with the requested choice.
In-session switching can be a later qualified capability; it must never happen
implicitly because an installation’s default changed. Record requested and
observed effective model in turn evidence, and detect native alias drift when
exact reproducibility is required.
nativeModel stays separate from CodingTurnSpec.modelAlias, which still
identifies the immutable role profile. Agent admission validates the typed public
coding.nativeModel field and projects it into the runtime envelope alongside
the server-owned claudePolicy. The worker binds the resolved native model into
its checkpoint. Do not replace coding-reviewer with opus or put the selection
only in the prompt.
The same design can later support Codex’s native model selection, but this note does not claim the current personal Codex worker accepts such an override.
Launch And Configuration Contract
The documented headless surface supports -p, --output-format stream-json,
--verbose, and partial-message streaming. It also supports stdin prompts and
reports a terminal result. Critically, --bare does not read subscription OAuth
credentials or the keychain; it cannot be the personal worker’s isolation switch.
Programmatic CLI usage.
Illustrative process shape, not a qualified production command:
<absolute-pinned-claude-executable>
-p
--output-format stream-json
--verbose
--include-partial-messages
--permission-mode dontAsk
Rust writes the bounded prompt to stdin, closes input for the one-shot profile,
and drains stdout and stderr concurrently. It uses tokio::process::Command
with explicit arguments, never a shell command assembled from prompt text.
A one-shot process is not a one-shot conversation: persistent jobs add the
explicit session selection described below and retain native session storage.
This example illustrates the constrained profile. The actual launch uses the
permission source and mode selected by policy; native inheritance omits the
--permission-mode dontAsk override shown here.
The CLI reference documents --safe-mode with authentication retained,
--restricted, --setting-sources, --settings, and --strict-mcp-config.
Settings overrides are not a guarantee that omitted settings disappear, and
managed policy can still apply. CLI reference.
For permissionSource: agent-policy, Phase 0 should first test --safe-mode
plus a trusted explicit policy with the native login intact. Qualify all relevant flag combinations rather than assuming
individual flags compose. For that Light-managed configuration mode, the required
outcome is:
- No automatic project/user hooks, plugins, commands, agents, remote sessions, MCP servers, or memory can execute or widen authority before admission.
- Approved repository instructions and centralized skills are supplied as reviewed context, without automatically activating their executable helpers.
- Managed configuration is inventoried and digest-bound or rejected when it introduces unapproved execution. Initialization telemetry is useful evidence, but is too late to prevent startup hooks.
- The native authentication store remains available only in the documented local-user context; changing the configuration directory must not be assumed to preserve authentication across platforms.
For permissionSource: claude-cli, normal native configuration discovery is
intentional, as described below; the preceding customization-suppression rules
do not apply. Qualify each source mode separately. If the pinned CLI cannot
satisfy the selected mode and native login, that mode fails qualification. Investigate the SDK configuration controls or
a stronger runner boundary. Permission bypass cannot substitute for configuration
containment or runner enforcement, even when explicitly enabled by policy.
Pin the installed artifact and version, detect updates before every launch, and
reject drift until requalification. Do not choose a rolling latest release.
Permissions And Interactive Features
Choose the permission source in agent policy
Support two explicit sources. For a dedicated trusted personal coding machine,
recommend claude-cli: reuse the installation the owner already maintains.
Choose agent-policy when centralized tool rules and reproducibility are needed.
This is a deployment choice, not a claim about measured user preferences.
| Policy source | Permission/configuration authority | Worker behavior |
|---|---|---|
agent-policy | Light supplies the admitted native tool and permission settings | Materialize the trusted configuration and suppress unapproved ambient customization |
claude-cli | The owner delegates native tool decisions to the installed Claude configuration | Use installed settings and admitted native customization within the qualified filesystem namespace |
Implemented fields inside agentPolicy.execution.codingProfile.claudePolicy:
{
"permissionSource": "claude-cli",
"permissionMode": "inherit",
"defaultModel": "sonnet",
"models": { "sonnet": "claude-sonnet-5" },
"tools": [],
"allowedTools": []
}
The published agent policy is delivered through the normal trusted configuration path, validated by Agent admission, and bound into the worker execution contract. The prompt or an arbitrary job field cannot change this delegation. Do not require users to duplicate native allow/deny rules in Light. In native mode, Light validates the selected source/profile rather than pretending it has a complete independently enumerated list of Claude tools.
permissionMode: inherit means omit permission-mode/tool override flags and let
Claude resolve its native rules. Do not add --safe-mode, --bare, restrictive
setting-source flags, or an empty MCP configuration that would defeat the chosen
behavior. Process framing, session selection, and explicitly admitted model
selection still use their own flags. Claude’s documented settings precedence
continues to apply; user settings are not the only source, and managed policy may
constrain the installation.
Native settings and precedence.
For a machine explicitly dedicated to unattended automation, the owner may use:
{
"permissionSource": "claude-cli",
"permissionMode": "bypassPermissions",
"defaultModel": "sonnet",
"models": { "sonnet": "claude-sonnet-5" },
"tools": [],
"allowedTools": []
}
This retains native customization while explicitly requesting Claude’s permission
bypass mode at launch. There is no need to approve each ordinary coding tool in
Light. If the installation already achieves the desired behavior, inherit is
sufficient. Inheritance alone does not mean allow-all, and unattended execution
does not imply consent. Never silently change inherit into bypass when a tool
is denied. A native managed restriction or external service authorization still
applies; fail clearly if the requested mode is unavailable.
With interactionMode: unattended, use the pinned CLI’s qualified no-prompt
behavior and return recorded denials or a needs-input outcome when work cannot
proceed. Do not emulate human input by writing y, leave a process waiting
indefinitely, or override a denial. Headless behavior can differ from the terminal;
qualify it explicitly. A future interactive mode requires a structured Light
interaction bridge before it can be selected.
Headless execution.
Native mode deliberately trusts configuration discovered in the admitted working directory, including repository-provided executable hooks. Bind the dedicated OS user, runner, and workspace to that delegation before launch. A staged checkout may not contain the original checkout’s ignored local settings or resolve its relative paths identically: qualify the runner’s workspace projection so native configuration behaves as promised. Do not silently import arbitrary files from other repositories or copy credential stores into workspaces.
Reuse current native configuration on each new invocation without requiring a new agent policy publication for every local rule edit. Record source, requested mode, and a secret-free configuration revision/fingerprint where observable; do not log raw settings or credentials. Native mode intentionally permits owner configuration drift, so it does not promise reproducible tool policy. Keep the permission-source and override choice fixed within a Light session. Qualify native live-reload behavior and log observable changes; a resident-process mode must not claim it freezes configuration if Claude reloads it. Light-managed mode retains its stricter configuration binding.
This delegation covers native tool decisions, not Light identity, session ownership, leases, deadlines, concurrency, result validation, or the role’s external filesystem boundaries. A read-only review stays externally read-only; if that role conflicts with the selected installation, reject the incompatible profile rather than silently remove the review guarantee. Publication authority is configured separately as below. Dedicated-machine mode makes no claim of protecting the owner’s credentials from same-user code.
Implementation must replace the current fixed-tool-set assumptions with a typed source-aware policy contract across admission, worker capabilities, launch, and checkpoints. Add tests for both source modes, inherited native denies, inherited bypass, explicit bypass, managed restrictions, configuration discovery in staged workspaces, local edits between turns, prompt-injected overrides, cancellation, and unattended needs-input outcomes. These are requirements for the Claude worker, not a claim that current binaries accept the example fields.
Tool authority and execution boundaries
The Light-managed profile uses an explicit tool surface; the native profile delegates that surface to the installed CLI. Both retain the outer execution boundaries admitted by Light. An auto-approval list alone is not the complete tool surface or a filesystem boundary. Read-only reviews need an OS-enforced read-only candidate tree and separately writable scratch; allowing Bash can otherwise restore writes even if Edit is disabled.
The personal worker must be able to perform real development: read/search code, edit files, execute builds and tests, install dependencies, fetch documentation, and use approved MCP tools and subagents. These are supported design goals, not blanket prohibitions. Enable them through an explicit execution profile and qualify their actual tool surface, network access, credentials, and cleanup. Background tasks may run within the active lease and must stop at cancellation or turn completion; they do not become independent durable Light workflows.
Permit bypassPermissions / --dangerously-skip-permissions in an explicitly
configured trusted-personal automation profile. This preauthorizes native CLI
actions for that profile; it does not grant new Light permissions. A constrained
profile may instead use dontAsk and tool allow rules. Choose the launch mode
from trusted policy: either an explicit override or deliberate inheritance of
native settings. Prompt text cannot select either. Record the effective
mode in execution evidence. Under bypass, native permission prompts and approval
callbacks are not an enforcement boundary: the runner and external tool services
must enforce the admitted scope. Same-user credential isolation remains weak.
A read-only reviewer still requires an externally read-only candidate tree.
GitHub writes, push, deployment, and publication are permitted workflow actions. The default remains the existing fixed-action path after artifact acceptance. If a workflow explicitly delegates these operations to the native worker, define a separate publication-capable profile with scoped credentials, exact targets, approval/artifact binding, and audit receipts. That is an extension to the parent harness publication contract, not implicit authority for every coding turn. Until that extension is implemented, the workflow performs them through fixed actions; the coding agent can still implement, test, and propose the result.
For the initial adapter, advertise supports_approvals: false. A denied tool may
allow the model to continue, but the worker must preserve that denial and reject
a claimed completion when required work was not performed. A task needing
broader authority returns to Light admission as a new bounded attempt.
For interactive approval support, prototype a Python SDK bridge using the published permission callback. Its local Light-owned bridge protocol would carry start, permission request/decision, event, interrupt, and terminal outcome, each bound to attempt identity and sequence. The Rust core retains the authority:
- Validate tool name and normalized arguments against the lease.
- Bind approval to the exact request digest, tool-use ID, workspace, policy, user, attempt, and expiry; request only a permissible narrowing of authority.
- Wait within the lease deadline; deny on disconnect, cancellation, revocation, timeout, duplicate response, or argument mismatch.
- Forward the decision only to the originating live callback.
Callbacks can be bypassed by already permitted operations; qualify the complete permission evaluation path, not just callback happy paths. CLI MCP permission handling is another candidate only if the pinned public contract proves the same semantics. Neither path is a first-release capability by assumption.
Event And Completion Contract
Implement a Claude-specific parser, not a renamed Codex JSON-RPC parser. Proposed normalization follows this table; capture exact vendor schemas in Phase 0.
| Native observation | Light interpretation |
|---|---|
| Initialization/session metadata | Bind native session to this attempt; validate observed configuration |
| Text delta | Bounded progress, never proof of success |
| Tool-use/tool-result content | Sanitized execution evidence with native correlation IDs |
| Permission denial | Recorded limitation; no implicit grant |
| Usage and final result | Advisory usage plus candidate terminal outcome |
| Nonzero exit, malformed stream, missing result | Failed or indeterminate attempt |
Enforce byte limits before parsing, bounded nesting and total output, stderr limits, and backpressure deadlines. Keep the existing 1 MiB runtime-event and 128 KiB inline-patch ceilings. Partial deltas and full assistant messages can represent the same content: choose one presentation stream and avoid duplicate text or usage aggregation. Whitelist fields before emitting events; raw verbose frames, tool output, settings, and native transcripts can contain secrets.
Pin mandatory event envelopes and terminal semantics. Unknown security-relevant or lifecycle messages fail closed; additive diagnostic fields may be ignored under an explicit parser policy. Require one terminal result, acceptable native status, child-process cleanup, and successful artifact validation before Light reports success. EOF or exit code zero alone cannot satisfy completion.
The runner computes the diff against the immutable base, validates protected paths, symlinks, Git metadata, size limits, and workspace identity, and emits the existing artifact schema. Model-generated diffs are proposals. Review findings must pass the existing structured schema and reference the accepted candidate; prose claiming that review passed is insufficient.
Cancellation, Recovery, And Sessions
Use an adapter state machine:
admitted -> launching -> running -> validating -> completed
\ \ \
----------> stopping -> failed/cancelled/indeterminate
Cancellation, deadline expiry, runner disconnect, or lease loss fences new work immediately. Attempt graceful interruption only through a qualified signal or SDK call, then terminate the whole process group/cgroup within a fixed grace period and escalate to forced termination. Drain bounded output and verify no child or background tool survives. Do not depend exclusively on CLI cleanup. Artifacts from cancelled or uncertain attempts are quarantined, not published.
Do not retry an uncertain native turn in place. Repository edits or external side effects may already have occurred. Preserve attempt evidence and let the workflow authorize a fresh reconstruction; process restart is not exactly-once execution. Local cancellation does not revoke the user’s subscription login.
Workflow-controlled Session Continuity
Session support is required in the initial Claude adapter, matching
Workflow Coding Thread Lifecycle.
The existing native Codex implementation already uses coding.thread, private
checkpoints, and workflow-controlled new, resume, and close. Claude must
implement that same public contract rather than introduce a different lifecycle.
A successful job ends a process/turn, not the workflow’s conversation.
Claude documents two continuation strategies: --continue chooses the latest
conversation in the working directory, while --resume selects a conversation
by ID or name. --name supplies a display label; --session-id accepts a UUID.
Persistence must remain enabled: reject --no-session-persistence and equivalent
environment overrides. CLI reference.
Use explicit UUID identity for the worker. Names may be optional diagnostic
labels, never authorization or lookup keys. Reject ambient --continue,
name-based resume, caller-supplied transcript paths, and --fork-session in the
normal resume path. A local interactive invocation must not redirect the next
workflow turn to a different conversation.
Illustrative Rust argument construction (the supervisor must still apply the qualified binary, environment, workspace, tool policy, streaming limits, and cancellation controls):
#![allow(unused)]
fn main() {
// The worker allocates and privately records this UUID for the workflow session.
// These are separate invocations, potentially in separate worker processes.
let first_turn_args = [
"-p", "Review the new main.rs changes",
"--session-id", native_session_id.as_str(),
"--output-format", "stream-json", "--verbose",
"--permission-mode", "dontAsk",
];
let next_turn_args = [
"-p", "Focus on memory safety of the thread split",
"--resume", native_session_id.as_str(),
"--output-format", "stream-json", "--verbose",
"--permission-mode", "dontAsk",
];
}
The launch probe must verify that the returned session_id matches the reserved
UUID. If a qualified release instead requires capturing a generated ID, persist
that returned ID before issuing a successful checkpoint. Either strategy uses
only the private stored ID for subsequent resume. --yes is not a documented
general print-mode tool-approval flag; do not use it in this contract.
Claude restores its persisted context on resume, so Light supplies only the new instruction and current authoritative inputs, not the complete conversation. The documented headless examples explicitly demonstrate this across invocations. Native compaction can summarize older history; persistence does not promise unlimited verbatim recall, zero input-token usage, or a prompt-cache hit. Headless continuation.
Public lifecycle and private checkpoint
Reuse the existing typed directive, for example:
{
"thread": {
"runnerId": "personal-claude-runner",
"sessionRef": "019a0000-0000-7000-8000-000000000001",
"stageId": "implementation-phase-1",
"mode": "new",
"closeAfterTurn": false
}
}
This is a fragment inside coding, not a complete request. The workflow persists
sessionRef before dispatch. For resume or close, it supplies the last
successful receipt’s checkpoint as expectedCheckpoint.
| Workflow operation | Claude adapter behavior |
|---|---|
new | Require unused sessionRef; allocate native UUID; create persisted conversation and return validated checkpoint |
resume | Lock and validate expected checkpoint; launch --resume <stored-native-UUID>; commit next checkpoint only after result validation |
close | Without a model call, durably mark the Light session closed and prohibit further worker resume |
closeAfterTurn: true | Commit validated result first, then attempt closure; return actual READY or CLOSED state |
A result uses the existing codingThread receipt with sessionRef, checkpoint,
and state. Preserve the existing runner envelope locations. A native session ID
is not a public resume credential. Advertise supports_session_reuse: true and
workflow-coding-threads-v1 only after the complete Claude suite passes; until
then the adapter remains unqualified, rather than shipping without this feature.
The private record binds trusted workflow scope, host/user, runner, stage, role, adapter contract, model policy, immutable repository/base, materialization manifest, writable roots, and allowed tools. Candidate patch digests may advance within the same reviewer stage and are validated per turn; they must not force a new conversation. Keep native transcript storage private, outside disposable attempt directories, on the pinned runner. Do not copy OAuth credentials to implement history persistence or copy transcripts into prompt/artifact storage.
Native history must remain available across CLI and worker restarts. Reapply trusted launch policy every time; prior conversation permissions cannot widen the next lease. Pin the working-directory strategy and test reconstruction into new attempt paths. If the CLI cannot resume safely after path changes, provide a stable runner-managed path for that session or fail qualification. Missing history, runner loss, version changes, and incompatible bindings require an explicit workflow recovery decision, never an automatic new session.
Lock one operation per session and validate expectedCheckpoint before launch.
Preflight failures leave the last checkpoint intact. Durably mark IN_FLIGHT
immediately before starting a CLI that can advance history. Only validated
success returns to READY with a new checkpoint; cancellation or an uncertain
crash after that point blocks ordinary resume. Use the existing durable result
replay/ack path to recover delivered results without repeating a model turn.
Claude closure is a Light admission tombstone, not an assumed vendor archive API.
No running process is needed between turns, and no logout or prompt such as
“forget this session” is sent. Native transcripts may remain subject to retention;
CLOSED prevents use through the worker, not manual use by the local account
owner. If a vendor archive becomes required, qualify it separately. As with
Codex, closeAfterTurn must preserve a successful result if closure fails and
return READY plus closeError; the workflow checks for CLOSED and can submit
an explicit close. A failed dedicated close must not report success.
Multi-turn review and workspace refresh
The workflow starts distinct implementer and reviewer sessions at stage entry, resumes each during remediation rounds, and closes both at stage acceptance. The next stage gets new references. A fresh final review is an explicit workflow choice. Fresh reviewer context means independent from the implementer and other stages, not discarded between the reviewer’s own turns.
Every reviewer turn receives a reconstructed read-only tree containing the exact current candidate, requirements, finding ledger, and test evidence. Retained history helps track findings, but prior observations are not evidence that the new candidate passed. Inform Claude which candidate and paths changed. Carry forward the implementer’s latest accepted patch for implementation turns; do not assume arbitrary shell processes or untracked state persist with conversation.
For example: review kickoff -> memory-safety follow-up -> implementer remediation in its separate session -> reviewer resume against the updated candidate -> close. A request to “fix the linting errors” belongs in the implementer session under this role policy; resuming a reviewer does not grant write access.
Persistent Process Option: Structured Input, Not a PTY
The process-per-turn design preserves conversation already. Keeping one CLI
alive is a separate latency optimization worth prototyping, not a prerequisite
for multi-turn context. Claude documents --input-format stream-json alongside
streaming output. This is the candidate transport for multiple inputs over one
process; -p does not inherently require spawning a new process for every input.
CLI reference.
Proposed launch shape for a qualification probe:
<absolute-pinned-claude-executable>
-p --input-format stream-json --output-format stream-json --verbose
--permission-mode dontAsk
Keep stdin open and write one qualified user-message envelope per admitted Light
turn. Capture the exact input schema and terminal-event behavior from the pinned
release’s documented interfaces and fixtures before implementing the writer.
Do not guess that arbitrary text, y, or SDK-private control messages are valid
input. Qualification must demonstrate two prompts and two distinct completed
turns without EOF between them. A JSON input flag alone does not establish full
SDK control-protocol compatibility.
Use an explicit execution-mode contract:
| Mode | Conversation | Process lifetime | Initial status |
|---|---|---|---|
| Resume per attempt | Persisted native session | One leased attempt | Required baseline |
| Resident session | Same persisted native session | Multiple admitted turns | Optional measured optimization |
A resident process needs a runner-owned session host, not an untracked child left
behind when light-agent-worker exits. This adds a deployment component/lifetime
beyond the initial diagram. The host owns one worker/CLI pair per active native
session and attaches each admitted job through an authenticated local channel.
It never dispatches a queued prompt merely because the preceding turn ended.
Required invariants before promoting resident mode:
- Workflow
new,resume, andcloseremain the sole conversation boundary controls. Each turn gets a fresh lease, identity, checkpoint check, deadline, budget, and permission decision. Keep only one turn in flight per session. - Between turns the process has no authority to run tools or mutate a workspace. Prove a quiescence boundary, including background tasks, hooks, and subprocesses. If that cannot be enforced, terminate after each turn and retain baseline mode.
- Runner concurrency counts resident resources. Initially admit at most one resident process on a single-concurrency personal runner; evict an idle process before placing another session there. Eviction ends a process, not a workflow session, and is permitted only after a validated checkpoint.
- Reconstruct the exact next candidate and scratch state before admitting the next prompt. Qualify stable paths and stale file/tool-cache behavior; process memory must not cause review of the previous candidate. A session host cannot bypass the existing fresh-workspace contract for convenience.
- After the result and artifact validation, atomically commit the checkpoint and mark the host idle. Bound stdout/stderr draining even while waiting for a client. Never stop draining a child merely because an approval is pending.
- Explicit close stops the resident process tree and closes the Light checkpoint. Disconnect/revocation during an active turn cancels and fences it. An uncertain crash requires workflow recovery; do not replay the prompt on a replacement process. Clean idle eviction can resume the saved native session later.
- Idle timeout, resource pressure, and shutdown may evict clean idle processes. They cannot silently allocate new conversations. Persisted transcripts remain the recovery mechanism; an in-memory process registry is not durable state.
Benchmark startup latency, first-token latency, memory, candidate-refresh correctness, and process cleanup against resume-per-attempt. Promote resident mode only when savings justify the runner/session-host complexity. No SDK is needed merely to prototype public JSON input; use an SDK bridge only if required controls are unavailable through qualified public CLI interfaces.
Questions, Permissions, And Corrections To Wrapper Examples
Distinguish three events: normal assistant text asking for clarification, a structured user-question interaction, and permission to execute a particular tool. A follow-up user message can answer conversational text on the next turn. It cannot safely stand in for a permission decision on a blocked tool invocation. A qualified interaction channel must retain exact tool arguments, request ID, owner, lease, expiry, and one-time decision binding.
For explicit human questions, integrate a published structured interaction
surface into Light’s interaction records, or return a bounded needs-input result
and let the workflow submit the answer on a later turn. Do not claim that
--input-format stream-json alone supports permission callbacks or
AskUserQuestion responses. Until qualified, those features remain unavailable.
The supplied wrapper examples contain several claims we should not adopt:
- Permission bypass is not mandatory in headless mode. The documented
dontAskmode denies actions that would prompt; an SDK permission host or the CLI’s MCP permission tool can provide a structured approval path. A PTY and regex matching of terminal prompts are not a dependable substitute. Headless permissions. - Automatically writing
yremoves the approval boundary. Turning that text into a synthetic function call does not restore argument binding, authorization, replay protection, or reliable CLI control. - A request for permission is not a client-side function invocation. Approval authorizes the harness to execute a tool; an ordinary function result reports work the client executed. Generic clients must not confuse those operations.
- The quoted monthly-credit claim is superseded by Anthropic’s June 15 pause. Do not hard-code that pool or a claimed universal concurrency threshold. Our serial admission is an isolation/resource policy, not a vendor limit assertion. Subscription update.
- The JSON examples use guessed fields: documented headless output uses
session_idandresult. Achat.completionobject does not implement the Responses API or Anthropic Messages API. Capture and validate real schemas. - The Rust PTY sketch is not a compilable async implementation: a blocking
std::io::Readobject does not implement TokioAsyncRead; holding a standard mutex guard across an await can block progress, and discarding the child handle loses lifecycle control. Prefer native Tokio subprocess pipes and bounded tasks.
CLI-backed Model Providers In llm-gateway
Deferred follow-up notes: Codex CLI Provider and Claude Code CLI Provider. The Claude personal coding worker remains the implementation priority.
A text-only compatibility adapter is technically plausible for both Claude and Codex, but it is a separate proposal from the coding worker. Do not present it as a transparent replacement for model APIs or enable it as part of this design. The current LLM Gateway API says the gateway returns client-side tool calls rather than executing them and excludes CLI credential caches. It also mentions possible owner-scoped native connectors; that allowance and the explicit CLI exclusion need a separate architectural decision before adding these providers.
Repository inventory does not establish an existing compliant route:
crates/model-provider/src/claude_code.rsalready provides a CLI wrapper, but flattens conversation roles into text, runs with permission bypass in agent mode, buffers output, and does not preserve client tool-call semantics. It is not a qualified Responses/Messages provider.crates/model-provider/src/codex.rsis an HTTP implementation targeting a Codex backend endpoint, not a Codex CLI subprocess provider. Do not confuse its name with the qualified App Server worker or reuse subscription-token handling to implement a native CLI connector.
For a Codex connector, prefer the documented App Server lifecycle over a PTY or repeated terminal-output parsing. It exposes harness threads, turns, events, and approvals, not a drop-in public Responses model endpoint. Codex App Server. Native ChatGPT login and API-key authentication are distinct modes; native integration documentation alone does not establish permission for a shared subscription-backed inference proxy. Codex authentication.
If pursued, use two explicit experimental provider types (illustrative names
claude-cli-personal and codex-cli-personal) behind an owner-only local connector.
Keep the gateway free of native credential stores: it forwards an authenticated,
owner-bound request to the local supervised connector, which starts the official
harness. Never pool these routes across users, forward subscription tokens, or
silently fall back to paid API routes. Review vendor eligibility separately for
each connector; Anthropic’s third-party subscription restrictions remain relevant
even if the token never leaves the host.
Anthropic credential rules.
Start with a declared text-generation subset, only after proving all of these:
| Concern | Required contract |
|---|---|
| Tools and side effects | No repository access, shell, MCP, hooks, plugins, or other native tool execution; enforce outside prompts. Reject client tool requests until a faithful implementation exists. |
| Roles and instructions | Preserve supported instruction/message hierarchy; reject unsupported combinations rather than concatenate arbitrary roles into one prompt. |
| Model selection | Governed alias maps to an entitled native model; reject unknown aliases. Claude --model supports model selection, but model substitution must be explicit. |
| State | Stateless Messages requests must not gain hidden prior history. Responses continuation needs owner-bound response-to-native-session mapping, branch semantics, retention, and replay rules; reject previous_response_id until qualified. |
| Output | Generate unique response/message IDs per request, distinct from native session IDs; implement the actual requested envelope, terminal statuses, and errors. |
| Streaming | Map native events into the selected API’s ordering, content-block/item IDs, deltas, completion, usage, and cancellation; do not wrap arbitrary JSON Lines as SSE. |
| Unsupported features | Reject embeddings, images, reasoning controls, structured output, tools, storage, or other fields unless individually qualified; never silently drop them. |
| Quota and billing | Advisory native usage, bounded queue/backoff, explicit rate-limit outcomes; no invented invoice amounts or guaranteed subscription headroom. |
| Operations | Pinned binaries, owner-scoped capacity, deadlines, full process-tree cleanup, disconnect handling, output bounds, and sanitized audit evidence. |
For an agent that requires standard function calling, text-only support will not
satisfy the requirement. Either use ordinary API-backed model providers or design
and qualify a real external-tool bridge. Returning a fabricated permission tool
call or custom HTTP 403 requires_action is not standard Responses/Messages
compatibility. A workflow-native coding action is the appropriate interface when
the harness itself owns repository tools and needs Light approvals.
Recommendation: implement and qualify the Claude coding worker first; evaluate resident-process mode as a measured optimization. Treat owner-only CLI model connectors as a separate feasibility effort with explicit API-subset tests and vendor eligibility, not as a shortcut to generic full-featured model routing.
Implementation Plan And Qualification
Phases 0 through 3 have implementation records below. Phases 4 and 5 remain optional proposals.
| Phase | Deliverable | Exit evidence |
|---|---|---|
| 0: Native feasibility | Pin CLI artifact/platform; record flags, events, auth/config behavior, persistence, and eligibility | Native login and exact-ID resume across processes work without unwanted startup execution |
| 1: Deterministic adapter | claude_code.rs in the worker, bounded parser/process supervisor, private session checkpoints, Claude capability and launch contract | Synthetic new/resume/close, stale checkpoint, malformed-output, cancellation, and policy-selected permission-mode tests pass |
| 2: Coding integration | Role selection, canonical patches, independent review, native-auth classification | Real multi-turn edit/test/review across separate worker processes; refreshed candidate, closure, cross-owner and profile-confusion checks pass |
| 3: Local distribution | Dedicated personal runner configuration and setup documentation in both local distributions | Portal-to-runner smoke on portal-config-loc/all-in-lt and light-portal-install; no gateway model traffic |
| 4: Optional interaction | Qualified SDK bridge or public CLI permission channel | Approval replay/expiry/cancel races pass before interactive capability enablement |
| 5: Optional resident process | Runner-owned session host and structured multi-input CLI probe | Cross-turn leases, candidate refresh, idle eviction, crash recovery, cleanup, and latency comparison pass |
Reuse the existing qualification framework without declaring Claude qualified by
copying the Codex evidence. Add a Claude manifest under
contracts/coding-adapters/, proposed gate
scripts/run-claude-personal-gates.sh, and explicit live-smoke opt-in. Validate
all 13 dimensions: protocol lifecycle, approval mediation, streaming, usage,
cancellation, resumability, canonical patch, review isolation, authentication,
workspace isolation, panic containment, dependencies, and licensing. For an
unsupported feature, evidence must prove safe rejection and truthful capability
advertisement; confirm that promotion policy accepts that constrained profile.
Required negative cases include:
- Wrong user/host, concurrent workspace access, stale lease, capability or binary digest mismatch, expired login, API credentials, and gateway injection.
- Repository startup hooks, user plugins, managed-policy surprises, hidden MCP configuration, unauthorized permission-mode overrides, and attempts to mutate protected files or reviewer input through Bash.
- Truncated/oversized JSON, duplicate terminal results, output flood, stderr secrets, missing usage, denial followed by a success claim, and exit without a terminal result.
- Cancellation during startup, tool execution, output backpressure, and artifact collection; orphan processes and post-cancellation writes.
- Implementer transcript leakage into review, malformed findings, native session collision, name ambiguity, ambient latest-session selection, stale checkpoints, concurrent resume, duplicate new, resume after close, and cross-user/stage resume.
- A conversation-only marker supplied in turn one is recalled in turn two through separate worker/CLI processes without replaying history; a third review turn evaluates an updated candidate while preserving prior findings. Verify native ID continuity, persistence under qualified configuration flags, missing-history failure, explicit close without a model call, and unchanged reviewer authority.
Live tests must report the pinned version and actual native authentication
class without personal account details. Record quota skips as unqualified live
cases, not passing tests. Verify absence of model traffic through llm-gateway
and absence of native credentials in Light events/artifacts. Do not claim this
proves same-user credential isolation.
Phase 0 Verification Record
Implemented September 12, 2026:
prototypes/claude-code-v1/phase0.py: bounded native subprocess probe using only Python’s standard library; no SDK or production Rust dependency.scripts/run-claude-personal-phase0-gates.sh: offline failure tests and explicit opt-in native smoke. An omitted live run prints NOT RUN, never a live pass.contracts/claude-code/v2.1.269/phase0.json: exact binary/platform pin, required public flags, limits, tested model mapping, and unresolved distribution decision.contracts/claude-code/v2.1.269/phase0-live-evidence.json: sanitized live evidence. The neighboring fixtures contain explicitly labeled synthetic parser input.
Live execution passed four turns on Linux x86_64 with Claude Code 2.1.269,
SHA-256 25e44883f54419569a3d739f38cbbdaebe83b09895da0f343e1b003710a4775b.
Authentication status was first-party native claude.ai; only the normalized
personal-subscription class is recorded. Credentials were not copied or logged.
Both agent-policy and claude-cli modes completed new/resume across separate
processes, preserving an exact UUID and recalling a random conversation marker
not repeated in the second prompt. Every turn requested sonnet and observed
claude-sonnet-5, overriding a different disposable-project model default.
A harmless project SessionStart hook was suppressed by --safe-mode and ran on
each native-mode invocation. Native mode inherited the project’s dontAsk
permission mode without a permission-mode flag. Managed mode used explicit
dontAsk, an empty tool surface, and observed no loaded MCP servers or plugins.
This proves the tested configuration combination, not suppression of every
possible managed hook or behavior of every installed customization.
The implementation includes bounded frame/output/stderr handling and process-group termination. Offline tests exercise malformed/duplicate/error output, identity mismatch, wrong authentication, conflicting environment, binary mismatch, output floods, and a child holding a pipe past the deadline. Phase 0 does not yet qualify production runner isolation, native editing, permission bypass, tool denials during execution, workflow model admission, checkpoints/close, or resident processes. Those remain later-phase work. Eligibility for supported distribution is explicitly unresolved, so technical feasibility does not promote the adapter.
Run from the repository root:
./scripts/run-claude-personal-phase0-gates.sh
LIGHT_RUN_CLAUDE_PERSONAL_SMOKE=1 \
LIGHT_CLAUDE_EXECUTABLE=/absolute/path/to/claude \
LIGHT_CLAUDE_NATIVE_MODEL=sonnet \
LIGHT_CLAUDE_PHASE0_REPORT=/tmp/claude-phase0-report.json \
./scripts/run-claude-personal-phase0-gates.sh
The live command consumes native subscription usage and leaves its probe conversations in the owner’s native history. It does not change user settings or remove existing sessions. Exact binary drift fails before a model turn.
Phase 1 Implementation Record
The candidate Rust adapter is implemented in
apps/light-agent-worker/src/claude_code.rs behind claude-prototype (also built
in unit tests). The default worker dispatch is unchanged, and even a build with
the candidate feature advertises only the existing Codex capabilities. The
Phase 5 optional-adapter guard permits this isolated candidate module and tests
that it does not become a production selection.
Implemented candidate contracts and lifecycle:
- Strict
ClaudeTurn/LaunchPolicytypes with distinct native model, permission source/mode, available tools, and pre-approved tool rules. Runner paths and trusted scope are supplied separately by the future admission integration. - Exact binary/version and native-subscription preflight; direct argv launch, controlled environment, native inheritance or explicit managed configuration.
- Bounded JSON Lines and normalized progress/tool metadata, sanitized advisory usage, terminal validation, stdout/stderr bounds, deadline, cancellation, event backpressure failure, and process-group cleanup.
- Shared private checkpoint locking and atomic writes for new/resume/close, model continuity, policy/scope binding, stale checkpoint rejection, and uncertain-operation fencing. Close is a Light tombstone, not vendor deletion.
- A native result is a proposal. Only trusted artifact acceptance commits its
checkpoint; unaccepted or failed work remains
IN_FLIGHT. Permission-denied work cannot be accepted as success. Phase 2 supplies canonical validation.
The launch contract is pinned in
contracts/claude-code/v2.1.269/phase1-launch.json and contributes to checkpoint
identity. Run scripts/run-claude-personal-phase1-gates.sh for offline Phase 0
checks, Rust adapter/shared-runtime regressions, candidate/default build checks,
production capability isolation, and the documentation build. Synthetic process
fixtures do not consume subscription usage. The existing live Phase 0 opt-in is
available separately through the composed gate.
Phase 1 did not enable public workflow dispatch. Phase 2 adds the integration below; the Phase 1 protocol fixture remains useful as a smaller regression gate.
Phase 2 Integration Record
A Claude coding Agent uses the existing light-agent executable with a Claude
coding profile. The workflow selects that Agent definition; it does not call a
worker directly. Agent admission validates the pinned claude-code-v1 contract,
personal authentication, explicit thread directive, server-owned claudePolicy,
and optional coding.nativeModel, then schedules the normal runner execution.
Each Agent process has one configured coding adapter. Deploy separate Codex and
Claude Agent definitions/processes when both are needed, with separate worker
pools. Both implement and review roles can use the same Claude pool.
The dedicated light-claude-worker executable is built with the historical
claude-prototype feature. It advertises only coding.claude-code-v1; the default
light-agent-worker continues to advertise only Codex. Configure the Claude
runner with claudeHome, claudeExecutable, the dedicated worker executable,
and its exact capability digest. Claude configuration cannot share a Codex or
enterprise-broker pool. The actual runner validates admission before spawning,
journals lease-fenced progress/artifact/terminal events, and verifies terminal
authentication and artifact evidence. Client input cannot override Claude policy.
The pinned contract is contracts/claude-code/v2.1.269/phase2-launch.json and its
technical qualification descriptor is phase2-qualification.json. The explicit
LocalQualified status requires the 12 technical dimensions and exact binary,
launch, capability, and evidence digests. It does not satisfy the existing
13-dimension production Qualified promotion gate: subscription integration and
distribution eligibility remain unresolved. Local admission is deliberately
limited to this exact contract, not a generic qualification bypass.
The coding implementation:
- Copies and hashes the immutable repository bundle before Git consumes it (128 MiB bound), creates a fresh checkout for every turn, and keeps Git metadata outside CLI-writable roots. Current manifests must have no packages or instructions; unsupported materialization is rejected.
- Restores only the accepted canonical implementation patch on resume. Remediation binds its prior artifact digest. Review uses a separate conversation and the newly admitted candidate. Structured review output must match the existing review schema, review ID, and artifact digest.
- Uses a required Linux bubblewrap namespace with a sparse filesystem. The
reviewer sees its candidate and Git metadata read-only, its own native state,
and scratch space. Other user repositories, implementation transcripts, and
Light checkpoints are not mounted. Implementers receive only their admitted
writable roots. System
/usrand/etcare read-only; networking is available. This is not isolation against an unsandboxed process running as the same user. - Keeps accepted Light checkpoints in owner-only
.light-claude-checkpoints-<native-home-path-sha256>and native conversation state in separate owner-only.light-claude-native-<scope-session-role-hash>sibling directories. Native credentials are mounted read-only from the owner’s existing installation, never copied into checkpoints or reports. - Supports
agent-policytool rules andclaude-cliconfiguration delegation, including explicitbypassPermissions. Native mode projects installed settings, hooks, plugins, skills, commands, agents, and the user config file read-only at their original paths. External customization dependencies outside the namespace need separate qualification. Provider/API credential overrides are rejected. The current native-auth profile requires file-backed login; keychain-only login and credential refresh requiring writes are not qualified. Refresh login outside the worker when necessary; there is no paid-provider fallback. - Preserves workflow-owned new/resume/close semantics and accepted-checkpoint fencing. Failed, cancelled, denied, malformed, or unvalidated work cannot advance acceptance. Cancellation kills namespace descendants. Close records a tombstone without pretending a fresh native authentication occurred.
- Reports native authentication as
native-claude-store; usage is advisory. It never manufactures test-command exit evidence from model prose.
Run the deterministic integration gate:
./scripts/run-claude-personal-phase2-gates.sh
Run the opt-in native proof through the real worker-process runner:
python3 scripts/run-claude-coding-smoke.py \
--claude /absolute/path/to/pinned/claude \
--native-home /absolute/path/to/existing/.claude \
--worker target/debug/light-claude-worker \
--runner target/debug/examples/claude-dispatch \
--permission-source claude-cli \
--report /tmp/claude-phase2-live.json
The driver exercises implement-new, review-new, implement-resume, review-resume-close, and implement-close. It checks unpredictable conversation markers in both resumed sessions, reconstructs the final accepted patch, and runs an independent fixed Python assertion. Reports omit prompts, credentials, native IDs, and transcripts. Failed attempts retain uncertain state rather than silently retrying an accepted turn.
The deterministic gate passed 152 Rust tests and 10 Python tests. Five existing
tests were skipped: four require a PostgreSQL test database and one is an unrelated
workspace namespace test. Builds, capability isolation, formatting, and mdBook
also passed. Sanitized native runner evidence is recorded in
contracts/claude-code/v2.1.269/phase2-live-evidence.json.
This completes local Phase 2 integration. The runner proof uses an authenticated local lease fixture and the real runner/worker transport; it does not claim a live Portal/controller deployment. Phase 3 below adds published Agent deployment and enrolled Controller dispatch. Production distribution eligibility remains an independent open gate.
Phase 3 Local Distribution Record
Both local distributions now contain light-agent-claude-personal and the same
host-native setup helper. portal-config-loc/all-in-lt/docker-compose.yml includes
the Claude service directly, on loopback port 8090. The installer keeps it under
an automatically selected profile after enrollment. Native runner health is on
port 9445; native login and private conversation files remain on the host.
The Portal publisher now validates the exact Claude local contract and
claudePolicy, preserving Codex’s production qualification checks. Publish the
profile through agent-policy-authoring / codingProfile; the runtime snapshot is
still publisher-owned. The workflow selects the Claude Agent definition, then
passes typed coding.nativeModel and explicit thread directives to that Agent.
Shared workspace support for both native adapters is described in
Shared native coding sessions. It uses explicit
workspace bindings and task IDs; prompts do not discover arbitrary host paths.
light-fabric/scripts/build-claude-personal-local.sh builds the dedicated worker,
host runner, static Agent, and local Rust Agent image. Build the Java publishers
through their normal build/release pipeline after rebuilding their light-portal
dependency. Use a clean Maven package for shaded publisher jars so stale dependency
classes cannot survive inside a reused fat jar. Select those images through
PORTAL_HYBRID_COMMAND_IMAGE and PORTAL_HYBRID_QUERY_IMAGE. The normal Agent Dockerfile now includes
the pinned contract resources needed by the shared Rust crate.
setup.py validates supplied service/runner identities, checks the native CLI
pin and bubblewrap availability, installs content-addressed binaries, and writes
an owner-only runtime directory and enabled systemd user unit. It preserves
existing journals and conversation state. Native credentials are not copied.
The service JWT is bound separately into the non-root Agent container; its parent
directory is owner-only. Expiring local issuer credentials must be renewed.
Normal portal-config-loc/scripts/deploy-local.sh lt regenerates a merged
Controller admission file from the installed Codex and Claude units, deduplicates
shared origins, rejects conflicting runner identities before replacing admission,
restarts both user services to load their installed configurations, and checks both
personal Agent containers. The checked-in
Compose selections include the local Claude Agent image and use the standard
Java publisher image settings; there are no Claude-specific Java image overrides. The installer
includes its tracked controller overlay and Claude profile after setup and checks
image capabilities before starting enrolled services.
Run scripts/run-claude-personal-phase3-gates.sh for configuration parity,
admission merge regressions, Java profile validation, source-pin parity, shell
checks, and mdBook. The live opt-in is scripts/run-claude-deployment-smoke.py;
install its pinned dependencies from scripts/claude-personal-test-requirements.txt.
See each distribution’s light-workflow-runner-claude-personal/README.md for build,
enrollment, publication, normal restart, and smoke commands.
The deployment smoke submits public Agent WebSocket coding requests, observes Controller scheduling and durable fenced receipts, resumes separate implementer and reviewer sessions, explicitly closes both, and independently tests the final canonical patch. It checks the local gateway audit count before and after; subscription model traffic must not add audit rows. Reports omit native session IDs, credentials, and transcripts.
The two distribution service layouts were qualified against the same local
Portal/Controller and host-native enrollment. The installer test replaces only
the Claude Agent service using the installer’s checked-in mounts and environment;
it is not a claim of a clean-room installer/database bootstrap. The final state is
restored to portal-config-loc/all-in-lt using its normal deployment command.
Sanitized evidence is under contracts/claude-code/v2.1.269/phase3-*-evidence.json.
Decisions To Close Before Implementation
The initial scope is Linux, trusted single-user execution, workflow-controlled new/resume/close sessions, workflow-selected native models with an agent default, and policy-selected constrained or trusted automation without interactive approval. Phase 0 pins a technically feasible CLI/configuration candidate. Before production promotion, complete deployment qualification and the eligibility determination for distribution. Until then, retain explicit local qualification. Session continuity is mandatory and does not itself require an SDK bridge. Revisit the bridge if interactive approvals or other required capabilities exceed the qualified public CLI surface.
This gives Claude a small independent adapter boundary while preserving the existing Codex worker and Light’s durable orchestration. It does not introduce a second worker policy engine or treat personal subscriptions as pooled providers.
Final Review Follow-up (Phases 0–3)
Phases 4 and 5 are deferred by owner decision. The final review identified three Phase 0–3 defects: ignored build output entered canonical patches, installer startup did not replace active old runners, and container health was reported as successful deployment without verifying host-runner readiness.
Canonical collection now honors ignore rules for new files while retaining tracked changes and independent protected-path checks. A regression compiles Python and writes a 2 MiB ignored build output, then validates only the source patch. Both distributions share a tested runner lifecycle helper: preflight all configured runners before restarting any, restart after full stack startup, and wait for healthy Controller connectivity with matching executable/configuration and backend compatibility. Missing configured units and stale/expired credentials fail explicitly. Entirely unenrolled optional runners are not required.
Operator guides describe renewal, pin upgrades, bounded restart drain, aggregate storage monitoring and conservative retention. No native transcripts or credentials are printed. CLOSED checkpoint pruning does not imply native transcript deletion; automatic native-state garbage collection remains unqualified. Empty native validation evidence still requires independent fixed build/test actions. The expanded native smoke uses a fixed unittest suite and a different candidate on reviewer resume, while checking that generated output does not enter the artifact.
Remaining qualification boundaries are explicit: full workflow-engine service-mode execution/restart, clean installer/database bootstrap, representative project build environments, release artifact publication and production eligibility. They must not be inferred from the local Agent WebSocket smoke or synthetic lifecycle tests. These are rollout prerequisites, not reasons to implement interactive approvals or a resident native process.
The final review follow-up checked the deployed Workflow store under the exact
workflow_ops, operational_meta search path used by light-workflow/src/main.rs.
agent_definition_t, agent_policy_snapshot_t, and agent_job_t are unresolved
there; the latter two exist in agent_ops, while the authoring catalog exists in
the Config Server database. TaskExecutor::execute_agent_call and
reconcile_agent_job still use unqualified catalog/job queries. A full service-mode
workflow test is blocked by that shared-store integration, rather than a missing
Claude session feature. Do not bypass the store boundary by granting the Workflow
runtime broad access to Agent tables or switching it back to Config Server.
Implement and qualify an explicit projection/job-bridge contract separately.
Read-only reproduction on the local stack:
docker exec postgres psql -U postgres -d operations -Atc "BEGIN READ ONLY; SET LOCAL ROLE operations_workflow_runtime; SET LOCAL search_path TO workflow_ops, operational_meta; SELECT name,to_regclass(name) FROM unnest(ARRAY['agent_definition_t','agent_policy_snapshot_t','agent_job_t','wf_definition_t']) name; ROLLBACK;"
The first three names returned null in this qualification; wf_definition_t
resolved. This is not a successful workflow-engine execution test.
Shared workspace extension
Claude and Codex can use the same registered task worktrees with separate native conversations. See Shared native coding sessions for permissions, workflow handoff, Chat controls, publication and verification. This extension does not require the deferred resident-process or interactive-approval phases.
Shared native coding sessions
Codex personal and Claude personal use one host-owned workspace store. Each has
its own Agent deployment, runner enrollment, native login, and conversation state.
The workspace registration grants both Agent service IDs. Each runner has its own
owner-only RunnerWorkspaceConfig pointing to the same store, and its own
published codingProfile.workspaceBindings selecting that runner. Matching
workspace names on different stores do not share files.
flowchart TD WF[light-workflow: selects Agent and controls stages] --> CA[Codex light-agent] WF --> CL[Claude light-agent] Chat[Portal Chat] --> CA Chat --> CL CA --> Controller[Controller scheduling and fenced receipts] CL --> Controller Controller --> CR[Codex personal runner and worker] Controller --> LR[Claude personal runner and worker] CR --> CT[Codex app-server dynamic task_workspace tool] LR --> LT[Claude CLI fixed MCP task_workspace bridge] CT --> Store[One workspace manager and task worktrees] LT --> Store
The manager holds the task lock across a native turn. Implementation tools may edit files with digest preconditions. Review and inspection tools cannot edit. The CLI namespace contains private native state and the fixed tool transport; it does not contain the workspace store or task checkouts. A completed tool session persists the exact returned checkpoint for subsequent review admission. This is a content checkpoint, not an approval or publication authorization.
Claude uses its pinned CLI and native subscription login, with a fixed stdio MCP
bridge over a private Unix socket. Only task_workspace is exposed. The native
CLI’s --tools "", --strict-mcp-config, and explicit MCP configuration keep
repository operations within the manager’s authority. See the official
CLI reference. Workspace bindings grant
this managed tool surface; bundle-mode inheritance of installed native tools,
hooks and MCP servers does not extend to shared tasks. This matches the Codex
shared-workspace boundary. Native permissions cannot override a reviewer’s
read-only manager session. Tests, commands, freeze/approval, Git operations and
publication remain separate fixed manager/workflow operations.
Task and conversation identities
workspaceId + task.taskId identify shared files. A thread.sessionRef
identifies one private native conversation belonging to one Agent, runner, task,
intent and stage. Implementer and reviewer use different conversation IDs; they
never share native transcripts. Both adapters support:
thread.mode: new: caller chooses a fresh UUID and stage ID.thread.mode: resume: caller supplies the previous conversation checkpoint.thread.mode: close, orcloseAfterTurn: true: caller ends reuse explicitly.nativeModel: optional native selection; Claude validates its published model map. Keep the model selection consistent when resuming a Codex workspace turn.
The native UUID checkpoint (codingThread.checkpoint) is different from the
file checkpoint (workspace.checkpointDigest, a SHA-256 digest). Neither replaces
the other. Changed policy, role, task or membership invalidates native reuse.
Uncertain turns stay fenced; a new request ID does not authorize automatic replay.
Legacy Chat requests without thread execute one fresh conversation and close it.
Workflow workspace jobs require explicit thread control.
Implementation and review sequence
- Select either Agent. Send
intent: implementand a new task/conversation. - Retain the returned task ID, workspace checkpoint and implementation conversation checkpoint.
- Select the other Agent for review. Send the same task ID,
intent: review, the exact workspace checkpoint, and a new review conversation. - For remediation, resume the implementation conversation against the current workspace checkpoint. Retain the resulting new checkpoint.
- Resume the reviewer with that new workspace checkpoint. The reviewer must reread the current task; native history is not evidence of current contents.
- Close both conversations when the caller finishes the review loop.
Review requires an existing task and expected workspace checkpoint. A stale checkpoint is rejected before native execution. Read-only review does not itself produce trusted acceptance evidence for commit, push or deployment. Existing freeze/approval/delivery contracts continue to own those transitions.
Chat exposes task selection, expected workspace checkpoint, new/resume/close, close-after-turn and native model. Successful results carry both checkpoints; Chat selects the existing task and resumed conversation for the next turn. Changing intent starts a separate conversation. To return to another conversation, use the ID and checkpoint displayed with its prior result. Reconnecting or switching Agents does not implicitly transfer conversation authority.
Workflow request shape
The Agent’s existing durable job dispatcher accepts the same typed input:
{
"profile": "coding",
"clientMessageId": "review-turn-2",
"text": "Review the updated implementation.",
"workspace": {
"schemaVersion": 1,
"requestId": "review-turn-2",
"workspaceId": "personal",
"expectedMembershipRevision": "sha256:<published membership digest>",
"task": {"kind": "existing", "taskId": "<implementation task ID>"},
"intent": "review",
"expectedCheckpointDigest": "sha256:<implementation checkpoint>",
"instruction": "Review the updated implementation.",
"thread": {
"runnerId": "personal-claude-runner",
"sessionRef": "<review conversation UUID>",
"stageId": "review",
"mode": "resume",
"expectedCheckpoint": "<previous review conversation checkpoint UUID>",
"closeAfterTurn": false
},
"nativeModel": "sonnet"
}
}
Place this payload in a service-mode agent call’s with.input; select the Agent
with with.agent. The dispatcher derives the workspace subject
workflow-agent:<agent-definition-UUID> from immutable server authority. Grant
that exact subject and Agent service ID in both the published and runner-local
binding for workflow use. A human Chat grant alone does not grant workflow use.
The browser or model cannot provide a host path, binding, store, or subject.
The local Workflow service’s existing cross-database Agent job/catalog bridge is
still an integration prerequisite: the workflow_ops search path does not expose
the required Agent job and catalog tables. This change adds workspace handling on
the Agent side; it does not claim that service-mode Workflow invocation already
works across those separate deployed databases. See the recorded prerequisite in
contracts/claude-code/v2.1.269/review-fixes-workflow-prerequisites.json.
Deployment and repeatable checks
Build updated workers, runner, Agent, Portal policy publishers and Portal UI.
Enable workspaceConfig on the Claude runner using the setup helper’s
--workspace-config option; subsequent setup preserves that configuration.
Use the same store/registration as Codex and publish each generated profile through
agent-policy-authoring.codingProfile, then publish a new immutable Agent snapshot.
Do not edit a compiled snapshot. The existing normal deploy-local.sh lt restart
regenerates admission and verifies installed runner identities.
Run Rust workspace/worker tests, Java ClaudeCodingProfileTest, and the Portal
workspace/Chat tests. scripts/run-shared-coding-smoke.py exercises Codex implement
→ Claude review → both edit/review resumes → conversational follow-ups → both closes on a disposable task with an
independent file-content assertion. scripts/run-claude-workspace-smoke.py checks
Claude edits across three disposable repositories. These consume native
subscriptions and exercise the real workers, not the Workflow engine or browser.
For a deployed read-only check, install scripts/claude-personal-test-requirements.txt
in a virtual environment and run:
python scripts/run-shared-workspace-deployment-smoke.py \
--token-file /private/path/owner.jwt --workspace personal \
--task existing-task-id --report /tmp/shared-workspace-deployment.json
The credential must identify an owner granted access to both Agents. The script uses an existing task, opens and closes one conversation per adapter, and checks both returned the same file checkpoint. Run on a quiet task; concurrent owner edits invalidate the comparison. URLs are configurable. It does not submit an implementation turn or execute a Workflow process.
The 2026-09-12 qualification passed the six-turn native implementation/review loop,
both deployed read-only turns through Agent/Controller/runner, both namespace
isolation tests, and a normal deploy-local.sh lt restart. Sanitized receipts are
in contracts/claude-code/v2.1.269/shared-workspace-qualification.json.
Completion, sandbox and retained state
A confirmed successful native turn may be conversational: it need not call a workspace tool or return nonempty text. Empty output gets a neutral completion message; public text is UTF-8-truncated at 64 KiB. These presentation decisions must not leave a completed conversation IN_FLIGHT. Protocol/transport errors, missing terminal results, cancellation and uncertain native execution still fail closed and require explicit recovery rather than automatic replay.
Codex uses native read-only sandboxing (including native tool network denial), while task_workspace writes remain authorized by the external workspace manager. The outer namespace exposes no checkout/store. Fixed native configuration and credentials are mounted read-only; only private native session state persists. Claude prepares its fixed MCP bridge, socket and credential mounts before opening a writable workspace tool session.
Successful legacy requests without explicit thread control delete their native
home after recording the CLOSED receipt. Explicit conversations and failed or
uncertain standalone histories are retained. personal-runner-lifecycle.py storage
reports shared workspace native directory count, bytes and checkpoint state counts
for either adapter without printing transcripts. Both runners report the same
store-wide totals; do not add those totals together. Automated deletion of retained
explicit/uncertain homes is not implemented. Operators must reconcile records and
drain runners before cleanup; age alone is not authority to delete uncertain state.
Conversation locks use an explicit-unlock guard on every exit path, including rejected opens and closed-record pruning. Acquisition tolerates a 100 ms inherited file-descriptor window and still rejects a live owner. Codex acquires the workspace tool session before marking the conversation IN_FLIGHT or issuing thread/start or thread/resume. Task contention therefore leaves the existing conversation checkpoint available for retry after the competing task turn ends.
Workflow-controlled coding threads
For native permission inheritance and coding.nativeModel, see
Personal Codex permissions and models.
The workflow chooses conversation boundaries. Keep an implementer thread and a separate reviewer thread through one stage’s implementation/review/remediation cycle. Close them when the stage is accepted, then allocate new session references for the next stage. An optional final reviewer can start fresh to reassess the whole accepted change without earlier review assumptions.
This replaces the requirement to start a fresh harness thread for every job.
Every job still has a new Light turn, execution lease, deadline, and result. The
worker process and Codex App Server process may exit between jobs: native
thread/resume restores the persisted Codex conversation in a new process.
No /clear command is sent. Within each job, Codex keeps context across its model
and tool calls.
Implemented scope
The native personal-subscription codex-app-server-v1 worker supports this
contract. claude-code-v1 remains an unimplemented adapter; its eventual
implementation must obey the same lifecycle and keep its review conversation
separate from Codex’s implementation conversation. This change does not claim
that Claude execution is available.
Enterprise workers reject persisted thread directives: their per-attempt Codex
homes contain ephemeral broker configuration and credentials. Persisted enterprise
thread storage requires its own isolated storage/credential design and qualification.
Existing requests without coding.thread retain the one-job ephemeral behavior.
Job contract
A workflow service-mode Agent job supplies the same typed input as an Agent coding
request: profile: coding, text, and the complete coding object. The existing
workflow expression resolver can populate the thread directive and its checkpoint
from workflow context. Do not place the directive inside the prompt.
Example directive inside coding:
{
"thread": {
"runnerId": "personal-codex-runner",
"sessionRef": "019a0000-0000-7000-8000-000000000001",
"stageId": "implementation-phase-1",
"mode": "new",
"closeAfterTurn": false
}
}
This is a fragment, not a complete coding request. The immutable repository, base revision, role, tools, bounds, and review/remediation input are still required. Use different UUIDs for the implementer and reviewer. The workflow persists them before dispatch and carries them through retries; it must not generate a new UUID on every retry of the same logical job.
| Operation | Required state | Behavior |
|---|---|---|
new | Unused sessionRef; omit expectedCheckpoint | Create a persistent native thread |
resume | Last successful receipt’s expectedCheckpoint | Resume that exact native thread |
close | Last successful receipt’s expectedCheckpoint | Archive the thread and close its checkpoint without a model turn |
closeAfterTurn: true | A new or resume turn | Attempt to close after its successful validated result |
A successful worker result includes a codingThread receipt:
{
"sessionRef": "019a0000-0000-7000-8000-000000000001",
"checkpoint": "019a0000-0000-7000-8000-000000000002",
"state": "READY"
}
Set expectedCheckpoint to that checkpoint on the next operation. A close
returns state: CLOSED. Native threadId is diagnostic evidence, not the public
resume credential; callers cannot choose an arbitrary existing Codex thread.
closeAfterTurn is an attempt, not a guarantee
closeAfterTurn rides on a turn whose real result is the patch or the review. The
worker commits that result to the checkpoint before it tries to archive the
thread, and the archive is bounded by whatever execution time remains after
reserving enough to deliver the result. When the lease is nearly spent the close is
skipped outright rather than risking the turn being killed at its deadline with an
undelivered result. If the archive fails, never answers, or is skipped, the turn
still succeeds and returns its implementation or review result, but the receipt
reports the thread as it can be proven to be — still open, with the reason
attached:
{
"sessionRef": "019a0000-0000-7000-8000-000000000001",
"checkpoint": "019a0000-0000-7000-8000-000000000002",
"state": "READY",
"closeError": "thread/archive did not answer within 10s"
}
A workflow that requires closure must check for state: CLOSED rather than
assuming a successful turn closed the thread. A READY receipt with a
closeError is a normal outcome: keep the result, and close the still-resumable
thread with an explicit mode: close turn using the checkpoint just returned.
state describes the conversation, never the cleanup — that is why a failed close
does not invent a distinct thread state.
A dedicated mode: close turn is the opposite case: archiving is the entire point
of the turn, so a failed or unanswered archive fails the turn, and such a turn only
ever reports CLOSED on success.
Implementation results are wrapped by the runner, so the receipt is under the
implementation result’s worker.codingThread; review and close receipts are
under codingThread. The Agent’s durable terminal result also wraps the execution
result. Inspect that envelope and map the receipt into the workflow context; do
not confuse a scheduling request ID with a checkpoint.
The workflow-to-Agent bridge now dispatches typed coding jobs from
agent_job_t through the same digest-bound execution scheduler used by the
WebSocket coding path. Invalid coding jobs fail explicitly. Ordinary service-mode
Agent jobs retain their existing behavior.
Scope, ordering, and workspace continuity
Agent supplies the trusted thread scope from persisted Host, service, policy, data boundary, and workflow process identity. For direct authenticated Agent coding requests, scope uses the Agent session and its principal instead. The request cannot supply a foreign workflow scope. Storage additionally binds stage, role profile, model alias, adapter contract, runner, immutable repository/base, materialization manifest, tools, and writable roots.
Session jobs pin runnerId and require workflow-coding-threads-v1. The native
runner’s admission and live registration advertise the same feature only when
the configured worker capability digest matches the thread-capable worker.
Upgrading the runner alone does not advertise reuse for an older worker. They must
also carry the exact rebuilt worker capability and binary digests.
Only one operation may hold a session’s lock. An old checkpoint, changed scope,
changed role, changed contract/base, duplicate new, or a closed session is
rejected before a model turn. Missing native history or an interrupted execution
never silently falls back to a fresh conversation. The workflow explicitly starts
a replacement session from accepted artifacts if recovery is needed.
Native history resides in the owner-only Codex home. Worker checkpoint metadata
and the latest validated implementer patch reside in its private
light-worker-threads directory. This directory is not a shared tenant store.
Do not edit or copy checkpoint JSON to manufacture a resume authorization.
The worker reconstructs a fresh repository for every job:
- Implementer: same immutable base plus the previous validated checkpoint patch. Remediation findings must refer to that patch’s digest.
- Reviewer: same immutable base plus the newly supplied candidate patch. Earlier reviewer memory is retained, but the candidate and evidence are refreshed. Only separate build scratch is writable; the candidate must remain unchanged.
Untracked caches and arbitrary shell state do not carry forward. Each prompt
identifies the current repository and warns that earlier paths/tool observations
may be stale. A stage or base change uses a new sessionRef and new immutable input.
The present implementation is a single-repository checkpoint, not the future
multi-repository WorkspaceSet implementation.
Storage is marked IN_FLIGHT immediately before the first native request that can
create or advance the durable thread, and not earlier. Opening a checkpoint locks
and loads it without writing, so a preflight rejection — a repository that will not
stage, a changed binding, remediation findings that do not match the checkpoint’s
patch — releases the thread untouched and the workflow may simply retry the same
resume. Only once the native operation is claimed is an interruption
indistinguishable from a lost one. Only successful validation advances the
checkpoint, so an interrupted or uncertain operation requires workflow recovery
rather than an automatic potentially duplicate turn.
Execution-result delivery still uses the runner’s existing durable replay/ack path.
A close archives native history; it is not secure deletion or a retention policy.
Workflow policy
sequenceDiagram
participant W as Workflow
participant I as Implementer thread
participant R as Reviewer thread
W->>I: new, stage A
I-->>W: Patch and checkpoint I1
W->>R: new, stage A, candidate
R-->>W: Findings and checkpoint R1
W->>I: resume I1, findings
I-->>W: Updated patch and checkpoint I2
W->>R: resume R1, updated candidate and evidence
R-->>W: Accepted review and checkpoint R2
W->>I: close I2
W->>R: close R2
W->>I: new sessionRef, stage B
Keep requirements, plans, accepted patches, review findings, and actual test results in durable artifacts independently of conversations. These are the recovery source and publication evidence. The reviewer never receives the implementer’s private conversation. Retaining the reviewer’s own history helps track finding closure; policy can add a fresh final review when warranted.
Upgrade and verification
Rebuild light-agent, light-agent-worker, and light-workflow-runner. Regenerate
the worker template and runner admission from the installed binaries, then
regenerate and publish the coding profile with the new capability/template/image
and qualification digests. Recreate Controller with that admission and restart
the runner. Version-string equality alone does not validate these artifacts.
Close active stages before upgrading: an adapter contract change intentionally
prevents resuming an old checkpoint. Start replacement stages from accepted
artifacts when a clean close was impossible.
Focused local tests:
cargo test -p coding-agent-runtime -p light-agent-worker -p light-workflow-runner --lib
cargo check -p light-agent
The opt-in live gate runs two real coding turns and a close through separate worker processes. It uses a temporary private copy of the existing native login, leaving the owner’s interactive threads and configuration untouched. It verifies that a marker supplied only in the first conversation is remembered by the second and that the first patch survives reconstruction. It consumes native plan usage.
cargo build -p light-agent-worker
python3 scripts/run-coding-thread-smoke.py \
--codex /absolute/path/to/qualified/native/codex \
--worker target/debug/light-agent-worker \
--codex-home "$HOME/.codex"
The gate reports observed usage, including cached input when provided. These native usage events are not gateway billing receipts. Resumption can reduce repeated exploration, but cost/cache improvements need measurement rather than an assumption that a persisted thread guarantees a cache hit.
Verification record (September 5, 2026)
The live native gate passed new, resume, and close through separate worker processes. The same native thread ID was returned on both model turns; the second turn reproduced the first turn’s conversation-only marker and preserved its patch. Both turns reported cached input. This is a continuity qualification, not a controlled benchmark of cache savings.
The disposable PostgreSQL test exercised workflow job creation, fair activation,
selection in RUNNING, native coding scheduling, runner pinning, trusted workflow
scope, duplicate-dispatch rejection, and guarded failure handling. It used the
checked-in Agent migration, not the application database. Unit tests cover
checkpoint concurrency, stale state, interrupted execution, scope/contract
changes, close behavior, protocol shapes, and typed workflow input. The affected
Rust tests and both documentation builds passed. Production deployment and a real
multi-stage Portal workflow run are separate rollout checks; the live worker
smoke does not claim to qualify an unimplemented Claude adapter.
Personal Codex permissions and models
The codex-personal-policy-v1 extension applies to immutable-repository coding
workflows using codex-app-server-v1, pinned to Codex 0.153.4. The enterprise
route and separate workspace-service tool bridge retain their existing contracts.
Publish this optional object under
agentPolicy.execution.codingProfile.codexPolicy, and add
codex-personal-policy-v1 to the profile’s requiredFeatures:
{
"schemaVersion": 1,
"permissionSource": "codex-cli",
"permissionMode": "inherit",
"interaction": "unattended",
"allowedModels": ["gpt-6-astra"]
}
Use exact model identifiers available to the owner’s account. Optional
defaultModel must belong to allowedModels. Omitting the entire object preserves
the existing Light-managed behavior and serialized spec.
| Source / mode | Native mapping |
|---|---|
agent-policy / managed | Existing role sandbox and approvalPolicy: never |
codex-cli / inherit | Omit thread and turn approval/sandbox overrides |
codex-cli / trusted-personal-unattended | approvalPolicy: never plus thread danger-full-access and turn dangerFullAccess |
The explicit automation mode requires published personal policy. Native managed
requirements and OS/service restrictions still apply; never alone does not
grant filesystem authority. Unattended approval requests are declined, with no
raw request payload in evidence. Interactive approvals are not qualified here.
Workflow coding input cannot supply permissions, sandbox settings, flags, or
policy objects, and prompt text cannot change the admitted source or mode.
Native mode requires root-owned, non-writable /usr/bin/bwrap. An outer mount/PID
namespace keeps the host and Git metadata read-only, hides Light checkpoints,
and permits admitted implementation paths or review scratch plus native state
storage and a private /tmp for shell helpers. Native config.toml, rules, skills, plugins, AGENTS.md,
managed_config.toml, and staged .codex policy files are mounted read-only during
a turn. Missing native surfaces are initialized as empty defaults so a turn
cannot install new permissions.
The worker retains canonical patch validation, protected paths, review candidate
immutability, owner/host binding, runner leases, cancellation, deadlines, and
checkpoint ordering. Publication still uses the existing workflow contract.
External tools require service-side role restrictions: a filesystem namespace
cannot enforce read-only access to a remote MCP service. Qualify those installed
tools for their intended role before enabling this profile.
Workflow model and conversation binding
Set optional coding.nativeModel to an exact admitted native catalog identifier.
On new, it overrides the published or native default. It is independent of
modelAlias; role alias validation and enterprise alias routing are preserved.
The worker checks the account catalog and the thread response’s model/provider.
Inference may still be rejected by the account; that error fails the turn.
Unknown/unavailable models, aliases, reroutes, and alternate API providers fail
without silent substitution or paid API fallback.
The selected model is bound into private checkpoint state and sanitized
nativeSelection result evidence. On resume, omission retains and explicitly
forwards the saved model even when the owner’s local default changed. An explicit
different model or changed published policy requires a new session. Existing
sessions cannot adopt the extension in place. Close uses the saved selection
without a new catalog lookup or model turn. Uncertain native operations remain
unrecoverable and are never automatically replayed.
Configuration discovery and reload
Each invocation launches a fresh App Server with the runner-projected
LIGHT_CODEX_HOME as CODEX_HOME. Native-mode HOME and PATH come from the
runner service. Native login, rules, configured MCP, skills, and applicable
customization remain available. Installed relative paths retain their locations;
commands run in the reconstructed repository (review uses scratch). Referenced
executables must be available under the service’s OS restrictions.
A Git bundle contains committed files. Ignored .codex/config.toml, untracked
skills, and other source-checkout local files are not implicitly copied. Codex
applies its own trust rules to staged project configuration; original-checkout
trust does not automatically transfer. The worker never trusts the whole spool.
Use installed user configuration and absolute tool references for settings that
must survive arbitrary staged paths.
Owner updates are observed on subsequent invocations without republishing native rules. Already-running processes are not promised to reload. Evidence contains source, mode, model, published-policy digest, a digest of opaque native layer versions, and an ignored-layer count. It excludes settings, credentials, paths, and disabled-reason text. Revisions reveal drift; model and published-policy bindings remain fixed across resume.
Upgrade and qualification
Rebuild Agent and worker together, regenerate image/executable, capability, adapter-contract, and qualification admission through the deployment generator, then re-publish source policy and start new workflow sessions. The extension changes both the worker capability digest and qualification evidence digest; the upstream schema and native CLI version remain pinned. Do not edit generated snapshots or substitute an arbitrary installed Codex version for the pinned binary.
python3 scripts/test-codex-personal-config.py --codex /absolute/path/to/codex
python3 scripts/run-coding-thread-smoke.py \
--codex /absolute/path/to/codex --model gpt-6-astra \
--personal-policy inherit
The first gate is credential-free: inheritance, explicit overrides, managed
requirements, project trust, and reload revisions. Rust tests exercise admission,
injection, model binding, uncertain checkpoints, and actual namespace write denial.
The live gate uses a temporary private copy of the native login and consumes plan
usage. It checks explicit selection, separate-process new/resume/close, remembered
context, patch continuity, and omission after a changed local default. Repeat
with --personal-policy trusted-personal-unattended on the intended runner.
Skipped, failed, or timed-out live cases are unqualified, not passed.
Remote-tool role enforcement and the deployed workflow require deployment-specific
qualification.
The mapping is checked against the pinned schemas and binary, alongside the App Server documentation.
The worker reports capability version 0.153.4-personal-policy-v1, independently
of the native CLI pin 0.153.4. A worker with the previous capability digest
cannot advertise codex-personal-policy-v1; regenerate admission after upgrading.
The pinned native config loader rejects attempts to redefine the reserved
openai provider. This is verified through config/read and thread/start,
since the serialized Config schema does not expose model_providers.
Cancellation is observed at native protocol waits. Active turns receive
turn/interrupt with a bounded grace period. Event frames and committed checkpoint
receipts finish delivery even if cancellation arrives during shutdown.
Codex App Server upgrade strategy
Light Agent Worker runs an independently installed, exactly qualified Codex executable. Upgrading the interactive CLI does not upgrade the worker. Promote the executable, worker contract, admission, and any published coding profile as one compatible release, with the previous release available for rollback.
Process and ownership
flowchart LR
controller[Controller execution service] --> runner[light-workflow-runner]
runner --> worker[light-agent-worker]
worker -->|spawn for an execution| codex[Qualified codex app-server process]
worker <-->|JSON-RPC over stdin and stdout| codex
codex --> provider[Selected model provider]
The worker launches codex app-server, performs initialization, creates a
thread (or resumes the workflow-selected thread) and turn, consumes events, and
terminates the child after the execution. Process lifetime does not determine
conversation lifetime; see Workflow Coding Thread Lifecycle.
It does not attach to an interactive CLI session or link Codex into the worker
as a Rust library. The codex-embedded-v1 prototype is a separate, unqualified
experiment and is not upgraded with the production App Server adapter.
For the personal subscription profile, the runner supplies an owner-only
CODEX_HOME and an absolute Codex executable path. Sharing the existing login
directory supplies authentication/configuration and persisted native history.
The worker resumes only its scope-bound workflow threads, never the interactive
conversation used to configure it. The native worker pool permits one concurrent execution.
Enterprise execution uses its separate broker, credentials, and isolation
profile; personal-subscription validation does not qualify enterprise routing.
The standalone run-codex-app-server-smoke.sh launches Codex directly. It
proves the native App Server path, not worker transport, runner enrollment, or
Portal dispatch. Those require their own validation layers.
Release policy
Review upstream stable releases weekly. Expedite qualification for a relevant
security fix, blocking regression, or required model capability. Do not follow
the latest npm tag automatically at service startup or upgrade an executable
in place while jobs are running. An upstream release is a candidate until our
qualification passes.
Resolve the stable version from the official npm package and compare it with the official Codex changelog. Record the date, package version, target platform, downloaded package provenance, native binary SHA-256, generated schema SHA-256, and qualification results. Binary hashes are platform-specific; a Linux x86-64 result does not qualify an ARM or macOS executable.
Install complete native package contents in a separate version directory, including the code-mode helper and bundled resources. Preserve the previous binary, worker, runner, configuration, and admission. Never copy the global npm launcher and call that the qualified native binary.
Qualification and promotion
- Prepare a candidate without changing the active runner. Download the
exact version, verify its package provenance, record the native hash, and
check
codex --version. - Generate the candidate’s protocol artifacts. Run
codex app-server generate-json-schema --out <directory>/jsonandcodex app-server generate-ts --out <directory>/typescript. Compare both trees with the previous version. Review changed request, notification, approval, usage, error, and terminal-event shapes. OpenAI states that these generated artifacts correspond to the exact executable version; see the App Server documentation. - Update the reviewed contract. Keep the binary hash, version, schema, qualification evidence, packaging, examples, and tests consistent. Do not bypass the checks to make a new executable run.
- Qualify the native process. Test initialization and authentication in
isolation, then one real personal-subscription turn with the requested
model. Require the returned model identity to match. Verify that direct
subscription traffic leaves the local
llm_audit_event_tcount unchanged. Keep unrelated gateway traffic quiet during this global-count check. - Qualify the worker integration. Run the coding-harness gates. Cover schema validation, approval mediation, cancellation/deadlines, terminal events, usage, patch validation, reviewer isolation, authentication separation, worker binary validation, and runner registration/recovery. A skipped database test is not live database qualification.
- Rebuild and stage the release. Rebuild affected workers/runners and,
where coding profiles are enabled, the agent service. Generate admission
from the exact deployed runner binary and configuration using
print-admission. Publish matching immutable coding-profile contracts through the normal Portal workflow when that path is in use. - Drain and activate. Stop new assignments and let active work finish or follow the supported cancellation procedure. Stop the runner before switching its binary/configuration. Reload the controller admission and restart the runner. Require successful registration, fresh heartbeats, readiness, and HTTP 200 on authenticated result polling.
- Qualify the actual product path. Before describing Portal coding as qualified, submit a controlled job through the enabled Portal agent and observe dispatch, execution, result acknowledgement, and final projection. Registration and a direct Codex smoke are narrower evidence.
Run the cumulative local regression gates from light-fabric:
LIGHT_CODEX_SMOKE_EXECUTABLE=/absolute/path/to/candidate/bin/codex \
./scripts/run-coding-harness-phase5-gates.sh
For a personal GPT-6 Astra qualification from the workspace root:
LIGHT_CODEX_NATIVE_EXECUTABLE=/absolute/path/to/candidate/bin/codex \
LIGHT_CODEX_SMOKE_MODEL=gpt-6-astra \
./portal-config-loc/all-in-lt/light-workflow-runner-personal/run-smoke.sh
Model access is an account/provider capability as well as a client capability.
An upgraded CLI cannot grant access to a model. For personal coding, the worker
currently uses native Codex model selection; it does not forward the Portal’s
enterprise aliases such as coding-implementer as native model names. Check
the effective CODEX_HOME model configuration. An explicit smoke-model override
tests a model without rewriting the user’s interactive CLI configuration.
Files that form the current release contract
| Artifact | Purpose |
|---|---|
crates/coding-agent-runtime/src/lib.rs | Accepted adapter version, binary, schema, and evidence digests |
contracts/codex-app-server/v<version>/ | Generated JSON/TypeScript schemas and provenance |
contracts/coding-adapters/codex-app-server-v1-qualification.json | Qualification metadata whose digest is checked by the worker |
apps/light-agent-worker/src/codex_app_server.rs | Runtime validation and generated-schema tests |
scripts/generate-codex-app-server-schema.sh | Reproducible pinned schema generation |
scripts/run-codex-app-server-smoke.sh | Exact version/hash and live native model qualification |
scripts/run-coding-harness-phase1-gates.sh, run-coding-harness-phase5-gates.sh | Schema/evidence checks and cumulative regression gates |
apps/light-workflow-runner/docker/Dockerfile | Exact native distribution shipped with the runner image |
| Controller admission and Portal coding profile | Approved deployed binary/configuration/capability identities |
The runtime currently accepts one hardcoded qualified version. Supporting a small set of approved versions via a signed or release-owned qualification manifest is a future improvement. Such a manifest would still require exact binary hashes and validation evidence; it would not authorize arbitrary versions or remove the need for compatible admission and Portal contracts.
Local deployment and rollback
The local deployment keeps private tokens, admission, binaries, and the
Compose overlay in
portal-config-loc/all-in-lt/light-workflow-runner-personal/.runtime/, which is
ignored by Git. Its user service is light-workflow-runner-personal.service.
Use that directory’s configure-local.py to regenerate the exact admission
after copying qualified binaries, then start.sh to reload the controller and
runner. The runner credential expires after 30 days; renew it using the same
documented setup procedure. Reapplying only the base Compose file can remove
the local runner overlay.
Before activation, save the old codex, companion files, light-agent-worker,
light-workflow-runner, runner.yml, admission.json, and private
compose.yml together in an owner-only rollback directory. Retain the durable
execution journal in place. For rollback, drain/stop the runner, restore the
matching artifact set, reload the controller, and reconnect. Do not restore an
old journal over newer execution evidence or restore only the Codex executable
while leaving new worker digests/admission active. Renew expired credentials
instead of restoring an expired token. Roll back any published coding profile
through its normal versioned publication process.
Upgrade record: 0.153.2 to 0.153.4
The npm stable version and official changelog were checked on 2026-09-05.
Release 0.153.4 was published on September 4 and fixes Astra’s bundled picker
visibility/default selection and async-question guidance. GPT-6 Astra support
was already introduced in 0.153.1; 0.153.4 improves its integration.
See the release notes.
| Evidence | Candidate value |
|---|---|
| Native package | @openai/[email protected] |
| Platform | x86_64-unknown-linux-musl |
| Native SHA-256 | 56ef98ab4032d317ab26e9b5e5a175650717351edb16ed9cde0cb6d1734d62da |
| v2 schema SHA-256 | d3eace08be5dca386bfd1f1e8df650058b4113f1e10870a284d775d75517576a |
| Schema comparison | All 1,010 generated JSON and TypeScript files are byte-identical to 0.153.2 |
| Requested live model | gpt-6-astra |
Verified on 2026-09-05:
- The native personal-subscription smoke completed a real turn with
gpt-6-astra; the returned model matched and the gateway audit-row count did not change. - The cumulative
run-coding-harness-phase5-gates.shcompleted successfully, including stages 1 through 4, unit/integration regressions, Clippy, formatting, schema/evidence checks, and mdBook. Existing Clippy warnings remain warnings. - The local release runner and worker were rebuilt and activated. Worker
capabilities report adapter version
0.153.4and capability digestsha256:80e4dcb3d624635cb871bcbdccc18641550a0906ac8d438f8e2ab49be167d12d. - Controller registration succeeded, runner readiness returned
true, the execution database reported a connected runner, and authenticated agent result polling returned HTTP 200. - The previous runtime artifact set is retained at
.runtime/rollback-0.153.2/; the active complete native distribution is at.runtime/codex-0.153.4/in the local personal-runner directory.
This evidence qualifies the local native model path and runner connectivity. No Portal coding profile was published, no end-to-end Portal coding job was submitted, and no enterprise image or remote deployment was promoted as part of this local upgrade. The updated Dockerfile pins the new distribution, but a production image build and enterprise live qualification remain deployment gates for those environments.
Subsequent enrollment verification found that the sibling Portal publisher
always emitted an empty coding profile. The local light-portal source now
supports a validated instance property: module agent-policy-authoring,
property codingProfile, type map, explicitly assigned to the instance’s
product version. It publishes the complete map into
agentPolicy.execution.codingProfile and binds it into policy digests. See
light-portal/db-provider/README.md, “Agent coding-profile publication”, for
catalog setup, validation, and the normal publication workflow.
Deploy the updated Portal query/command services and align the account-agent container with the upgraded worker contract before publishing that profile. The source fix and Java/Rust digest tests do not constitute a deployed profile or an end-to-end Portal coding test.
Personal Development Workflow Orchestration
Status: architecture corrected September 30, 2026 by owner direction. External API integration belongs in Gateway configuration and workflow definitions. The GitHub capture/provider implementation track under #427 is withdrawn; the independent #425 artifact fix is retained. This is a documentation/plan revision, not a runtime rollback, deployment or live qualification.
The immediate direct API pilot and Ownership sections supersede conflicting September 29 capture/provider instructions. Broader development/review/workspace contracts remain for separate Phase 1 review; they do not authorize implementing external integrations inside the Workflow engine.
This document covers a developer using codex-personal and claude-personal
with their own subscriptions on a local machine or dedicated VM. The separate
Enterprise Development Workflow Orchestration
design covers codex-enterprise inside a corporate network with llm-gateway.
Decision And Scope
Use Codex for requirements, design, optional planning, implementation, and fixes. Use Claude to review the design, plan, and every completed implementation phase. For implementation scope, after all phases pass, run a Codex final review and then a Claude final review. Start each final reviewer fresh once; resume its own conversation for subsequent fix verification. Workflow definitions invoke registered APIs and generic workspace tools to publish documents, update comments and commit/push accepted changes. External-system integration does not introduce subsystem-specific engine tasks.
User, Application, And Workflow Authorization provides the shared authorization foundation. Apply its current contracts together with Workflow Invoke And Tool Binding Publication and the execution-authority rules below. Personal subscriptions do not authorize Portal/API access. Historical credential-broker enrollment and database-bridge instructions are not prerequisites for the current native MCP path.
The pilot uses independently started top-level stage workflows connected by accepted artifact references. A parent coordinator and durable child workflows come later. The personal pilot does not depend on enterprise sandboxing, billing ledgers, new Chat interaction cards, a CLI, or full installer conformance. Phase 1 still requires the storage/recovery qualification for the local and personal installer variants specified below. Extend the existing Workflow Admin and Worklist surfaces for stage starts and human decisions; their current generic controls do not yet implement every development-feature transition.
The first pilot admits one active feature per VM. Shared-runner concurrency and coordinated routing across multiple VMs are Phase 4 qualifications; the capacity design below is not a claim that the first pilot supports them.
Keep these business contracts compatible with the enterprise profile: accepted requirements, optional planning, phase review, finding identity, review coverage, publication receipts, and completion criteria. Changing execution profile does not waive a lifecycle gate; enterprise authorization/accounting add profile gates.
Ownership
| Component | Responsibility |
|---|---|
| Developer and workflow author | Define requirements, review steps, data mappings, pagination, retry/error branches and coding-agent inputs |
| Portal UI and backend | Author/register API tools and workflow definitions, configure grants and approvals, and manage existing product/instance/configuration snapshots |
| Gateway / mcp-router | Verify user/application credentials, enforce configured ACLs, route registered APIs/tools and apply the configured upstream credential |
| Identity/OAuth | Supply original-user and demand-driven LONG exchange behavior; reject ineligible renewal |
light-workflow | Execute generic tasks and transitions, persist ordinary workflow context, enforce existing execution/lease/cancellation rules and invoke registered tools/agents |
| Registered external APIs | Own external resource semantics; GitHub REST/GraphQL is reached directly through Gateway configuration |
light-agent and Controller | Admit and route coding-agent work through existing execution mechanisms |
| Personal runner / workspace manager | Host workers, workspaces, generic Git commands, checkpoints and validation tools |
| Existing artifact storage | Retain artifacts for existing consumers; issue input does not create a special capture store, receipt or admission subsystem |
The API integration boundary is generic. Adding GitHub, Bitbucket, ServiceNow, Salesforce, filesystem or network tools does not add subsystem-specific parsing, credentials, policy catalogs or task kinds to the Workflow engine. API-specific schemas and workflow expressions are configuration/definition assets. A real missing generic capability or defect requires a separate, reviewed change and an integration-independent regression.
The earlier light-github-action-provider architecture is withdrawn. That app
currently has branch/PR and document-publication consumers; its retirement must
migrate or explicitly retire those consumers before removing deployment wiring.
It is not extended with an issue reader. Existing broader development-stage
implementation is historical context, not authorization to expand engine-owned
GitHub orchestration.
Component Diagram
flowchart LR
U[Developer] --> P[Portal editor / API catalog / approvals]
P -->|workflow_start| G[Gateway / mcp-router]
G --> W[Generic workflow engine]
W -->|API/tool call with user and app credentials| G
G -->|ACL checked; configured GitHub credential| H[GitHub REST / GraphQL]
H -->|API result| G
G -->|result or error| W
W --> V[Workflow context and definition expressions]
V -->|next task / explicit retry or failure branch| W
W -->|existing coding-agent task; context input| A[Agent / Controller / runner]
W -->|only when a call needs renewal| I[Identity / LONG exchange]
The user token authenticates to Gateway; the configured GitHub token authenticates the outbound GitHub request. Secrets stay in the established protected Gateway configuration/secret mechanism, not tool arguments, context, returned JSON or model input. The selected upstream and credential are fixed by published API/tool configuration, not by an arbitrary URL supplied with a credential.
Workflow definitions map raw API responses into context and coding-agent input. Pagination loops, limits and retry/wait/failure decisions are definition behavior using supported generic capabilities. There is no autonomous GitHub polling or snapshot/receipt creation in Workflow. REST is the initial proposal; GraphQL can be selected at the API-contract checkpoint using a configured query and variables. Neither choice adds a GitHub-specific engine path.
The dispatch scheduler’s shared fairness and multi-VM routing are Phase 4 work;
the initial pilot still admits one active feature per VM. Optional future
feature-delivery parent/child orchestration, the deferred light-cli TUI and
enterprise llm-gateway are outside this personal pilot diagram.
The Agent never advances a workflow stage based on conversational text. The workflow selects an admitted Agent and sends typed jobs; it does not launch the native CLI directly. GitHub and native conversations are not workflow state.
operations.workflow_ops owns Workflow runtime processes, tasks, feature records,
stage claims, Agent dispatch intents, and effect receipts. Agent owns its separate
agent_ops.agent_job_t admission/execution records. Registered Agents pull bounded
jobs and report results through the authenticated job transport. The shared job
identity correlates these records; it does not grant either service access to the
other’s database. Portal owns definition authoring and its publication ledger,
while Workflow persists the acknowledged definitions and execution policy it
uses. Portal runtime projections and old started events are not execution authority.
Invocation And Execution Authority
Development stages use asynchronous workflow_start through Gateway MCP. The
Editor currently reaches it through Portal’s StartWorkflow command, which
acknowledges definition/grant synchronization and supplies
expectedDefinitionDigest. A development stage start additionally carries its
typed claim and pinned inputs. Every entry point must reach the same
claim-and-start application transaction; using the Editor must not bypass the
feature ownership check.
flowchart LR
UI[Workflow Admin or Editor] --> P[Portal start adapter and definition sync]
P --> G[Gateway MCP and ACLs]
O[Other authorized async callers] --> G
G --> S[Workflow workflow_start]
S --> C[Atomic feature claim and run admission]
C --> J[Workflow-owned Agent job intent]
J --> A[Agent admission and execution record]
A --> R[Controller and runner]
R --> A
A --> J
An Agent invoking a workflow-backed Tool calls the published Tool on Gateway.
Gateway internally calls synchronous workflow_invoke, which enforces its
published binding and returns output, terminal failure, or a bounded timeout.
Clients do not call workflow_invoke directly. This path does not replace the
asynchronous development-stage lifecycle or implement durable child workflows.
Native development Agent turns currently require asynchronous claimed stages.
At Gateway ingress, Authorization must contain the acting user’s bearer.
An optional caller X-Scope-Token is independently validated as an application
token and never substitutes for the user. Gateway forwards the user bearer and
supplies its own application token to Workflow. Workflow retains Host/owner,
definition-grant, feature-version, assignment, and execution-authority checks.
Gateway-to-Workflow JWT authentication does not require mTLS; the separately
configured Agent/Controller/runner transports retain their own authentication.
Async runs use Workflow-owned LONG authority for work that outlives the initiating bearer. New stage admission must establish that run’s authority exactly once; replay must retain the existing binding. Synchronous Invoke uses its bounded run credential and does not register LONG. Neither route revives the retired credential broker. Native workers receive no refresh credentials. Published workspace Agent identities and complete runner bindings remain independently required for native admission.
Persist feature and stage deadlines and consumed budgets independently of token expiry, HTTP waits, and credential renewal. Human/capacity waits do not consume a remediation round, but do not reset the feature deadline. Revocation or unavailable authority prevents new execution; recovery cannot silently widen authority or replay uncertain work. Any reauthorization must use the supported explicit authority flow, retain consumed budgets and evidence, and recheck current policy.
Async start acceptance proves durable admission, not completed execution or stage
acceptance. Status/result/wait operations attach to the recorded instance. A lost
response retains the same start/operation identity for reconciliation; a new
transport ID is not permission to repeat work. A timeout is not proof that an
execution or external effect stopped. Preserve structured failure, retryability,
and afterEffect evidence through the UI and fixed actions.
Workflow Composition
| Top-level workflow | Pilot responsibility | Output |
|---|---|---|
feature-intake | Normalize an existing issue or agreed Codex requirement draft | Feature issue and accepted requirement version |
feature-design | Codex authoring, validation, Claude review/fixes, optional human sign-off, document publication | Accepted design, commit permalink, and PR reference |
feature-plan | Optional detailed plan with the same review/publication loop | Accepted plan and phase manifest |
feature-implement | Implement and review one selected phase | Accepted phase checkpoint and evidence |
feature-finalize | For implementation scope, full-feature reviews and delivery; for design-only scope, fixed verification of the already-reviewed design and its delivery | Evidence for the pinned delivery scope and completion target |
For the pilot, an operator starts the next top-level stage with the preceding
stage’s accepted output. Start feature-implement once for each declared phase,
in dependency order. Each phase performs its own review/fix loop. No stage
starts an untracked background workflow, and a stage’s acceptance does not
declare the whole feature complete.
Persist a FeatureRun record across these stage instances. Inputs include the
feature/run and issue identity, stage/phase ID, expected previous accepted
version, applicable artifact digests, workspace/task and repository bases,
Agent/runner bindings, deadlines, and remaining budgets. Only one stage owns
mutable task execution at a time. Starting a successor atomically claims the
expected accepted predecessor/version; duplicate starts return the same instance.
A deliberate replan or reopened phase gets a new stage execution identity and
retains the same feature ledger and consumed budgets.
A stage returns an immutable StageResult containing its input/output digests,
acceptance status, findings, validation and publication receipts, and native/file
checkpoints and diffable snapshot references where applicable. The next stage
validates this result instead of trusting an issue comment or copying an entire
conversation.
The pilot uses bounded stage-local iteration and reliable Agent job completion. Pinned definitions can materialize a finite set of review-round task slots; a new logical turn uses a distinct durable task/job identity, while a retry keeps its identity. The current service-call idempotency key includes process/task IDs, so reusing one completed task ID is not a new remediation turn.
Later, feature-delivery can automate these handoffs and use implement-phase
children. That extension must add durable child start/wait/result/cancel,
definition/input pinning, parent budget accounting, and restart recovery.
run.workflow is modeled but currently rejected by the executor; it is not a
prerequisite for the standalone design pilot.
Enforced Stage Handoffs
Use the existing FeatureRun store in the light-workflow operational
database and its claim-and-start transaction behind native MCP. This is required
even while an operator selects and starts each top-level stage. Persist the current feature
version, accepted artifact/phase versions, active stage owner, consumed budgets,
immutable StageResult references, and stage-claim receipts.
The start transaction locks the feature record and first checks for an existing
claim to replay. For a new claim it validates the expected accepted predecessor
and all current input versions, and checks that the selected stage is an allowed
successor with no conflicting owner. It atomically records the
claim, workflow instance/process and initial task, pinned inputs/definition, and
new stage owner. development_store.rs::claim_and_start already composes feature
checks with invocation/process creation. Ordinary invocation idempotency alone
does not implement the feature-version/ownership check. A claim followed by a
separate unprotected workflow start is insufficient.
The application result must distinguish a newly created claim from a replay. For new work, admission supplies the configured artifact store, verifies its readiness, and establishes the run authority before any task can dispatch. Persist the authority binding with the admitted instance. If external authority registration is uncertain, reconcile its stable registration identity before execution; never leave an unauthorized runnable process or mint another binding on a blind retry. A replay returns the original claim/run and does not renew a completed run or reacquire a released VM. Readiness checks for new work must not prevent authenticated retrieval of an already-committed historical receipt.
Give each logical transition a store-owned identity, including the feature, predecessor version, and selected stage/phase. A unique claim and normalized request digest make concurrent identical starts return the same instance, even with different transport request IDs. Changed inputs for an existing claim conflict; an unclaimed stale predecessor cannot start. Retrying an older completed claim returns its historical receipt and never launches a new stage. A new replan/reopen transition must be explicitly recorded by the state machine.
Workflow Admin stage starts use this MCP-backed transaction. Dispatch and fixed effects require the claim bound to their process; a generic workflow start cannot bypass it. Stage acceptance verifies durable output/snapshot receipts and current input bindings, then atomically records the result and releases ownership for the next stage. Document supersession/replan invalidates affected input bindings in this store. Unknown worker execution keeps ownership fenced until reconciled. Restart after a committed start but lost response recovers the recorded instance; rollback leaves neither a claim nor a runnable process. Feature budgets survive handoffs.
Feature Operations And Human Decisions
Expose development-feature operations through native Workflow MCP on Gateway, using shared application handlers rather than a separate REST workflow for the UI. The following are semantic operations; missing Tool names and wire schemas must be specified and qualified in the implementation plan before publication.
| Operation | Required binding and result | Current interface boundary |
|---|---|---|
| Start a stage | Pinned definition, claim, predecessor and inputs; return durable instance/claim receipt | workflow_start exists; development admission integration remains open |
| Read process/feature and VM holder | Host/owner filtering, current versions, active instance and release state | Native list/get Tools and Workflow Admin views exist |
| Accept a stage | Expected feature version, operation ID, validated output/review/sign-off/publication evidence and allowed successor | Application handler exists; native MCP exposure remains open |
| Replan or reopen | Expected version, operation ID, affected input versions and acceptance invalidation; preserve budgets/fences | Application handler exists; native MCP exposure remains open |
| Publish a fixed intent | Pinned operation and target, approved retained content, stable effect identity; confirmed or unresolved receipt | Finalize-only HTTP dispatcher exists; lifecycle and native MCP integration remain open |
| Record design sign-off | Assignment identity/version, feature/stage and exact design digest, authorized decision and durable evidence | Generic human-task Tools exist; development sign-off producer remains open |
| Finalize design-only delivery | Pinned scope/terminal definition, unchanged approved contents and required delivery receipts | Application handler exists; native MCP exposure remains open |
| Cancel feature and release VM | Expected feature version, store-resolved reservation generation and fencing evidence | Native workflow_cancel_feature and UI control exist; full release qualification remains open |
Mutations retain an operation identity and normalized request digest. Identical retries return the recorded outcome; changed requests conflict. UI clients retain pending operations after transport uncertainty, refresh authoritative state, and distinguish durable decision recording from executor continuation. Legacy private HTTP adapters may delegate to the same handlers during migration, but completing an HTTP-only diagnostic flow does not qualify the Gateway/UI boundary.
End-To-End Lifecycle
flowchart TD
I[Existing issue or Codex requirement dialogue] --> R[Accepted requirements and feature issue]
R --> D[Codex design, Claude review, and fixes]
D --> S{Design sign-off}
S -- disabled or approved --> DP[Publish accepted document revision]
S -- changes requested --> D
S -- rejected --> H[Human resolution required]
DP --> DS{Pinned delivery scope}
DS -- design-only --> DF[Fixed design finalization and delivery checks]
DF --> DD[Complete design-only run and fence VM release]
DS -- implementation --> P{Separate plan needed?}
P -- yes --> PL[Codex plan, Claude review, fixes, and publication]
P -- no --> PH[Start implementation phase]
PL --> PH
PH --> C[Codex implementation and fixed validation]
C --> V[Claude phase review]
V -- findings --> C
V -- accepted --> N{More phases?}
N -- yes --> PH
N -- no --> CF[Codex full-feature review]
CF -- findings --> CX[Codex fixes and validation]
CX --> CV[Resume Codex reviewer to verify changes]
CV -- findings --> CX
CF -- accepted --> CL[Claude full-feature review]
CV -- accepted --> CL
CL -- findings --> FX[Codex fixes and validation]
FX --> RV[Resume Claude reviewer to verify changes]
RV -- findings --> FX
CL -- accepted --> G[Review coverage and publication gate]
RV -- accepted --> G
G -- coverage missing or broad change --> OR[Resume required other final reviewer]
OR -- findings --> OF[Codex fixes and validation]
OF --> OR
OR -- accepted --> G
G -- complete coverage --> COMMIT[Commit exact accepted changes]
COMMIT --> PUSH[Push GitHub task branches]
PUSH --> PR[Create PRs and post delivery links]
PR --> DONE[Verify completion target]
The stage-to-stage edges are operator handoffs in the pilot. Review loops run inside their stage. All loops are bounded. Scope expansion, stale inputs, or missing review coverage route to the appropriate additional review or replan before publication, as specified below.
The coverage gate selects every required reviewer of each saved delta, including the original reviewer again if an escalated review causes a further broad fix. The other-reviewer edge resumes an existing final conversation; it does not restart the initial full reviews. Document revisions discovered during phase or final work follow the revision loop below before that work can resume.
Requirement Intake
Immediate Pilot: Direct API Issue Input
The owner supplies an issue URL to a workflow started with asynchronous
workflow_start. Reuse the demonstrated demo3-workflow editor/tool setup and
approval path using call: http and registered lightapi:// API targets.
Native binding-less MCP is not required by this flow. Obtain its saved definition and API configuration during the new
plan’s first checkpoint; do not infer that a custom capture dispatch path is
required. Workflow-backed tools continue to use workflow_invoke with the target
tool’s existing setup; the calling native workflow needs no outer Tool binding.
The definition performs these ordinary steps:
- Validate/derive the registered API arguments from the input URL using supported expressions. The Gateway upstream remains fixed by configuration.
- Call registered GitHub REST operations (or a reviewed GraphQL query) through Gateway for the issue and all required comment pages. Authentication, ACLs and upstream credentials follow the normal Gateway configuration path.
- Accumulate returned data in ordinary workflow context. Use supported CEL or another already-installed generic transformation facility; jq/JavaScript availability must be demonstrated rather than assumed.
- Use the standard HTTP task
retry:policy for bounded attempts and delay when desired; the engine durably schedules eligible retries vianext_attempt_ts. No retry policy or exhausted retries fails the workflow. P01 Wait is for polling loops over successful responses, not a prerequisite of failed-call retry. Try/Raise and status-specific error branches are not v1 requirements. GraphQL HTTP 200 can carry errors; inspect those if selected. - Pass the structured context to the existing coding-agent task, then expose the design output for owner review. If the existing agent handoff is incomplete, report its exact generic integration gap before changing the engine.
Issue bodies/comments and attachment links are input data, not executable instructions or authority. Specify attachment metadata versus downloaded bytes at the API checkpoint. Download only through approved routes without forwarding GitHub credentials to unrelated attachment hosts. Do not claim that one response contains every comment page or every attachment’s contents.
There is no feature-design-input capture subsystem, capture/status API family,
READY capture row, immutable creator/task projection, receipt admission, policy
currentness lookup or capture retention job for this flow. The issue URL and API
results are regular context values. Existing artifacts used by other features
are unaffected. A second start reads the API again unless the definition explicitly
chooses another behavior.
Normal local/timer execution does not inspect token expiry or renew a token. When an operation needs a user token, use the current credential if sufficiently valid; otherwise exchange through LONG. Gateway enforces ACLs. Definitive renewal denial fails the operation/workflow; transient handling follows the existing bounded generic contract. There is no continuous Portal membership check or cross-database membership fence. Preserve original user identity through existing LONG behavior without adding a special creator table for GitHub.
The former GitHub Issue Input C0/S00/P01–P05/S01–S06 implementation track under
issue #427 is withdrawn.
Its successful component and owner-reported deployment tests remain historical
evidence, not approval of that architecture. The replacement implementation plan
is implementation/light-agent/2026-09-30-DirectApiDesignWorkflowImplementationPlan.md
in the sibling implementation repository. Old plans are archived and must not
be executed. Broader Phase 1 delivery requires separate revalidation.
Keep issue #425: it fixed a
pre-existing shared-object deletion defect affecting all artifacts. Its A00
commit 6bd01e0 (merged by df64d06) and migration 0020 are independent of the
withdrawn GitHub capture design. Do not revert them.
The #427 correction keeps P01 (677a3bb) as the generic durable DSL wait
capability, including migration 0021, nonblocking wake, cancellation and deadline
handling. Remove its capture-profile-specific absolute preparation cap and use
neutral polling fixtures. Current wait support is literal whole-second PT<n>S
for 1–600 seconds inclusive, outside fork branches. Definitions can reuse this
capability for explicit retry delays; it is not a GitHub subsystem.
Revert P02 (5a41529) capture-specific runtime and tests, while carrying its
independent generic fixes in a small separate change: guarded TLS stream support,
configured client mTLS loading, LONG fixture repair, inspected action-ID forwarding
and the timer’s nullable binding read (needed for native async waits). Preserve
working demo3, async start, sync Invoke, LONG and unrelated changes. Keep deployed
0022 history immutable; after quiesced definition/reference/row-count checks,
remove its three unused verified-context/result tables with a new forward
migration (0023 if still free). Nonempty tables block cleanup for review.
No runtime or database rollback is performed by this design revision.
The adjusted P02 revert must retain those generic fixes and the original 0022 SQL/package/staging entries before it is committed. Do not merge a broken revert and repair it in a later PR. Prove native NULL-bound wait and guarded TLS at the revert commit as well as on the final cleanup tree; the implementation plan lists the exact migration assets and safe commit order.
Broader Intake Lifecycle
Support both entry paths:
- Existing issue: a fixed read action loads the selected issue and relevant requirement comments. Preserve source IDs and a content snapshot. Codex identifies material gaps; unresolved scope decisions block acceptance.
- New requirement: the developer works with
codex-personal, which drafts the title, body, andRequirementArtifact. After agreement, a fixed action creates the issue and persists its repository, number, and URL.
The artifact records scope, non-goals, acceptance criteria, affected repositories
(or explicitly unknown ownership), compatibility/migration concerns, source
references, unresolved decisions, and a version/digest. Intake also selects a
completion target: pr-ready, merged, or deployed, with its required checks.
Also pin a delivery scope: design-only or implementation. This is a
required contract addition, not a claim that the current wire schemas already
expose such a field. Design-only ends after intake, reviewed design, optional
sign-off, document delivery and fixed finalization. Implementation continues through any
plan, every declared phase, and both final reviewers. The completion target
applies within that scope: a design-only pr-ready result needs its document PR
and declared checks, and cannot claim implementation delivery. Unsupported
scope/target combinations fail admission. A scope change requires an explicit
versioned replan, not selection of a weaker Finalize definition.
Use one feature issue by default; add linked repository or deferred-work issues only when separately owned tracking is needed. Issue creation is idempotent by feature and tracking purpose, not document revision. Existing issue content is requirement data, not authority to change worker permissions or run arbitrary commands. Changes after freeze create a new version and an explicit impact/replan decision; they are not silently folded into the current implementation.
For the new-issue path, reserve a durable intake identity and its authorized issue creation intent before GitHub I/O. Permit an issue-pending intake record until the confirmed issue receipt is bound; no design or implementation successor may start in that state. Retry reconciles the same intent. The current intake seed requires an existing issue reference, so this pending state needs implementation rather than a fabricated issue number or untracked pre-intake write.
Design And Optional Plan
Choose one canonical design path, normally in light-portal-doc/src/design
or light-fabric/docs/src. Codex authors against the accepted requirements.
Fixed validation runs document/navigation checks; Claude reviews the current
artifact and finding ledger. Codex fixes, validation reruns as needed, and the
same Claude reviewer verifies changes until the closure contract passes.
A per-run requireDesignSignoff setting defaults to false. When enabled,
Claude’s acceptance creates a Worklist decision for the exact design digest
before publication and implementation. Record the human’s approval, rejection,
or requested changes. A new design digest requires a new sign-off when the
setting is enabled. Requested changes return to design authoring/review;
rejection enters HUMAN_RESOLUTION_REQUIRED until an explicit revise or cancel
decision. Routine review rounds do not add other human gates.
Workflow creates the assignment from the pinned sign-off policy and retained
design candidate. Completion verifies current assignee/role eligibility, claim
and assignment version, current feature/stage, and the exact design digest.
Persist the human-task decision, DesignSignoff authority evidence, and the
resulting feature transition atomically, or through a durable idempotent
continuation that keeps acceptance blocked until its evidence is committed.
An identical completion retry returns the same decision; a stale or changed
decision cannot approve a superseding candidate. Generic task completion or a
Tool-binding access approval is not a design sign-off receipt. Implement the
producer as well as the existing acceptance-side validation.
For implementation scope, decide once after design acceptance whether a separate plan is necessary. Cross-repository work, migrations, contract changes, and several dependent phases usually benefit from one. Codex saves it in the implementation repository; Claude reviews and verifies fixes using the same stage contract. A small feature can use one phase derived directly from the accepted design. A sufficiently detailed design can supply several phases without a duplicate plan.
The accepted phase manifest defines scope, owning repositories, dependencies, validation commands, and exit criteria. Implementation consumes it; it never starts by generating a second plan.
Document Publication And Revision
The pilot uses a separate immutable publication task and branch for each
accepted document revision, with a PR targeting that repository’s develop
branch. This works with the existing absent-or-identical-ref push rule and does
not depend on implementing updates to an already-published ref.
For example, a feature may publish design v1 and v2 through distinct tasks/refs
derived from featureRunId + documentId + revision. The actual ref uses the
manager’s task-branch naming convention; the revision identity is durable and
is not regenerated on retry.
For each accepted design or plan revision:
- Freeze the accepted files, including any navigation changes, with validation and review evidence and optional design sign-off.
- A fixed action creates a dedicated publication task on a recorded base, materializes only that accepted document file set, verifies it, and commits it.
- Push its unique task ref and create a document PR. Persist the commit SHA, ref, PR number/URL, and file permalink. The feature issue links to the commit, not just a local path or moving branch.
- Retain the revision ref/commit for the feature’s artifact-retention period. Record acceptance and supersession in the ledger/status comment. Publishing a document does not automatically merge its PR or close its feature issue.
Here, an accepted revision means an immutable candidate approval containing
its exact snapshot, checks, review coverage and any required sign-off. Persist
that approval while the design/plan stage still owns the feature. Fixed document
publication consumes it; only after required publication receipts are confirmed
does stage acceptance emit the final StageResult and enable a successor. This
avoids requiring an accepted StageResult to publish while simultaneously
requiring publication to accept that same stage. Publication failure retains the
approval and stage owner for reconciliation without another author/reviewer turn.
Delivered document bytes and modes must match the approved file set, including navigation edits. Keep recovery markers in effect metadata, commit/PR metadata, or comments; publication steps must not append markers to approved document contents. Any intentional content transformation must happen before snapshot validation and review. A GitHub Contents API write alone is not an immutable revision-task, branch, PR and verification receipt.
Later implementation discoveries use the same path: propose a new document version, assess its effect on requirements/plans/accepted phases, run the applicable author/reviewer/sign-off cycle, then publish a new revision task, ref, and PR. No old commit, review record, or permalink is rewritten. If an old revision already merged, base the replacement on the current recorded integration revision; if it has not merged, the new snapshot must contain the complete desired document state. Mark the old unmerged PR as superseded in the feature ledger; closing it is a fixed action under the run’s current authority and pinned publication policy.
The latest accepted document revision is authoritative for subsequent stages.
Canonical design/plan paths are delivered by their document PRs; implementation
PRs must not carry competing edits to those paths. If implementation edits one,
route the new contents through document acceptance/publication and reconcile
the implementation candidate before final review. Declare the latest document
PRs in the final delivery manifest. A pr-ready target may leave them open;
merged/deployed require their integration as declared by the feature.
Base changes caused by merging documents receive the same impact/validation
treatment as other integration changes.
Snapshot primitives and basic publication dispatch exist. Their integration into this complete revision-publication path, including candidate approval, supersession and document PR receipts, remains pilot work. Reusing distinct manager tasks avoids its existing one-push/one-PR-per-task receipt conflict without weakening that constraint.
flowchart TD
W[Phase or final work discovers a document change] --> R[Record REPLAN_REQUIRED and affected input versions]
R --> D[Re-enter affected design or plan stage]
D --> A[Author, validate, review, and fix revision]
A --> S{Required design sign-off}
S -- changes requested --> A
S -- rejected --> H[Human resolution required]
S -- disabled or approved --> P[Publish new revision task, ref, and PR]
P --> B[Supersede prior revision and reconcile implementation]
B --> C[Claim affected stage with current inputs and retained budgets]
C --> W2[Resume work and required review coverage]
Record the revision request and reconcile any active worker before handing off the stage owner. Resume the earliest invalidated phase or final stage selected by the impact decision; publication of a new document alone does not restore invalidated implementation acceptance.
Diffable Candidate Snapshots
The current checkpoint::capture in crates/task-workspace/src/checkpoint.rs
records file hashes/modes, HEAD, and index/status hashes. It retains no prior
file contents. workspace.checkpointDigest detects change; it cannot supply
historical review content or a diff after later edits overwrite those files.
Use manager-owned CandidateSnapshot receipts, implemented by
crates/task-workspace/src/snapshot.rs. Capture the initial
stage baseline, every candidate submitted for review, and every resulting fix
candidate before another edit can overwrite it. Accepted phase checkpoints
retain their snapshot references. This applies to design/plan fix rounds as
well as implementation and final review; saving accepted phases alone is too
late to verify intermediate fixes.
Under the task’s exclusive lock, the manager writes a Git tree from a temporary private index containing the checkpoint’s complete admitted file set. Preserve exact bytes, executable modes, additions, and deletions, including non-ignored untracked files; do not apply content-changing Git filters. Keep the current checkpoint restrictions on symlinks and submodules. An unsupported or out-of-scope change blocks acceptance instead of silently disappearing from the diff. Verify the captured bytes against the checkpoint manifest and recheck the workspace before completing the receipt. Drift produces no accepted snapshot.
Pin each per-repository tree in a unique manager-only local ref, such as
refs/light-workflow/<featureRunId>/snapshots/<snapshotId>, without changing HEAD,
the task branch, or the real index. These are local tree snapshots, not feature
commits; fixed push actions never include the private refs. The receipt binds
the feature/stage/task, repository bases, file checkpoint digest, per-repository
tree IDs, and content-manifest digest. Native conversation checkpoints remain
separate. light-workflow exports the contents and manifest supplied by a fixed
manager read into the immutable artifact store before a review or stage result
can depend on them, so recovery does not depend on a surviving local Git object
database. The storage and transfer requirements are defined below.
The manager derives each complete delta from the saved before/after trees;
light-workflow persists its artifact/digest with both snapshot references.
Retain binary contents and mode/deletion evidence as well as the textual diff;
reviewer tools must be
able to read either saved version. This is a fixed read operation, not native
shell access or a model-authored patch. Workflow impact routing consumes this
evidence. Missing snapshots, failed reconstruction, or digest mismatch block
review/publication rather than falling back to an implementer summary.
Retain all baseline and transition snapshots referenced by StageResult or
ReviewCoverage for the feature’s evidence-retention period. Garbage collection
cannot prune referenced trees/content. Final commit verification compares the
delivered file contents with the accepted snapshot; these snapshots do not move
implementation commit/push ahead of final closure.
Personal Artifact Storage And Export
The personal pilot uses a filesystem-backed durable artifact store owned by
light-workflow. DurableArtifactStore now supports filesystem and S3 backends;
the separate workflow.fixedActions.artifactRoot scratch directory does not
enable the durable store or its publication/recovery contract.
The supported settings are workflow.artifact.backend: filesystem and
workflow.artifact.filesystemRoot: /var/lib/light-workflow/evidence, alongside
the existing artifact prefix/retention settings. The template accepts these keys;
qualify their effective published values and dedicated persistent volume in
portal-config-loc/all-in-lt and the personal stack in light-portal-install.
Only the Workflow service mounts this volume; it is separate from runner task
worktrees, fixed-action scratch, and the container’s writable layer.
The workspace manager produces a snapshot package and its manifest through a
fixed authenticated runner job. Transfer follows runner → Controller execution
results → Agent result reporting → Workflow’s result reconciler; it requires no
inbound manager listener.
light-workflow retrieves bounded chunks
bound to the feature/task/snapshot identity, checks lengths and digests, and
publishes the contents through its artifact service. The snapshot_transfer.rs
path exists; qualify it through current stage admission and recovery in Phase 1.
A local runner pathname is not a transferable artifact.
The runner and native agents receive neither store credentials nor write access
to the artifact volume. Review/recovery reads use fixed authorized operations
against the stored artifact identity and digest.
The filesystem backend preserves content-addressed, tenant-scoped references and the existing stage/metadata/promote/verified-binding contract. Use durable temporary writes and atomic promotion, verify any existing destination on retry, and recover interrupted promotion before binding a receipt. Clean abandoned staging files and apply retention without deleting referenced evidence. A receipt is usable only after the complete package is durably stored and verified. Feature admission checks that the configured store is available and writable; a disabled store, full volume, or failed export blocks dependent review/stage acceptance without replaying the author turn.
Qualification must recreate the Workflow container and recover after removing the runner’s snapshot refs/object cache. The persistent artifact volume and Workflow database must survive that exercise. Recovery from loss of the volume itself requires a backup of both evidence and metadata; this personal backend does not provide automatic replication or VM failover. The enterprise profile selects its own qualified store while preserving the same evidence contracts.
Phase Implementation And Review
Run each phase before its dependent successor. Codex starts a fresh implementer
conversation for every phase. This is intentional: each feature-implement
run consumes the accepted requirements/design/plan, phase manifest, prior phase
results and snapshots, finding ledger, and validation evidence. It does not
depend on the phase-1 conversation surviving into phase 2. Resume the same
implementer and reviewer conversations within that phase’s fix loop.
- Codex implements the declared scope and returns a summary. The manager captures the candidate checkpoint/snapshot, and fixed validation operations run the phase’s required checks. A check that changes deliverable files requires a new snapshot and validation bound to that candidate.
- Claude reviews the manager-derived delta from the previous accepted phase snapshot (the recorded implementation baseline for phase 1), with access to the accumulated candidate, requirements, design/plan, ledger, and actual test evidence.
- Codex addresses findings and returns a mapping from canonical finding IDs to changes and evidence. Rerun affected checks and resume Claude to verify the updated candidate.
- When closure passes, record the accepted checkpoint/snapshot and next-stage inputs. Publish the concise round summary and update feature status.
Claude reviews every completed phase. Do not wait until all implementation and Codex final review are finished before involving Claude. Judge each phase against its declared exit criteria; deliberately scheduled later-phase work is not a current-phase omission. It cannot excuse a failed current-phase check.
Acceptance requires all declared repository checks and cross-repository contracts. An environment-skipped required test is unqualified, not passed. Phase acceptance records a checkpoint; implementation commit/push happens after final closure. Shared-workspace native tools provide file operations: tests/commands run through fixed manager/test operations, not unsupported native shell access.
Final Review And Fix Verification
For implementation scope, create a complete immutable candidate manifest covering original repository bases, accumulated changes, accepted requirements/design/plan versions, phase results, and validation evidence. Final review checks full-feature behavior, integration, migrations, documentation, and requirements coverage.
- Start a fresh Codex final-review conversation, separate from its implementer. Review the whole candidate once. Codex implementation turns fix its findings; the same Codex reviewer resumes with the before/after snapshot references/digests, manager-derived complete diff, finding IDs, and validation evidence until accepted.
- Start a fresh Claude final-review conversation over the resulting whole candidate. Codex implementation turns address its findings. Resume that Claude reviewer to verify the changed content and close the cited findings.
- Routine final fixes remain inside
feature-finalize. Run affected phase and integration checks, but do not automatically reopen a phase workflow or repeat a separate Claude phase review for a fix Claude is already verifying. - Evaluate the review coverage contract below before commit/push.
Each resumed reviewer focuses on the delta and finding closure while retaining access to the whole current candidate. It may report a new evidenced regression; “delta review” does not mean ignoring an effect outside the changed lines.
Reinvoke the other final reviewer when a change crosses the recorded scope of the active fix, affects multiple phases or cross-repository contracts, changes requirements/design/acceptance criteria, alters security/authorization, migrations or public APIs, or cannot be shown to preserve earlier coverage. The workflow owns this routing decision using declared phase/path/contract mapping, saved snapshot deltas, required-check results, and the active reviewer’s scope assessment. Codex’s fix summary alone cannot declare its own change harmless. Uncertain impact requires the broader review. Reuse the other reviewer’s session when its binding remains valid; a new session is for changed binding, missing history, explicit recovery, or a deliberately restarted design scope.
The publication gate consumes a ReviewCoverage record:
- retain original full-review verdicts and their exact candidate digests;
- record each subsequent candidate transition, its before/after snapshot references and full delta artifact/digest, canonical finding closures, impact decision, checks, and reviewing role;
- explicitly carry forward coverage only for unchanged/unaffected scope;
- require an unbroken accepted transition chain to the current manifest and no open actionable findings without an authorized disposition.
For example, Codex accepts A; Claude fully reviews A and finds F; Codex fixes F to produce B; Claude resumes and accepts A-to-B with the required tests and a confirmed local scope. Publication may accept B using that recorded coverage chain. It must not relabel Codex’s original verdict as a review of B. A broad A-to-B change also requires resumed Codex review before closure.
The chain itself is a proposed higher-level contract. Existing fixed publication validators must be explicitly integrated and qualified to consume it; never reuse an old verdict with a forged current digest to bypass an existing exact-candidate check.
Keep the first final reviews sequential. The current workspace manager takes an exclusive task lock even for read-only review. Parallel final reviews would require independently frozen read-only copies or qualified shared-reader access, plus findings reconciliation; that optimization is outside the pilot.
Finding Identity And Closure
light-workflow owns canonical finding IDs. The reviewer owns the semantic
assessment of whether an observation is new, an existing finding, or a regression
of a previously closed finding. Do not rely on model-generated DESIGN-003
labels or automatic fuzzy text matching for identity.
Allocate the review ID before dispatch and bind it to one logical review turn; retries keep that ID. Canonical finding IDs remain unique for the feature run across stage instances and reopened phases. The model cannot choose a new review identity to evade an existing result or finding mapping.
Every reviewer receives the active ledger and a compact closed-finding history with access to full records. Its schema separates:
existingFindingId: a reference to a canonical ID supplied in that ledger, with disposition/evidence such as still-open, verified-resolved, or reopened;newFindings: turn-local IDs, repository/location, severity, concrete failure, evidence, and required resolution;- proposed
duplicateOfreferences with an explanation when two observations describe the same failure.
The workflow validates scope, referenced IDs, and required fields. It assigns a
canonical ID once for each admitted new item and persists the mapping keyed by
featureRunId + reviewId + localFindingId; replay returns the same mapping.
A validated duplicate relation preserves the old canonical ID and an alias/audit
record rather than deleting history. A cross-feature ID, conflicting disposition,
or ambiguous duplicate claim cannot silently close a finding: return it to the
reviewer for clarification, then use human resolution if disagreement persists.
Codex remediation references canonical IDs and records changed paths/evidence. It may dispute a finding with evidence but cannot waive or self-close it. The reviewing role verifies fixes; an authorized human records waivers/deferrals with reasons and tracking references. A reviewer must explicitly cite a closed ID to reopen the same failure with new evidence.
Stage closure requires a schema-valid accepted verdict/coverage, no unresolved actionable finding without an authorized disposition, required checks passing on the relevant candidate, valid input/policy/base bindings, and any required sign-off. Keep advisory suggestions separate. Empty native text, transport success, or malformed structured output never means “no findings.”
Use three remediation rounds after the initial review as the default, plus
configurable turn and wall-clock limits. Count validation-fix and output-repair
turns against those limits. Budgets count dispatched logical work, not the number
of unique findings: rewording or duplicating a finding cannot reset a budget.
Final remediation consumption survives reopened phases, new stage instances, and
reviewer replacement. Deterministic non-progress indicators include repeated
still-open canonical IDs with no accepted resolution; semantic disputes remain
explicit decisions. Exhaustion or unresolved disagreement enters
HUMAN_RESOLUTION_REQUIRED, never automatic approval.
Conversations, Tasks, And Recovery
Codex and Claude for a feature use the same workspaceId and task ID, with
different thread.sessionRef values. The workflow controls new, resume,
and close. Persist the native codingThread.checkpoint independently of
the file workspace.checkpointDigest; neither authorizes the other.
Keep role conversations through their stage’s fix loop. Close them on stage
completion and verify the close receipt; closeAfterTurn alone is not proof of
closure. The next top-level stage, including the next implementation phase,
uses new conversation identities and consumes accepted artifacts and snapshots.
A review has read-only access to the exact expected candidate and saved delta
versions, and must reread changed files on resume.
A task has one workflow-stage owner and one active manager turn. Stale checkpoints or changed runner/policy/task bindings are rejected. Missing history requires explicit replacement from accepted artifacts. Uncertain native execution stays fenced until reconciliation; changing a request ID, stage, or VM is not permission to replay edits. Completed execution with invalid output needs a bounded output-repair/recovery step against saved state, not a blind rerun.
GitHub Actions And Comment Volume
Codex/Claude write content; fixed workflow actions own GitHub effects. Reuse the
workspace manager’s gh delivery boundary. Pass structured arguments and body
files/stdin; do not interpolate model text into shell source. The current
run.shell path disables network and credentials, and shared-workspace workers
do not receive arbitrary host gh access.
The following are product-level publication outcomes. Implement API calls and reconciliation as definition/tool behavior using generic facilities, not new GitHub operations or provider policies inside Workflow. The former provider consumer migration must account for each required outcome:
| Operation | Eligible lifecycle point | Required content/identity and receipt |
|---|---|---|
| Create feature issue | Intake, before issue identity is established | Agreed requirement artifact, feature/tracking-purpose identity; issue number and URL |
| Create/update status comment | Each active stage and terminal delivery | Feature-wide identity, monotonic version and stored comment ID; confirmed remote body/version |
| Create round comment | Completed review round in its owning stage | Stage execution, round and immutable summary digest; comment ID/URL |
| Publish document revision | Design/plan candidate approved, before stage handoff | Approved snapshot, revision task, recorded base and unique ref; commit, verified ref, PR and file permalink |
| Publish implementation delivery | Final review coverage and checks complete | Approved multi-repository manifest and per-repository commit/ref/PR receipts |
Pin permitted operations and destinations in the stage policy, then enforce the operation’s lifecycle precondition. Merely relaxing the current Finalize-only dispatcher check would allow premature publication. Status and round summaries can describe unaccepted work, but their saved content cannot substitute for candidate approval. New external writes require current authority; uncertain intents retain their identity and reconcile before any further write. Finalize verifies earlier document receipts rather than republishing accepted revisions.
The retiring provider currently implements issue/comment creation and a document Contents API write. Inventory its consumers before removal. Required status edits and document delivery must be migrated to registered API calls and generic Git tools, or explicitly retired with owner agreement. No replacement provider receipt schema or GitHub-specific engine dispatcher is authorized by this design.
Persist immutable author/reviewer/remediation results before publication. Default to one status comment per feature run, edited in place, plus one short comment per completed review round. The status points to the current stage, candidate, open finding IDs, latest design/plan PRs, delivery state, and full artifacts. A round comment summarizes implementation or fixes, validation, review disposition, and links to full results. Limit its prose, for example to 200 words plus links; large finding sets stay in the artifact store.
This preserves each response in durable artifacts and represents it in the issue without pasting full responses repeatedly. Update status after completed author, review, and remediation events; append the compact round comment when the round is complete. Design acceptance and final delivery also publish explicit links. The developer can choose a more verbose comment policy per run.
The workflow owns serialized GitHub action intents and their remote IDs:
- status identity:
featureRunId + status, with monotonically increasing version; - round identity:
featureRunId + stageExecutionId + round, bound to its immutable summary digest; - document/delivery identity: feature, artifact/revision, repository, and action.
Edits target the stored comment ID, never “edit my last comment.” Serialize updates so an older retry cannot overwrite newer status. Record intended body digest/version and reconcile uncertain results by fetching the known ID or enumerating matching markers. Markers aid recovery; they do not replace an atomic local effect claim. A failed comment retries the saved content without another model turn. Required publications must complete before the corresponding stage or feature is declared delivered. GitHub deletion/editing cannot change the internal acceptance ledger.
Commit, Push, PRs, And Completion
For implementation scope, after final review coverage and required checks pass, fixed actions:
- Freeze and verify the accepted manifest, authorized repositories/bases/refs, review coverage chain, and required publication authority. Unexpected changes return to review; a model’s completion text is not publication authority.
- Run declared pre-publication checks. Any deliverable modification from checks or hooks creates a new candidate transition requiring review.
- Commit exactly the accepted files in each changed repository, including intended new files/deletions. Verify the tree and record SHA/parent; no-change repositories are explicit. Exclude unrelated changes, caches, and credentials.
- Push each recorded SHA to its unique GitHub task branch and verify the remote
ref. The normal integration target is a PR to
develop; no direct feature push todevelopormaster. - Create the required PRs and verify head SHAs/target branches. Include the latest accepted document PRs in the delivery manifest. Update the feature issue with per-repository commit, branch, PR, and validation links.
- Complete the declared CI/integration gates.
pr-readyrequires the intended PRs and checks;merged/deployedalso require their authorized integration and runtime verification. Opening a PR alone does not close the feature issue.
Record per-repository commit/push/PR/comment receipts. Multi-repository delivery is not atomic: retain successful effects and reconcile only incomplete actions. A lost push response requires checking the remote SHA, not rerunning Codex or creating another commit. Do not overwrite a conflicting remote ref.
For the pilot, a revised published implementation candidate uses a new immutable publication task/ref/version, just as documents do; record which prior PRs it supersedes and require current review coverage. Transparent advancement of an already-published task ref can be added later with reviewed ref-update handling. It is not required to publish design v2 or to recover a post-publication fix.
For design-only scope, a pinned fixed Finalize stage verifies the accepted design, required sign-off, unchanged retained document contents, confirmed revision PR receipts and the selected completion-target checks. It may then record terminal completion and request the normal generation-fenced VM release without running implementation or the two full-feature final reviews. This exception is permitted only by the scope frozen at intake and a design-only history. It cannot terminate an implementation feature, waive a required reviewer, or publish modified content under an earlier approval. The existing fixed finalization handler is a reusable component; scope enforcement and complete delivery evidence remain to integrate.
Record delivery scope, completion target and satisfied checks in the terminal result and display them in Workflow Admin and the issue summary. A design-only completion does not claim the designed feature has been implemented. Work that later expands scope requires a tracked transition/new run with explicit lineage; never reinterpret the old terminal receipt or reset budgets within a reopened run.
Concurrent Instances And VM Placement
The Phase 4 target allows independent features to remain active concurrently; their phases remain sequential. Each feature gets separate worktrees, branch names, conversations, findings, and budgets even when it changes the same repositories.
The current personal workspace runner requires maximumConcurrency: 1.
With one Codex runner and one Claude runner, Codex can implement feature B while
Claude reviews feature A. light-workflow owns fairness: its dispatch
scheduler persists ready turn intents and chooses eligible feature runs in
round-robin order per required runner, FIFO within a run, with bounded admission.
Persist queue order and the last-served run so restart cannot favor one feature.
workflow_agent_job_t persists Workflow dispatch intents; Agent’s agent_job_t
persists its downstream admission/execution/results. Neither replaces the
Workflow fairness policy. Controller and runner remain authoritative for actual
capacity.
The workflow releases its dispatch reservation only after a terminal/reconciled
execution receipt; a confirmed pre-execution capacity rejection returns work to
the ready queue. A lost admission response requires reconciliation before release.
Waiting for capacity or human input consumes no remediation round, while the
configured overall deadline still applies. Fairness is among workflow-owned
turns; unrelated interactive traffic still competes for real runner capacity.
A second VM can host another independently enrolled Codex/Claude Agent set and private workspace store. Both agents on one VM share that VM’s store; matching workspace names across VMs do not share files. Pin runs to their admitted Agent/runner/store bindings. Moving an uncertain run is explicit recovery after fencing, never copying live native state. Separate installed control stacks must avoid assigning the same issue to competing owners.
Use globally distinct task refs, explicit inter-feature dependencies, and separate test resources (or serialized access to shared resources). Local locks cannot coordinate remote merges across VMs. Recheck target-base movement and revalidate/review affected changes before integration. Another VM provides execution capacity; it does not remove integration or upstream account limits.
The first pilot enforces one active feature per VM through a light-workflow
admission reservation retained across stage handoffs and uncertain execution.
Another separately operated VM can host its own pilot and Agent set, but shared
scheduling, coordinated multi-VM routing, and automatic failover are not qualified
by that arrangement.
They remain later milestones, not reasons to postpone the standalone design loop.
Pilot VM Slot Release
light-workflow persists the reservation’s VM/runner binding, feature owner,
generation, acquisition time, and release status. Workflow Admin shows the
holder with its issue, stage/state, waiting reason, outstanding execution, and
available resume/cancel actions. READY_FOR_NEXT_STAGE and
HUMAN_RESOLUTION_REQUIRED retain the slot until the operator resumes or cancels
the feature; they cannot hide an indefinite reservation from the operator.
On feature COMPLETED, CANCELLED, or FAILED, release the slot idempotently
once all admitted turns/fixed effects have terminal or confirmed fencing
receipts. Record the release against the same owner/generation in the feature
store. Completing an individual stage does not release it. A timeout, missing
heartbeat, or terminal stage/process label alone is insufficient proof that the
VM is safe to reuse.
Provide an authorized Cancel and release VM action for a stalled feature,
bound to its expected feature version and reservation generation. It blocks new
dispatch and successor claims, cancels queued work, requests cancellation of
active work, and fences outstanding worker/effect generations. Wait for
Controller/runner/manager confirmation that old execution cannot continue, and
reconcile already-issued external effects before completing cancellation and
releasing the slot. Clearing a database owner field alone is not fencing.
If confirmation is unavailable, show
VM_RELEASE_PENDING and the unresolved execution; the VM remains unavailable
until an authorized stop/reconciliation supplies the missing evidence.
Persist the action intent and release receipt so a lost response can be retried without freeing a newer owner’s reservation. Late old-generation dispatches, results, and publication actions cannot reacquire the slot or advance the cancelled feature. Retain its artifacts, findings, budgets, and publication receipts; any later explicit reopen must reacquire a slot and stage claim.
State And Evidence
| State or substate | Required transition evidence |
|---|---|
ISSUE_PENDING | Reserved intake identity, agreed requirement artifact and unresolved issue-creation intent; successor admission blocked |
REQUIREMENTS_FROZEN | Accepted requirement version, pinned delivery scope/target and confirmed issue reference |
DESIGN_ACTIVE / PLAN_ACTIVE / PHASE_ACTIVE | Candidate, validation, review, and bounded remediation receipts |
DESIGN_SIGNOFF_PENDING | Exact-digest human decision when enabled |
DOCUMENT_PUBLICATION_PENDING | Immutable candidate approval and pending revision task/effects; confirmed commit, ref, PR and permalink receipts required to leave this state |
READY_FOR_NEXT_STAGE | Durable accepted StageResult and snapshots; successor claim/start transaction validates the current feature/input versions |
FINAL_REVIEW | Primary/independent reviewer substate, finding ledger, and current review coverage chain |
COMMIT_PENDING / PUSH_PENDING / PR_PENDING | Verified per-repository publication receipts |
POST_PUBLICATION_VALIDATION | Required completion-target checks |
REPLAN_REQUIRED | Revised inputs, affected-scope decision, and explicit acceptance invalidation |
HUMAN_RESOLUTION_REQUIRED | Dispute, exhausted budget, sign-off rejection, or ambiguous recovery decision |
REAUTHORIZATION_REQUIRED | Explicit recovery of current user/run/action authority with policy recheck and preserved budgets/fences |
VM_RELEASE_PENDING | Cancellation/release intent and outstanding stop/fencing receipts; reservation remains held |
COMPLETED / CANCELLED / FAILED | Terminal feature result with delivery scope/target and separate VM release evidence; a successful stage alone is not COMPLETED |
These are conceptual lifecycle states/substates. Some are represented by separate
task/effect records rather than the current FeatureState enum. New persistence
and wire mappings require contract changes and gates; this table is not a list of
already-implemented enum values.
The workflow and artifact store retain feature versions/owners and claim receipts, input versions, snapshot contents/deltas, finding mappings, review coverage, tests, budgets, sessions, VM reservation/release receipts, queue state, and external receipts. Restart resumes this state. Post-publication failures retain the published history and require a tracked remediation/version; they do not rewrite old acceptance or reset the feature budget.
Current Implementation Boundary
Source reviewed September 29, 2026 against light-fabric a67d424,
portal-view 0a1efba, and light-portal 81c8d3aee. This is a source baseline,
not deployed-image, database, effective-configuration or runtime verification.
Historical test/live evidence remains in the
Phase 1 progress record; its
September 14–15 results must not be carried forward as qualification of the
later native MCP integration. No application/runtime gates were rerun for this
design revision.
Workflow source filenames below refer to apps/light-workflow/src in
light-fabric; other paths and repositories are identified explicitly.
| Capability | Current source support | Remaining integration or correction | Qualification boundary |
|---|---|---|---|
| Async and sync invocation | rule_api.rs implements native workflow_start; invoke_api.rs implements binding-scoped synchronous Invoke; Portal start synchronizes definitions/grants | Connect development admission to the current native path; preserve the distinct execution/credential contracts | Ordinary Editor/Tool behavior does not qualify development stages |
| Atomic stage claims | apps/light-workflow/src/development_store.rs persists features, VM reservations and atomic claims with replay/version checks | Native start passes no artifact store; its development branch also labels new claims Replay, bypassing the Accepted branch that registers LONG authority | Storage evidence exists historically; native admission regression and runtime gates required |
| Runtime and Agent ownership | native_jobs.rs, job_authorization.rs, agent_job.rs and Agent domain.rs implement Workflow-owned intents and Agent polling/reporting into separate stores | Qualify current authority, exact Agent bindings, transport/result recovery and cancellation | Do not restore obsolete cross-database catalog/job access |
| Personal turns and capacity | Shared task inspect/implement/read-only review, session controls and exclusive task locks exist | Integrate bounded stage rounds and current runner policy; shared fairness remains Phase 4 | Prior shared-session evidence does not qualify the whole feature lifecycle |
| Snapshots and artifact storage | crates/task-workspace/src/snapshot.rs, snapshot_transfer.rs and artifact_store.rs implement retained trees, package/delta reads, transfer and filesystem/S3 backends; workflow.yml accepts backend/root settings | Carry the store into native admission; qualify effective settings, persistent volumes, failure and recovery on each required stack | Historical local evidence exists; complete current-path and installer gates remain required |
| Feature operations and UI | Acceptance/replan/publication/finalize HTTP handlers exist; native feature/process list/get/cancel and Portal ProcessInfo.tsx VM controls exist | Add missing native MCP feature transitions and connect existing UI controls to them | Generic UI or HTTP-only diagnostics do not qualify a full feature run |
| Human sign-off | Generic native human-task claim/complete operations exist; development_handoff.rs validates persisted sign-off during acceptance | Produce digest-bound development sign-off evidence and route all three decisions; qualify stale/revoked/duplicate cases | A consumer of DesignSignoff is not an implemented producer |
| Existing publication consumers (retirement inventory) | publication_dispatch.rs and the retiring provider have existing consumers; task-workspace/src/delivery.rs supplies Git primitives | Migrate required effects to definition-driven registered APIs/generic tools or explicitly retire unused consumers; do not extend the provider/engine with GitHub operations | Historical mocked-provider evidence is not qualification of the direct API replacement |
| Completion and VM release | development_finalize.rs has a restricted reviewed-design terminal path; acceptance/cancellation use reservation fencing | Pin explicit delivery scope, verify complete delivery receipts, expose finalization through MCP, and qualify interrupted cleanup/late-generation cases | Existing terminal/storage fixtures do not prove full-feature delivery or safe live VM reuse |
| Feature rules and shell boundary | development-workflow-contract provides typed review/finding/budget/handoff rules; run.shell admission forbids network/credentials | Extend contracts for revised lifecycle semantics and integrate them into fixed actions | Phase 0 evidence covers the recorded rules only; new rules require new contract tests |
| Child workflows | executor.rs rejects run.workflow; native Invoke rejects nested parent-action admission | Durable child composition and shared multi-VM scheduling remain later work | Synchronous Tool support does not qualify parent/child orchestration |
Revision Record — September 29, 2026
This document remains the authoritative design. The following stable review IDs map the source review to the contracts above and the additional Phase 1 gates below. They identify design-review findings, not runtime finding-ledger IDs. The separate Phase 1 completion implementation plan is unchanged by this revision and must be reconciled after design review; none of these items is marked closed merely because its intended behavior is now documented.
| Review ID | Design disposition | Implementation or evidence follow-up |
|---|---|---|
| DW-R01 | Async MCP stage admission uses one atomic feature/run path and distinguishes creation from replay | Supply artifact access, establish LONG authority exactly once, and test both through native MCP |
| DW-R02 | Fixed effects have operation-specific lifecycle gates; candidate approval precedes publication and stage completion | Add intake issue-pending support, ordered status updates and document task/ref/PR delivery with exact approved bytes |
| DW-R03 | Feature transitions use native MCP; design sign-off requires an authorized digest-bound producer | Expose missing transitions, extend existing UI, and integrate human decisions with acceptance |
| DW-R04 | Workflow and Agent own separate runtime/job stores with authenticated transport | Qualify current Agent admission/result/fencing paths without restoring database bridges |
| DW-R05 | Implemented components, remaining work and qualification are recorded separately | Reuse stores/snapshots/transport; refresh evidence against exact images, configuration and migrations |
| DW-R06 | Current user/application and LONG/Invoke authority contracts replace retired broker prerequisites | Test identity, revocation, expiry/recovery and retained deadlines/budgets on the selected stack |
| DW-R07 | Delivery scope is explicit; design-only finalization cannot satisfy implementation completion | Extend scope contracts, pin them at intake and verify scope-specific terminal receipts and VM release |
Delivery Plan And Acceptance Gates
Phase 0: Contracts And Deterministic Review Rules
Implemented and qualified by bash scripts/run-development-workflow-phase0-gates.sh.
Wire fixtures and worker JSON Schemas live in contracts/development-workflow/v1.
See the qualification record for
the exact boundary between pure rules and Phase 1 runtime enforcement.
Define FeatureRun, stage-claim identity/version rules, StageResult,
CandidateSnapshot/delta receipts, finding/remediation schemas, review coverage,
budgets, VM reservation/release receipts, document/publication revisions, and
comment intent identities.
Exit gate: fixtures prove that replay preserves canonical finding IDs; reworded findings citing an existing ID retain identity; invalid/ambiguous duplicate references cannot close a finding; disputed findings require reviewer/human disposition; and reopened phases/new instances retain consumed budgets. Also prove a local final fix can close with resumed review and current coverage, while a broad change or missing transition cannot publish. Test optional design sign-off and stale sign-off rejection. Define fixtures for duplicate/stale handoffs and missing snapshot evidence; Phase 1 must enforce these contracts in the stores and start path. These gates require no model calls.
Phase 1: Standalone Stage Execution And Fixed Actions
See the Phase 1 implementation progress for verified slices and remaining runtime gates. Phase 1 is not yet complete.
Prerequisite: qualify the shared authorization foundation as used by the current native MCP and LONG paths described above. Include current user/app identity, revocation and authority recovery; historical broker-only or ordinary Editor/Tool results do not satisfy development-stage gates.
Integrate and live-qualify the Workflow-owned Agent job transport and
exact workflow Agent bindings. Each separate workflow Agent instance needs its
own service ID in codingProfile.workspaceBindings[].agents and the matching
runner-local RunnerWorkspaceConfig.bindings[].agents. Publish and install the
same complete binding, including its authorization/membership revisions,
subjects, intents, runner, host, and environment: the runner compares the full
binding for equality. Authorize the separate workflow Agent service identity
explicitly; do not reuse the interactive Agent’s identity or weaken that
comparison. Reuse the implemented components and complete their integration:
- the durable
FeatureRun/StageResultstore and atomic native MCP stage admission, including version/owner checks, dispatch enforcement, and acceptance/replan transitions; wire operator starts and feature mutations through Gateway MCP; - manager-owned candidate trees, private retention refs, immutable content artifacts, and fixed delta/version reads tied to checkpoint receipts;
- the personal filesystem artifact backend, persistent volume and published
settings in
portal-config-loc/all-in-ltandlight-portal-install, readiness checks, and Workflow-owned export/recovery through fixed manager reads; - bounded stage-local round slots, schema validation, native/file checkpoints, cancellation, and restart recovery;
- fixed tests, issue creation, status/round comments, and immutable document publication tasks/PRs;
- one active feature per pilot VM, Workflow Admin holder visibility, terminal release, and authorized cancel/release with confirmed execution fencing.
Contract additions in this revision, including delivery scope, candidate approval and issue-pending intake, need explicit schemas, persistence and deterministic tests before integration. They are not covered by the historical Phase 0 result.
Do not implement run.workflow for this gate. Exit gates:
- Each workflow Agent service ID succeeds with matching published and runner-local workspace bindings. Missing IDs on either side, mismatched binding revisions, and an unauthorized interactive Agent identity are rejected before a native turn starts. Qualify the separate instances using their actual service IDs.
- A top-level design stage performs Codex authoring, fixed validation, Claude review, and a resumed fix round using a manager-derived before/after delta. After later edits and a restart, the earlier contents still reconstruct exactly, including new/deleted/binary files and executable modes. Exercise artifact recovery without the local snapshot refs. Missing/corrupt evidence blocks review/acceptance; the real index, HEAD, and task branch stay unchanged.
- Each personal stack variant publishes and restores evidence with its configured filesystem store after Workflow container recreation and removal of the local snapshot refs/object cache. Test disabled/unwritable/full storage, interrupted export and promotion, and retry without duplicate artifacts or another author turn. The runner receives no artifact-store credentials or writable store mount.
- Transactional store/start tests prove simultaneous identical successor starts return one instance; changed inputs conflict; stale/superseded inputs and unclaimed dispatch cannot run; commit followed by a lost response replays the same instance; rollback leaves no orphan claim or runnable process. Acceptance enables the next valid stage and preserves feature budgets and recovery fences.
- Completing a feature frees its VM slot. Cancelling a stalled feature in
READY_FOR_NEXT_STAGEorHUMAN_RESOLUTION_REQUIREDfrees the slot for another feature. An uncertain turn keeps release pending until confirmed fencing; afterwards a new feature can claim the slot. A replayed release or late old execution cannot free the new reservation or publish old results. - Restart or lost result delivery repeats neither an uncertain model turn nor a GitHub effect. Durable handoffs are enforced, not left to operator discipline.
The September 29 review adds the following boundary cases to these exit gates:
- DW-R01: Through Gateway/native MCP, a real development definition admits with the configured artifact store and creates one claim/process/authority binding. Concurrent starts, restart and lost responses replay that identity. Disabled storage or denied authority creates no runnable work; uncertain registration reconciles without duplicate authority. Historical completed receipts remain readable without reacquiring a VM or starting new work.
- DW-R02: Intake creates or adopts an issue; design publishes its approved revision before successor admission; status and round effects run at their declared lifecycle points. Delayed status retries cannot overwrite newer text. Verify all delivered document bytes/modes, commit/ref/PR/permalink receipts and restart reconciliation. No marker injection, duplicate issue/comment/PR, or Finalize-only workaround may satisfy this gate. Optional plan publication adopts the same contract when its stage is added in Phase 3.
- DW-R03: An operator drives stage start, sign-off, acceptance/replan and design-only finalization through existing UI surfaces backed by Gateway MCP. Approve/request-changes/reject each produces the correct durable transition. Changed digests, stale assignment/feature versions, revoked eligibility and changed duplicate decisions fail closed. A recorded human decision is not displayed as completed execution before continuation is confirmed.
- DW-R04 / DW-R06: Current user-only and user-plus-app Gateway requests follow the identity contract; missing user, invalid supplied app, wrong Host/owner and unauthorized Agent bindings are refused. Exercise Agent polling/results and revocation across the separate stores. Async LONG renewal/recovery preserves deadlines, budgets and fences; synchronous Invoke never becomes a substitute for async stage admission or receives a LONG registration.
- DW-R07: A scope-pinned design-only run completes only after required document delivery/checks and releases its VM only after confirmed fencing. Implementation scope cannot use the review-free design terminal path, even with an otherwise valid definition. Reject scope substitution on retry and preserve the original terminal receipt.
For DW-R05, retain a qualification matrix identifying each gate’s repository revisions, deployed images, effective configuration/definition digests, migration baseline, fixture/run IDs and evidence. Report component tests, disposable database gates, mocked-provider tests and live application results separately, including executed/skipped/failure/error counts. Required skipped or blocked checks remain unqualified. Update the implementation plan from this design only after design review; do not infer Phase 1 completion from the component table.
Phase 2: Personal Design Pilot
Exercise both intake paths and operator handoff to standalone feature-design.
Publish accepted design v1, then revise it through the same review/sign-off
contract and publish v2 on a new task/ref/PR with an updated issue link.
Exit gate: both immutable permalinks remain valid, the ledger selects v2, no existing remote ref is overwritten, and the Phase 1 claim API rejects a new next-stage start with stale v1 inputs. Status edits cannot regress on retry; round comments stay concise and link to all stored author/reviewer/remediation results. A design-only run then completes through its fixed Finalize contract; an implementation run remains ready for its next stage. This is the first complete personal design pilot.
Phase 3: Optional Plan, Phase Work, And Final Delivery
Add the optional plan and repeat standalone implementation by phase. Implement initial full final reviews plus resumed fix verification, explicit review escalation, commit/push/PR delivery, and completion checks.
Exit gate: implementation phase 2 cannot start before phase 1 passes and starts with a fresh Codex conversation reconstructed from accepted artifacts/snapshots; the no-plan path creates no duplicate plan; a Claude final finding is fixed and verified without unconditionally restarting both reviews; a broad fix invokes the other reviewer. A multi-repository partial push failure reconciles receipts without duplicate commits/PRs. Latest document revisions and code PRs agree on delivered content.
Phase 4: Shared Capacity, VMs, And Optional Parent Coordinator
Qualify workflow-owned fairness and two independently enrolled VM agent sets.
Then add feature-delivery and durable child orchestration if automated
composition is needed, preserving the standalone contracts.
Exit gate: two features retain isolated results and refs; capacity waiting and restart preserve fairness; cancellation affects only its feature. Parent/child start, result, cancellation, and recovery are qualified separately and do not change review closure or budget semantics.
Revision Record — September 30, 2026
- Replaced the capture subsystem with direct registered API calls and ordinary context data; reused demo3 setup as the required baseline.
- Withdrew the capture-policy configuration framework, receipt UI/admission, provider extension and additional creator/membership authorization requirements.
- Planned retirement of light-github-action-provider after consumer migration.
- Kept independent #425 and generic P01 waits; planned removal of P02 capture machinery with its generic fixes retained separately and deployed migration history preserved. No rollback executed in this revision.
- Owner reports P02 redeployment and light-portal-test
make allsuccess; this remains owner-reported general regression evidence, not source/image provenance or proof of the proposed direct API flow.
References
- Workflow Invoke And Tool Binding Publication
- User, Application, And Workflow Authorization
- Phase 1 Implementation Progress
- Enterprise Development Workflow Orchestration
- Shared Native Coding Sessions
- Shared Task Workspaces
- Chat And Workflow Integration
- Workflow Coding Thread Lifecycle
- Coding Harness Integration
Development Workflow Phase 0 Qualification
Phase 0 of issue #392 is
implemented in crates/development-workflow-contract. The parent
personal workflow design remains the
lifecycle specification. This milestone makes no model calls and does not start
or deploy the personal workflow pilot.
Implementation
The new crate defines requirements and phase manifests, feature/stage identities, claims and results, snapshots/deltas, findings/remediation, review coverage, budgets, sign-off, VM reservation/release, document/publication revisions and comment intents. Strict Serde contracts reject unknown worker fields. Checked-in JSON Schemas cover review and remediation outputs.
Canonical finding IDs are deterministic, domain-separated SHA-256 identities over the feature, preallocated review and local finding ID. Replay returns the saved mapping, with a digest conflict for altered output. Alias history is kept; ambiguous/cross-feature references and simultaneous alias/closure are rejected without partially mutating the ledger. Implementer dispute evidence cannot close a finding; reviewer or authorized human disposition is required.
Coverage preserves the original full-review digests and a contiguous chain of saved candidate transitions. Codex may fix and resume before Claude starts its first full review. A later local Claude-verified fix can carry Codex’s unaffected coverage explicitly. Broad or uncertain changes require every already-started final reviewer; subsequent full review covers a reviewer not yet started. Changed sessions, missing snapshots/deltas, skipped checks and forged verdicts block publication.
Claim preflight checks historical replay before current feature version/owner; a changed request conflicts. Budgets charge logical dispatch once and retain consumption across replacement/reopened stages. Snapshot/result acceptance checks current inputs/claim and persisted reviewer evidence. VM release needs matching owner/generation, terminal state and confirmed execution/effect evidence.
Verification
Run from the repository root:
bash scripts/run-development-workflow-phase0-gates.sh
The gate runs formatting, 20 deterministic tests, strict Clippy and whitespace checks. Tests include:
- review replay and serialization, explicit rewording/duplicates, ambiguous references, disputes, reopen and non-progress without retry inflation;
- local/broad final fixes, Codex fixes before Claude full review, missing coverage, invalid snapshots/deltas, skipped checks and replacement-session rejection;
- optional/exact-digest sign-off, stale/rejected approval and cumulative budgets;
- matching JSON handoff fixtures, duplicate historical claims, stale/changed inputs and conflicting stage state;
- stage acceptance with persisted review evidence, missing durable snapshot, fenced VM release, immutable publication tasks and monotonic status comments;
- JSON Schema/Rust parsing of worker fixtures and rejection of unknown fields.
mdbook build docs also passes. No required test is environment-skipped.
Phase 1 Enforcement Boundary
These are pure reference rules and typed receipts, not a database implementation or an authentication boundary. Phase 1 must verify artifact contents and receipt provenance; load trusted review allocations, policies, clocks and budget scopes; and persist state atomically under feature/effect locks. A worker must never supply its own accepted ledger or scope-impact decision.
Claim, invocation/process/initial task and ownership must commit together. Tests of simultaneous starts, rollback and lost responses belong to that store path. Dispatch/effects must require the claim. Snapshots need durable export and fixed runner → Controller → Workflow transfer; native conversation history is separate. Phase 1 also implements the filesystem artifact backend, actual workflow Agent jobs, sign-off controls, fixed GitHub actions and confirmed VM fencing/release.
No current Workflow listener or publication action is wired to the new crate in Phase 0. Do not infer runtime enforcement or pilot readiness from these tests.
Phase 1 implementation progress
Status: in progress, not qualified for the personal pilot. This is a partial implementation record for issue #392, not a Phase 1 completion announcement.
The 2026-09-14 credential-broker checkpoints below are historical. Workflow Invoke later retired the enrollment routes and broker grant flow. Current Gateway-to-Workflow invocation uses the acting user’s bearer. Asynchronous
workflow_startruns can register LONG for work that outlives that bearer; synchronousworkflow_invokedoes not. See Workflow Invoke.
Live continuation update (2026-09-15)
Latest continuation: historical UNKNOWN retirement is qualified (see the operator
retirement section). The dedicated Claude runner now uses a private, stable copy
of the already-qualified 2.1.269 binary, verified against the canonical SHA-256
pin; the auto-updated host installation and published CLI contract are unchanged.
Workflow review turns request schema-constrained Claude output and retain the
original raw answer alongside reviewResult. The CLI-facing schema uses Draft 7,
which the pinned CLI accepts; the canonical Draft 2020-12 schema and Workflow’s
binding/finding-ledger validation remain unchanged. Ordinary conversations are
not forced into the Workflow result schema.
Worker tests: 75 passed, with the host bubblewrap isolation test and no-credential exact-CLI schema preflight additionally passing. Workflow’s native structured result regression checks raw-answer preservation and rejection of substituted bindings. Installer tests: 32 passed. The provider-container lost-response gate again passed with exactly three mocked issue/comment/document writes.
Fresh feature 322ae203-36bd-475b-8cb0-0b1edaca60d4 completed intake and reached
the design fix/re-review cycle. Workflow was recreated during that run using the
installer publication override and a separate local qualification image based on
2.3.5-dev.20260909.2338; release tags and existing data volumes were preserved.
The provider now runs separately with the persistent
all-in-lt_github-publication-state volume rather than a manually launched process
in Workflow’s container layer. Final live publication/release and full installer
storage-failure/cache-loss qualification are still pending.
The owner-authenticated dispatcher is now connected to fixed GitHub issue, comment and immutable-document publication routes. It selects policy and source files from the pinned Finalize definition and verified retained candidate, commits the effect intent before I/O, and retains provider verification with confirmation in one transaction. Acceptance verifies every required publication slot rather than trusting a caller-supplied receipt list. The fixed reviewed-design terminal handoff uses the normal acceptance and generation-fenced VM release transaction. Gateway publication, finalization and cancellation routes have targeted tests.
The provider’s lost-response test commits a remote issue but returns an error, reopens its SQLite journal, reconciles the receipt and verifies exactly one write. Opt-in Compose wiring now gives both local and installer stacks a persistent provider journal and loopback-only endpoint; full installer qualification remains open.
Provider-only container qualification subsequently passed using the built
light-github-action-provider:phase1-local image and
light-portal-install/tests/publication-provider-container-gate.mjs: mocked
GitHub commits each issue/comment/document effect but returns an error; container
recreation preserves the intent, status observes the receipt, exact replay returns
it and changed input is rejected. Exactly three writes occurred across the three
destinations. This is not application-level FeatureRun or full installer proof.
The installer base-plus-publication Compose override validates successfully.
Current checks pass: Workflow 94 unit tests, runner 53 unit tests, gateway
lifecycle route/schema test, extended provider lost-response test, and the fresh
PostgreSQL stage-store gate (phase1_review_terminal_1789498244909).
The earlier qualification incident started a new feature after owner cancellation released the
expired prior run. The new run completed author/snapshot work but stalled in
Claude intake review: the pinned CLI 2.1.269 no longer exists, while the installed
CLI is 2.1.272. The runner returned UNKNOWN with failed cleanup because no durable
native process identity exists. The run is not replayed and its VM remains held.
Admission now revalidates CLI paths before staging or starting, preserving a
durable NotRequired-cleanup rejection for this failure on future attempts. It does
not retroactively rewrite the uncertain attempt’s evidence.
During that incident, the local runner configuration pointed to the installed real 2.1.272 CLI
path, not a mutable symlink. The running runner and uncertain execution are not
restarted or automatically replayed by this configuration edit.
Current qualification is not complete: no live dispatcher issue/document/ comment receipt or new terminal VM-release receipt has been obtained, and installer application-level failure/recreation qualification remains outstanding. The earlier paragraphs below describe preceding contract/journal-only milestones, not the current dispatcher wiring status.
The connected PostgreSQL finalization qualification now passes on fresh database
phase1_review_terminal_1789498725303. It accepts a reviewed retained design,
starts a pinned fixed Finalize claim, rejects missing and uncertain publication,
reconciles the lost provider response without a second write, verifies required
publication evidence, finalizes the design and releases the VM. Exact finalization
replay succeeds; changed predecessor version and nil operation ID fail. After a
replacement feature reserves the next generation, old finalization replay cannot
release its VM. These are application-core storage fixtures, not live native runs.
This gate found and fixed two production gaps: the claim-derived finalization snapshot ID exceeded the manager ID bound, and acceptance always required a new review even for a fixed already-reviewed design. The bounded snapshot ID is now deterministic. Review-free finalization requires a pinned reviewed-design terminal definition, reviewed design-only history, unchanged retained repository contents/ trees/modes/checkpoint and the sole design package output. Declared reviewers and implementation finalization cannot use the exception. Contract regressions and all 94 Workflow unit tests pass; the musl Workflow/Gateway release builds pass.
Publication persistence now has a PostgreSQL implementation in
apps/light-workflow/src/publication_journal.rs, reusing workflow_task_effect_t.
Unlike the generic task-effect claim helper, its insert result distinguishes
the first writer (Dispatch) from an unresolved retry (Reconcile). Confirmed
results are immutable. The caller must authenticate and validate the active
stage, pinned policy and retained content in the same transaction, commit before
dispatch, and retain provider verification evidence before confirmation.
These are storage primitives, not a connected publication dispatcher.
The fresh-database stage-store gate passed with the new journal tests: rollback
before dispatch, concurrent first claims (exactly one dispatch), lost-response
reconciliation, changed request rejection, confirmation rollback, immutable
confirmation replay and conflicting/null result rejection. Evidence:
/tmp/phase1-design-cycle.R5XWTj/phase1_review_terminal_1789495941563.log;
disposable database phase1_review_terminal_1789495941563 was retained.
No live publication, feature finalization, or VM-release mutation was performed.
Publication preparation now has a typed contract in
crates/development-workflow-contract/src/publication.rs for issue, comment and
document destinations. Trusted policy pins repositories, document branches and
paths; issue/comment capabilities are independent. Plans bind the accepted
candidate and retained content reference. A feature/slot/revision effect identity
does not change when content or destination changes, so altered retry input
conflicts instead of allocating a duplicate write. Prepared effects can dispatch
once; in-flight/unknown effects require remote reconciliation, and confirmation
is immutable. Four new tests and all 20 existing contract tests pass.
This is a contract-only component, not a connected dispatcher or live GitHub
qualification. The host must still verify retained bytes and feature ownership,
persist transitions atomically, and validate remote evidence. No publication,
feature finalization, or VM-release mutation was performed in this slice.
Separate intake and design author/review/fix qualification passed. Feature
0bc64472-861a-4323-a8ba-4ac6771b6d6e completed intake invocation
01a0a634-f881-7352-b4ae-9fc67ff2a58e and design invocation
01a0a636-253c-79e0-bad9-921168964175. Both were accepted through the authenticated
Gateway stage API, reaching feature version 5, ready for finalize.
The first design review rejected one blocking finding; resumed Codex remediation
changed the retained candidate, and resumed Claude independently marked canonical
finding sha256:3f415d158436b2a622d71eae159a7c5dccf971f487893a458a75e19439414b64
verified-resolved. The final review accepted and fixed validation had no failures.
Before package digest: sha256:66d5281d3611de6ff91b7d4a03b1de556f10df2fa94bdea85ff6dbcd834c1b8d;
accepted package digest: sha256:7f6667d8e3e3577d296d8a2a3ecead98340cb8b1051e899cbff68a6047455695.
All 15 Agent jobs succeeded with CONFIRMED native cleanup. No native work remains
running, but VM personal is still reserved to this feature at generation 30 for
finalization. Cancelling completed stage IDs does not cancel a ready successor;
do not describe this reservation as released or reset it through SQL.
Current fixture catalog is catalog-v1.0.7.json: intake 1.0.6 and design 1.0.7.
The latter binds fix/re-review thread.expectedCheckpoint to the saved author/
reviewer native receipts. The previous fixture omitted those fields and was
rejected before fix dispatch. The standalone validator now checks both bindings.
Gateway snapshot 01a0a633-a55e-7d23-9858-e03fc54cc27b is active, permissions
unchanged. Detailed local report: /tmp/phase1-design-cycle.R5XWTj/cycle-report.json.
The diagnostic driver is not yet a durable daily-test implementation. Publication,
finalization/VM release and installer qualification remain open; Phase 1 is not
complete and no phase-completion GitHub comment has been posted.
Malformed native review delivery is now terminal: after authenticated peer,
subject, cleanup, claim and execution-fence checks, invalid review JSON/binding
or finding-ledger output commits the immutable original report and consumed turn,
sets the Workflow job FAILED with NATIVE_REVIEW_OUTPUT_INVALID, and creates no
approval receipt. Exact report retries return acknowledgement after cancellation;
conflicting reports remain rejected. Database failures remain retryable transport
errors. There is no automatic native model replay or arbitrary prose extraction.
The real disposable PostgreSQL gate covers preamble, substituted binding, unknown
finding reference and valid review, including report replay/conflict and identity
guards; it passed, along with all 94 Workflow unit tests. The local Workflow
container binary is updated, with backup in
/tmp/phase1-design-cycle.R5XWTj/terminal-review-deploy/. The successful live cycle
above uses this binary; malformed-output failure cases were exercised in the
disposable PostgreSQL gate, not induced in the successful model run. This does
not close Phase 1 or qualify a rebuilt image.
Fresh-workspace admission and provisioning now compare the actual file checkpoint, not the optional saved review checkpoint. Execution-time checks remain in place. The workspace suite passed 31 tests (2 ignored); worker tests passed 74 (2 ignored) on a serial rerun. Both native worker binaries and their Controller admission digests were updated locally without changing owner, repository or intent grants.
Workflow supplies the exact review binding in verified review material. The durable Agent bridge consumes and cross-checks that Workflow-only envelope before strict interactive-message parsing; interactive callers cannot supply it. Agent tests passed 46, Workflow tests passed 94, and the disposable PostgreSQL gate passed, including the stored review-material binding assertion. Review parsing accepts raw JSON or one exact JSON fence, while retaining strict schema and binding checks and preserving the original answer. General malformed-review report retry handling was subsequently hardened as described above.
Intake invocation 01a0a5c9-8187-71c1-9c1a-f6ecf5a7f337 completed native authoring,
retained snapshot transfer, fixed checks and accepted Claude review. Stage handoff
then failed because Gateway lacked workflow_get_feature@call access mapping.
Normal cancellation released the VM at generation 27; no SQL state reset was used.
Two owner-only mappings (workflow_get_feature@call, workflow_accept_stage@call)
were added through the versioned Portal command, preserving all existing entries.
Snapshot 01a0a5d4-7762-7713-8b3b-f66c81aa01f6 was activated and Gateway reloaded;
authenticated feature read succeeded. Subsequent cycle results are recorded above.
Intake 01a0a5d5-06f4-7ec1-bd1f-021e106de9ea subsequently passed authenticated
stage acceptance. Design admission exposed a 70-task fixture exceeding the
runtime’s 64-task bound; design 1.0.5 reduces bounded snapshot slots to 62 tasks.
Its author then stopped before a model call because intake and design shared a
native session UUID. Cleanup was confirmed and the VM released at generation 28.
At that point cycle fixtures were 1.0.6: 1.0.3 isolates native request IDs by stage,
1.0.4 supplies stage-specific review criteria and the strict result contract,
and 1.0.6 keys native sessions by stage transition UUID while preserving same-stage
resume. The standalone validator now checks this session isolation and the
64-task bound; both latest definitions pass. Gateway snapshot
01a0a623-c2c8-7994-b531-bfc28748d5f5 was activated, with permissions unchanged.
The fresh 1.0.6 intake 01a0a625-eb42-7012-b4a7-eaed03158f41 completed authoring
and snapshot transfer, but Claude returned a prose preamble before its review
JSON. Strict parsing rejected the result; the completed Agent job’s report
remained PENDING and retried. Normal owner cancellation ended the qualification
and released VM personal at generation 29. This failure motivated the terminal
handling fix above; no answer rewriting or automatic native turn replay was used.
The subsequent successful run is recorded above.
Imported historical files remain unchanged. Local Agent/Workflow/Gateway binary
overrides are container-layer deployments, not rebuilt release images. Publication
and installer qualification remain open; the design cycle subsequently passed.
The historical checkpoints below describe earlier slices, not current gates.
Earlier live continuation: membership aligned, checkpoint gate pending
The authenticated Gateway lifecycle now exposes workflow_get_feature and
workflow_accept_stage through the existing owner/user/service-authenticated
Workflow channel. It does not permit direct SQL handoffs or worker-selected
identity headers. Routes validate UUIDs and bounded acceptance bodies; Workflow
still verifies owner, current claims, version, candidate bytes, fixed checks,
review ledger and pinned successors. Gateway tests: 451 passed, 5 ignored;
Workflow unit tests: 94 passed. Both local musl binaries are deployed in their
container layers and must be included in the next image build.
Resumed review can name reviewArtifacts.previousReviewTask. Workflow supplies
the saved immutable review receipt and canonical finding IDs only when its
feature, stage, reviewer and before-candidate match. Caller-provided evidence
does not replace that ledger receipt. This path still requires live success.
Separate intake/design definitions and owner-only Tools were imported and
published from light-portal-event/workflow/20260915-design-cycle/. The corrected
intake invocation 01a0a59c-bc26-78b2-8b76-5f1e025a435a reached the native worker
but was rejected before a model call: the registered repository catalog hashes
to sha256:6869f6609d0563e3412bad1de2633c9e4f226a7ba3d1acff6b02d8874d2152fc,
while the then-published Agent and runner bindings pinned a different digest.
After operator approval, both Workflow policies were republished, both runner
workspace pins aligned, and both Agents reloaded. Owner, grants and authorization
revisions are unchanged. Runner-generated admission documents match the Controller
pins and both runners report ready. Immutable cycle 1.0.2 and daily snapshot 1.0.6
fixtures are published; the actual Rust membership/typed-definition validator passes.
Fresh intake 01a0a5af-6c60-7050-bdd7-7c11023503e8 passed membership validation
but failed at checkpoint precondition failed before any model request.
Execution 01a0a5af-7827-7b73-8d4f-c8b671a6c336 reports cleanup CONFIRMED.
The checkpoint gate is the next unresolved blocker; the full cycle remains
unqualified. No membership or checkpoint checks were bypassed.
See that event directory’s README and validate_design_cycle example for the
exact next step and retained run evidence. No phase completion comment was posted.
Candidate handoff evidence and fixed document checks (2026-09-15)
The live manager snapshot export now publishes a complete CandidateSnapshot
receipt and a durable JSON SnapshotRepository manifest for each repository.
Package, manifests and the accepted job result share one metadata transaction.
Manifest identities are deterministic and domain-separated by Host, process,
snapshot and repository, so an immutable report retry uses the same identities.
No worker gets artifact-store credentials or a writable artifact-store mount.
Live report light-portal-test/reports/snapshot-qualification/
daily-ecbe1d67-2efb-4ca0-b31d-59463934f203.json passed native transfer and
confirmed cleanup. The database held exactly two VERIFIED / BOUND / RETAINED
artifacts for this one-repository run. The daily driver independently checked
both file hashes/sizes and equality of the manifest with the repository in the
package, including exact feature/stage/task/checkpoint binding. Its 35 adapter
and lifecycle tests pass. This manifest addition is deployed only in the local
Workflow container; the previous binary is saved in
/tmp/phase1-candidate-deploy.62DmwL/light-workflow.original.
A subsequent source-only addition evaluates pinned
document.metadata.developmentWorkflowDocumentChecks against the verified
package. Each named check has shape
{"kind":"document-lines","repository":"repo","path":"design.md","requiredLines":["# Design"]}.
It performs bounded, exact UTF-8 line checks; it does not execute a shell, trust
model validation claims, or claim semantic design correctness. Unknown kinds,
empty or oversized policies and unsafe paths fail closed. Missing files, invalid
UTF-8 and missing lines produce failed checks. The export records each result
as durable {candidate,check,passed} bytes and returns a ValidationReceipt
plus validationFailures; only passing checks enter passedChecks. Existing
acceptance still requires its pinned check set and independently verifies the
evidence bytes. Definitions without this metadata gain no fabricated checks.
All 94 Workflow unit tests passed, including the two new fixed-check tests.
This check evaluator still needs deployment and live qualification with the
author/review fixture; the passing transfer report above predates it.
Concrete remaining integration work
- Integrate the now-passing separate intake/design cycle into durable daily-test wiring, preserving exact-run reconciliation and native checkpoint boundaries.
- Complete finalization and owner-authorized release of the successful feature’s VM reservation; the completed-stage cancel calls do not release ready successors.
- Integrate Workflow-owned fixed GitHub issue/comment/document publication and
its durable effect identities with the existing task-workspace delivery
primitives. The helper primitives and contract tests alone are not this live
path. Use the approved
networknt/light-agentqualification destination. - Qualify lost responses/restart without duplicate GitHub effects, and both local and installer application evidence recovery after recreation/cache loss plus disabled/unwritable/full/interrupted storage. Existing filesystem-helper tests alone do not establish installer application qualification.
No new events were imported during this slice. No Configserver data was reset. Issue #392 remains open; do not post a Phase 1 completion comment yet.
Durable review dispatch verification (source complete, live gate pending)
Workflow startup now supplies its existing configured artifact store to the task executor. Only pinned development-review turns require it; the existing enqueue entrypoint without a store fails closed for reviews. No store credentials, paths, or writable mounts are passed to Agent or runner.
Review input must include reviewArtifacts.candidate: {id, digest} and may include
reviewArtifacts.before: {id, digest}. A resumed workspace thread requires before.
Both artifacts must be VERIFIED, BOUND, RETAINED, unexpired (or legally held), and
owned by the authenticated Host and current stage process. Workflow reads and
re-hashes actual bytes, independently reconstructs Git trees, and verifies the
feature/stage, repository set, candidate digest, workspace/task identity and exact
workspace.expectedCheckpointDigest. Before/after deltas additionally require
matching scope and base commits. Reconstruction runs off the async executor thread.
These checks run before turn reservation, review allocation, or Agent-job insertion.
Verified review material, artifact reference and optional delta are embedded in
the immutable job instruction; caller-supplied verifiedReviewMaterial is rejected.
The default inline delivery embeds candidate contents. Explicit
reviewArtifacts.contentDelivery: "checkpoint-workspace" instead embeds the
verified manifest and UTF-8 Git delta. Repository contents are read through the
existing read-only task_workspace session, locked to the exact checkpoint.
Both delivery modes independently verify the complete retained packages before
dispatch. The existing 64-KiB instruction bound remains: oversized material is
rejected, not silently truncated. Unknown delivery modes fail closed.
The real disposable-PostgreSQL gate tests missing storage; wrong Agent/review IDs; missing, digest-mismatched, unverified, unbound, deleted, expired, wrong-process, missing-byte and corrupted-byte artifacts; wrong workspace/task/checkpoint; forged verified material; oversized input; and missing before-evidence for resume. Failures leave zero jobs and turns. Exact successful replay retains one job, one turn and one review allocation with verified candidate/delta material.
The 20 contract tests, 92 Workflow library tests, disposable database gate and
all-target compilation pass. Only light-workflow needs a rebuild/redeploy for
this slice; no migrations, Config Server resets, or runner rebuilds are needed.
After the operator’s rebuild, the live snapshot gate passed in report
daily-952c61ff-3d4b-4656-9bce-69577b3491f7.json. Both delivery modes then passed
the disposable PostgreSQL gate, 20 contract tests and 92 Workflow unit tests.
The checkpoint-workspace addition was deployed as a local musl debug binary to
the Workflow container only; it must be included in the next image rebuild.
The original binary is retained in /tmp/phase1-review-deploy.9hrgHP/.
The live author/validation/review/resumed-fix fixture, publication, and installer
qualification remain pending; the passing snapshot gate does not prove that cycle.
Review-cycle wiring continuation (source-only)
Pinned development-review dispatch now binds generated feature/stage/review/job
identities before computing the immutable Agent request digest. Callers may omit
these server-owned values; explicit conflicts fail. The selected reviewer and
Agent must match the definition, workspace intent must be read-only review, and
the model instruction carries the exact Workflow-owned binding. Replaying the
binding step leaves the request unchanged. Author turns and manager snapshot reads
do not use this normalization.
Native workspace adapters return their answer in finalMessage. Review result
reconciliation now parses that bounded answer as strict ReviewResult JSON and
requires equality with the persisted job binding before completing the turn and
applying the finding ledger. Invalid JSON, Markdown wrappers, unknown fields,
oversized answers, and substituted identities are rejected. The normalized ledger
output is separate from the original authenticated execution payload and digest.
SnapshotPackage::verified_delta reconstructs both retained trees in a disposable
private Git object database, checks package digests, feature/stage/task/workspace,
repository sets and base commits, and derives a bounded binary/full-index diff.
It does not read or modify the live task checkout, HEAD, index, or retention refs.
The regression compares its result with the existing manager delta before and
after removing the before-snapshot ref and package cache; binary/deleted/new
executable files survive and corrupt or substituted inputs fail closed.
Verification: 92 Workflow library tests, Workflow all-target checking, and all 32 workspace integration tests passed (30 in the default suite, plus both Linux user-namespace tests explicitly run). These changes are not deployed or live-qualified yet. Durable artifact loading and verification must be wired into review dispatch, followed by the pinned author/validation/review/resumed-fix live fixture. These helpers alone do not complete that gate. GitHub publication and installer application end-to-end qualification remain open.
Latest qualification: unfinished native cleanup interruption (2026-09-15)
The live kill-during-cleanup gate passed at 14:16 UTC. Execution
01a0a56d-7552-7d12-8980-5fde479d23a8 reached the qualification-only barrier
inside native containment cleanup with STARTED / REQUIRED and no terminal result.
The automated test killed only the dedicated Codex runner’s main process. While
it was down, Controller retained lease 01a0a56d-7552-7d12-8980-5fed4f64dacf,
fencing token 1, and VM generation 15 owned by feature
58f121fd-19a3-4f43-8fbc-816d0f9e0ca6. After restart, positive containment cleanup
produced UNKNOWN / CONFIRMED for that same lease and fence. Workflow cancellation
released the VM to generation 16 with no owner. No SQL recovery writes, worker
redispatch, fabricated success, or configserver reset were used.
The snapshot report is intentionally FAILED / cleanup CONFIRMED: interruption
prevents a successful snapshot, while the surrounding fault gate verifies safe
recovery. Report: light-portal-test/reports/snapshot-qualification/
daily-7b0c5556-1125-4a3c-9e57-306434927de6.json. Automated driver:
light-portal-test/scripts/qualify-native-cleanup-interruption.mjs.
The local before/interrupted/recovered record is
/tmp/phase1-native-crash.2Wjh2J/evidence.json (temporary evidence, not a release artifact).
A preceding barrier-expiry run exposed a distinct late-cleanup gap. Controller
now accepts a separately fenced RunnerLeaseCleanupCompleted receipt from the
same authenticated enrollment after reconnect. It updates resource cleanup only;
the original normalized result and outcome remain immutable. Cancellation routes
unconfirmed terminal cleanup to that reconnected enrollment. Native session
cleanup also resolves the persisted execution rather than calling the mock backend.
An acknowledgement alone remains insufficient: local reclamation requires fresh
positive containment proof. Regression tests cover missing evidence, stale fences,
wrong enrollment, replay, and unchanged terminal payloads.
That preceding execution 01a0a53c-cb1c-7f12-be7e-2fc257e5e36c recovered as
UNKNOWN / CONFIRMED with a separate native containment receipt. The supported
snapshot resume reconciled its timed-out report to FAILED / cleanup CONFIRMED,
without submitting another invocation. It is not claimed as a crash inside the
barrier: its barrier had already expired before the manual kill.
Verification: 51 runner library tests and the expanded Controller repository test
against a disposable real PostgreSQL execution schema passed. Controller and both
native runners use the updated code locally. Normal runner binary SHA-256:
05305d6a0792b80ea845a2126deb0dd992f1268017a9b54714fa2afdedeb79ae.
The fault-only binary and Restart=no override were removed after qualification;
approved per-unit Delegate=yes remains. Controller’s local binary deployment is
container-layer qualification and must be included in the next image rebuild.
After restoring normal settings, the automatic-login snapshot smoke passed:
light-portal-test/reports/snapshot-qualification/
daily-394fc262-56e4-42cd-ba9c-cbe55c0cdfff.json. All three native chunks reported
SUCCEEDED / CONFIRMED. Final checks also passed the runner WebSocket integration,
four execution protocol tests, and 26 snapshot qualification adapter/lifecycle tests.
Phase 1 remains open for the full author/validation/review/fix cycle, workflow-owned GitHub publication/deduplication, and installer application end-to-end qualification. Scheduled/hours-long renewal remains explicitly waived. Historical sections below record earlier states and do not override this latest cleanup qualification.
Approved historical recovery (2026-09-15 02:00 UTC)
The operator approved administrative recovery of the exact stranded execution
01a0a2b2-0969-7ce2-bd80-4788c8001c2e. The incident-specific operator script and
evidence are retained in sibling controller-rs/scripts/recover-phase1-prejournal-rejection.sql
and its companion Markdown record. This script is not a migration or automatic
cleanup fallback. Default rollback rehearsal and mismatched-fence rejection
passed before application; committed replay was idempotent.
Local execution audit ID 2 records ADMIN_PREJOURNAL_RECOVERY, actor
operator:steve, and workerReported=false. Outcome remains UNKNOWN with
inspect-required retry classification and no normalized worker result. Resource
release is administratively CONFIRMED; no worker cleanup receipt was fabricated.
The normal Agent cancellation poll reported the job at 02:00:28.241877 UTC;
Workflow recorded CANCELLED and released the personal VM, now generation 3 with
no feature owner. No direct Agent/Workflow/VM/configserver edits were used.
The historical fence no longer blocks a fresh qualification feature. Successful native turns, snapshot/restart recovery, in-flight cancellation, installer application qualification, and GitHub publication qualification remain unfinished.
Runner deployment continuation (2026-09-15 01:57 UTC)
The explicit Workflow-origin allowlist and durable pre-start rejection fixes are
now deployed to both local personal runners. The full runner library suite passed
38 tests, including an expanded admission-document assertion retaining the
interactive origin and adding only the configured Workflow origin. Four broker
socket tests required an unsandboxed rerun; that full rerun passed. The release
build and git diff --check passed.
Controller is healthy. Its live registry confirms Codex generation 15 and Claude
generation 14 CONNECTED with binary digest
sha256:ac4bf700aaf7d7ba3d81152ceeea55691ad03b990d1c189773cf4102b6ffd24c.
Effective config digests are
94cc3b9b0f83f3aeeb7eab06d9f0fad0dfe731dad52e3c004786b74eb2c7faa5
(Codex) and 8f779363c68a51398f043801f2742458103bad084adf4d3fb5af9f710f07295d
(Claude). The existing single-slot limits and Controller origin list were not
broadened. Previous binaries, runner configs and admission JSON are backed up in
/tmp/phase1-runner-deploy.RZ0GsD. These remain local qualification deployments.
The stranded execution is 01a0a2b2-0969-7ce2-bd80-4788c8001c2e;
01a0a2b2-015e-7fc0-8a66-4150b2d76be3 is its request ID, not execution ID.
Read-only inspection of operations.execution_ops.execution_attempt_t still
shows LEASED / REQUIRED. The runner’s SQLite journal has no row for that execution;
the original 01:32:51 UTC log records origin rejection before journal admission.
No historical worker cleanup receipt was fabricated, no ownership fence cleared,
and no replacement feature started. Explicit administrative recovery of this
legacy attempt remains unresolved; deploying the preventive fix cannot supply
missing historical evidence.
The user authorized networknt/light-agent for issue/comment/document publication
qualification, with a dedicated qualification/phase1-publication branch. No
GitHub test writes or implementation commits were made during this continuation.
Native success, snapshot recovery, in-flight cleanup, and installer application
qualification remain open; Phase 1 is not complete.
Live admission continuation (2026-09-15 UTC)
The signed-in Portal accepted qualification invocation
01a0a29c-2630-71d0-9830-def7ede52f05. Its Codex job reached Agent admission,
but no native turn launched: Workflow-only startup had not persisted its accepted
policy snapshot. Cancellation through the same Portal session subsequently
returned CANCELLED; the Agent reported not-dispatched, the feature became
cancelled, and its personal VM generation was released. This proves pre-dispatch
cancellation only, not in-flight native cleanup.
Corrections implemented during this continuation:
- Admit only pinned, async, bounded development Agent tasks through the existing invocation validator; general/unpinned Agent calls remain rejected.
- Compare typed stage claims so an omitted optional
phaseIdmatches the same typed claim without weakening immutable raw invocation replay checks. - Add forward Workflow migration
0010_workflow_action_runtime_privileges, granting the runtime only SELECT/INSERT/UPDATE on its six action tables. - Persist and verify the accepted Agent policy before enabling Workflow polling.
- Include workspace-only jobs in native dispatch; decode the internal manager
snapshot envelope separately from strict interactive
ClientMessageinput. Missing text/request IDs derive from the typed workspace request; explicit conflicting values, unknown fields and missing thread authority are not waived. - Carry the authenticated invocation owner through the typed job transport and
Agent-local immutable admission. Forward Agent migration
0005_workflow_job_ownerpersists this identity; session and workspace policy checks use the user rather than a synthetic Workflow Agent subject. No workspace subjects or consent broadened. - Wire review allocation and successful result recording into native transactions. End-to-end review-input construction and bounded review-loop qualification remain unfinished; these hooks alone are not a completed design loop.
All 88 Workflow library tests, two job-contract tests, three coding-dispatch tests, the Agent store inventory test, and Agent/Workflow all-target checking passed. A real disposable PostgreSQL Agent test additionally proved first-start policy persistence, exact replay, changed-owner replay rejection, user-owned sessions, and selection of workspace-only jobs. The installer fresh/repeated-bootstrap gate passed with 28 migrations across three isolated Host databases. Earlier in this continuation, expanded stage-store and snapshot-transfer component gates passed. These are not installer application end-to-end qualification.
Both additive migrations are staged in canonical/local/installer bundles and
applied to the three local operational databases, without wiping configserver.
Local qualification images are now networknt/light-agent:phase1-owner-20260915
and networknt/light-workflow:phase1-owner-20260915. Workflow is healthy and both
Workflow Agents registered. The release pin remains unchanged. Recreate overlay
and a pre-deploy Workflow binary backup are in /tmp/phase1-owner-deploy.KqZcoR;
these temporary paths are not durable release artifacts.
Qualification definition 1.0.2 explicitly pins each native runner/conversation
and closes it after the inspect-only turn. Its digest is
sha256:9c760f680c03d45bf175846d454cc6e6e7a3e58d9c760d0cef87e7fde40f0207.
Events are retained in sibling light-portal-event/workflow/20260914-phase1-native-qualification.
The four thread-fix events were imported; the active Tool create event was a
projection no-op, corrected by the subsequent thread-tool-update-events.json
forward update. Verified definition version is 1.0.2/aggregate 3 and Tool version
1.0.2/aggregate 4. Publication was staged with the existing explicit owner and
rule; snapshot activation and a successful native invocation are still pending.
Phase 1 remains open: native success, snapshot transfer/recovery, in-flight restart/cancellation cleanup, installer application qualification, and remaining operator/review/fixed-effect integration must still be demonstrated. GitHub-effect qualification additionally needs an explicitly selected disposable repository and branch. Scheduled/hours-long renewal remains waived, not passed or reinstated.
Subsequent live dispatch and publication findings
The legacy Java publication runtime advanced all three ConfigInstance projections
without their companion events. This continuation’s publication reproduced that
defect: endpointRules projection 5 versus stream 4. Three forward reconciliation
events in config/20260915-gateway-publication-reconciliation recorded the exact
materialized values, restoring history parity without SQL projection writes.
The existing working-tree Java companion-event fix was built (six publication
tests passed) and deployed to hybrid-command; the server and genai-command JAR
backups are in /tmp/phase1-owner-deploy.KqZcoR. This Java server update is still
container-layer qualification, not a rebuilt release image.
The ordinary Portal editor then removed only workflow_authorize@call, advancing
endpointRules to version 6. Snapshot 01a0a2b0-68f6-799c-bbba-95a7195cca2f was
captured and activated through Portal. Comparing every property with previous
snapshot 01a0a222-a6dc-7288-9845-f709ead64d8e found only tools changed; access
rules and every other property were identical. Gateway restarted successfully.
Invocation 01a0a2b1-fec7-79b2-91c6-3e7e1dec3d62 then reached Agent admission and
Controller dispatch. Process: 01a0a2b1-ff0a-7562-8a2d-050a4b406c8b; Agent job:
01a0a2b1-ff0a-7562-8a2d-0513e8a170f3; execution request:
01a0a2b2-015e-7fc0-8a66-4150b2d76be3. The personal Codex runner rejected the
lease before worker launch: agent lease origin is not admitted by this runner.
Its worker configuration pins only the interactive Agent service ID, although
Controller admission already includes the approved Workflow Agent ID.
Runner source now supports bounded explicit agentWorker.additionalOriginServiceIds
(default empty), used consistently by worker lease checks and generated admission
documents. Unknown IDs and wildcards remain rejected. Targeted allowlist and
admission-document tests passed. This runner change is not yet deployed.
The second invocation was cancelled through Portal. Agent job and scheduling request are CANCELLED, but the execution attempt remains LEASED with cleanup REQUIRED and Workflow has not received a cleanup report. The runner rejected the lease before journaling it and reconnected with a new generation. Therefore do not treat this cancellation response as VM release or cleanup proof, or bypass the ownership fence to start another feature. Durable pre-start rejection and reconciliation of this exact attempt are the next cleanup work.
Pre-start Agent rejection now persists a terminal policy failure before sending any acknowledgement, with cleanup not required because no inputs were staged or worker launched. A new regression test passes for lost response, restart, exact failure replay, unchanged capacity, and no active execution. The preceding full runner library suite passed 37 tests; the additional rejection test passed separately. These new runner changes remain source-only and do not retroactively resolve the old live lease or establish native success.
Implemented slices
task-workspaceretains exact candidate bytes, Git trees and private refs; tests cover binary content, deletion, executable mode, later edits, digest rejection, lost refs, and interruption between ref creation and record write.- Workflow supports a filesystem artifact backend, bounded digest-verified reads, fsync on writes/promotions, and a write probe for stage admission.
development_storecreates owner-scoped pristine feature records and VM generations, and atomically stores stage claims together with the existing invocation/process/initial-task transaction. Replays bind invocation inputs, policy, budget, deadline and authority, not a newly generated transport run ID.- Migration
0008_development_workflow.sqladds feature, VM, stage and turn records. A deferred process guard rejects unclaimed development starts, including legacy event inserts; a task guard rejects new dispatch after ownership fencing. - Turn reservations persist budget charges and a dispatch intent. An unresolved
intent returns
Uncertain, never permission to resend. Completed results replay exactly, and conflicting or late results are rejected. - Trusted completed review turns populate a durable finding ledger; worker output cannot allocate its own review identity. Stage acceptance uses pinned definition policy and successors, verifies retained artifact bytes and independently rebuilds candidate Git trees, and requires completed execution with cleanup evidence for remote tasks. Fixed-check and publication evidence binds the exact candidate/check or publication receipt.
- Acceptance and limited replan transitions have immutable operation receipts. Replan can reopen an accepted stage with its still-current original inputs, dropping downstream inputs without deleting history or resetting budgets.
- Controller result reconciliation records matching execution cleanup fences and completes bound turn results in the task transaction before acknowledgement. Exact committed-result replay recovers acknowledgement loss without republishing. Actual stage dispatch does not yet call the turn-binding API.
- Owner-scoped cancellation releases idle features whose claims all have accepted
results, using the VM owner and generation. Unsettled execution remains
vm-release-pending; recording cleanup does not yet finalize that release. - A pinned terminal finalize stage can complete the feature and release the VM atomically with acceptance. Unresolved fixed effects block acceptance. Replaying the old completion cannot release a replacement feature’s VM generation.
- Migration
0008is staged in the canonical operational bundle and both personal deployment bundles. Workflow startup requires its tables and migration ledger. Both Compose files mount a Workflow-only evidence volume; the Workflow image installs Git and owns its evidence directory as its non-root runtime user. Localall-in-ltnow runs the rebuilt qualification image and migration; the installer changes remain staged and unqualified.
The HTTP routes inherit the existing invocation identity/grant boundary:
| Method | Path | Payload/result |
|---|---|---|
| POST | /v1/workflow-invocations/development-stage | {claim, invocation}; ordinary invocation status |
| GET | /v1/workflow-invocations/development-features/{feature_id} | Owner-scoped feature and VM holder |
| DELETE | /v1/workflow-invocations/development-features/{feature_id} | {expectedVersion}; cancelled or release-pending feature |
| POST | /v1/workflow-invocations/development-features/{feature_id}/accept | Typed AcceptStage; current invocation claims and durable evidence required |
| POST | /v1/workflow-invocations/development-features/{feature_id}/replan | Typed ReplanStage; owner and historical invocation claims checked |
AcceptStage.nextStage is a selector for a handoff or null for completion.
Completion requires a finalize stage with pinned
document.metadata.developmentWorkflowTerminal: true and an empty
developmentWorkflowSuccessors array. No caller-supplied flag overrides this.
Development definitions require document.metadata.developmentWorkflowStage
matching the StageSelector. invocation.input.stageClaim must equal the supplied
claim. The definition digest, definition ID, workspace binding and stage deadline
are checked. The ordinary authenticated invocation endpoint also accepts a
marked development definition with input.stageClaim; unclaimed development
starts remain rejected. Initial intake may create pristine feature state in the
same transaction under pinned definition policy, as described below. There is
no public API accepting arbitrary serialized feature state.
Verification
Run against a fresh disposable database:
export DEVELOPMENT_WORKFLOW_TEST_DATABASE_URL='postgresql://.../fresh_test_database'
bash scripts/run-development-workflow-store-gate.sh
The gate refuses an existing workflow_ops schema. It creates migration roles,
installs the base and development migrations, then exercises the store as
operations_workflow_runtime. It does not silently skip PostgreSQL.
Passed locally on PostgreSQL 17: concurrent identical starts produce one instance; changed inputs conflict; lost-response replay; owner isolation; rollback without orphan claims/processes; generic/event bypass rejection; dispatch fencing; turn intent/result replay without a second charge; stale completion rejection; initial VM release and late-release protection. The gate also runs all 20 Phase 0 contract tests and all 83 Workflow library tests.
The expanded PostgreSQL gate also covers review reservation/completion, corrupted
artifact rejection, candidate tree mismatch, pinned successor checks, acceptance
and replan replay/conflicts, budget/history retention, idle release after accepted
history, and runner fence rollback, identity/policy/cleanup checks and exact
acknowledgement recovery. These are database fixtures, not a live Controller run.
The terminal-stage fixture also proves completion releases the slot, uncertain
fixed effects block it, a replacement feature increments its generation, and old
completion replay leaves the replacement reservation intact.
Candidate verification requires Git in the Workflow runtime image; the Dockerfile
now installs it and includes contracts/ needed by the workspace dependencies.
The new image runs as the non-root workflow user and has Git available.
Six generated catalog/mapping/instance events are in sibling repository
light-portal-event/config/20260914-development-workflow-evidence/events.json.
Live conflict checks passed; after explicit approval all six were imported and
published to local snapshot a229fbf8-da56-48be-baf9-37c96c65ecee. They target the
verified local Workflow instance, not arbitrary installer tenants. The
event-generation skill kept generation separate from the approved live import.
Local deployment qualification, 2026-09-14
- Configserver and operations were backed up before mutation in private directory
/tmp/development-phase1-qualification.RKqqlF/; no databases were reset. - Migration
0008and its checksum ledger were installed in one transaction. The Workflow runtime role can access all six new tables. - Running image:
networknt/light-workflow:phase1-local-20260914(manifestsha256:7157c4162e79d44e6a6f94ee3321d269ca15ab08a769d19b25da2bbac3ac5f48). The pinned release image/env file was not overwritten. Subsequent normal deployments still select the pinned image unless this override is supplied. - Only Workflow was recreated; existing broker, action and credential overlays
were retained.
/readyis ready with Controller connected and the new remote snapshot active. The feature route rejects an unauthenticated request (403). examples/development_artifact_gate.rs, compiled in the same Debian builder, exercised the actual backend as UID/GID 999 inside Workflow: stage, recreate, promote/retry, recreate, exact binary read, bounded reads, tenant isolation and corrupted-object rejection. Isolated test objects were deleted afterward.- Separate network-disabled containers proved read-only storage is rejected during initialization and ENOSPC makes the write probe fail on a 16 KiB tmpfs. The live volume was never filled. Six artifact-store unit tests also passed.
- Workflow alone mounts
/var/lib/light-workflow/evidence; neither Workflow Agent has that mount. No artifact credentials were sent to a runner.
This qualifies the local storage slice, not runner → Controller → Workflow transfer, metadata/promotion crash recovery, a live design loop, installer parity, or full owner/grant HTTP authorization. Phase 1 remains incomplete.
This proves the store slice, not live worker execution or HTTP authentication qualification. Runtime-role testing does not stand in for the separate runner and Workflow container gates.
cargo check -p light-workflow --all-targets passed. Strict all-target
workflow-store Clippy is blocked by the pre-existing redundant matches!
assertion in tests/postgres_binding.rs:45; that unrelated test was not edited.
The separate binding/restart PostgreSQL test was not run in this session.
Remaining Phase 1 work
Portal invocation entry continuation (2026-09-14)
portal-view now has a separate Invoke Workflow Tool row action and dialog;
the API-endpoint invocation path is unchanged. It loads the current Gateway’s
authorized live catalog, matches the Tool name, shows its published input schema,
and accepts bounded JSON arguments and an optional existing grant UUID (never a
token). It uses ordinary session cookies/CSRF and an in-memory MCP session on a
fixed /mcp route. The dev proxy now forwards that path. Submission requires
explicit confirmation and is attempted once; failures leave submission disabled
with an uncertain-outcome warning. Closing does not cancel or release a VM.
Verification: 19 targeted client/dialog/fetch-wrapper tests passed, targeted ESLint and whitespace checks passed, and the production build succeeded (existing dependency/chunk warnings remain). Repository-wide TypeScript checking still reports errors in other files and is not claimed green.
Browser qualification reached the new dialog through the real Tool row. Its
MCP initialization was rejected with HTTP 401; Invoke remained disabled and no
tools/call was sent. Gateway audit evidence at 20:55:12 UTC confirms
POST /mcp, 401, no authenticated principal. The active mcpChain is
[exception, cors, security, mcp], unlike the Portal routes that process the
session through stateless.
With operator approval, current local snapshot
01a0a1b9-5082-71a6-8908-40098f781c3a adds stateless before security in
mcpChain. Snapshot comparison changed only chains; all other chains and
Tool permissions remain unchanged. The Gateway restarted at 21:01:54 UTC,
registered with Controller and retained access-control revision
809fcf46c220fa12d0647036dd87349fe2b59bf079e1343240bdf655cb66fe63 with default-deny
enabled. Anonymous initialization still returns 401. The signed-in dialog now
returns 403 and remains disabled. The MCP-specific originAllowlist is absent
from the active snapshot and its catalog default is []; MCP preflight rejects
browser origins not explicitly listed. The existing WebSocket Origin allowlist
does not cover MCP.
With subsequent explicit operator approval, snapshot
01a0a1c0-e753-7bea-b5b0-add24ed27ffa is now current for portal-bff-loc.
Comparison against the preceding snapshot confirms only originAllowlist
was added, with value ["https://localhost:3000"]. Gateway restarted at
21:10:00 UTC with the same access-control revision and default-deny enabled;
anonymous initialization remains 401. The signed-in dialog now identifies the
403 as Workflow caller authorization was denied, after MCP Origin preflight.
The Gateway action context authenticates dual identity even for initialization;
the specific failing identity condition still needs diagnosis. No additional
caller permission was granted and no native invocation was submitted.
The client exposes only fixed HTTP-error diagnoses, not arbitrary response
bodies; focused client/dialog/fetch tests pass (20 tests), and focused client
ESLint plus diff whitespace checks pass. This is not full Phase 1 qualification.
Follow-up diagnosis: the mounted Gateway workflow-actions.yml admits only
com.networknt.workflow-1.0.0 with origin: workflow, bound to its approved
mTLS peer. It has no interactive caller registration. The MCP handler calls
action_gateway::Runtime::context even for initialization; this calls strict
dual_identity::authenticate, requiring a transport peer and an application
X-Scope-Token as well as the user token. Workflow-origin callers additionally
require an action reference. The Portal client sends session cookies/CSRF, and
its Vite /mcp proxy supplies neither application credentials nor an upstream
client identity. Session routing and the Origin exception therefore cannot make
this browser path satisfy the service-only contract. The generic live 403 does
not distinguish which individual check failed first.
The missing implementation is an explicit trusted Portal ingress/forwarding profile with server-held application credentials and authenticated transport, preserving user/Host, CSRF, Tool ACL and root-grant checks. Do not put app secrets in the browser, relabel a Workflow caller as interactive, or skip dual identity for browser requests. Activation requires approval of the new ingress access; no caller-policy or credential changes were made during this diagnosis.
Approved local Portal ingress (2026-09-14)
With subsequent operator approval, the local Vite server now has an opt-in
server-only MCP ingress (portal-view/server/workflowIngress.mjs). A dedicated
one-day app identity, com.networknt.portal.workflow-ingress-local-1.0.0, is
registered as interactive in the local Gateway action policy with exact mTLS
peer fingerprint 5f40f4a0d659586579dc87a29ed4c07d1fd56a2d2886c841b8ff93dbfaa28805.
The existing Workflow entry and its CA trust are preserved. Gateway restarted
at 21:37:33 UTC with unchanged Tool ACL revision and default-deny. This is an
overlay activation, not another configserver snapshot or release-image change.
The browser receives no app credential. The ingress rejects supplied service identity/action headers, enforces the exact local Origin/path/method, forwards the ordinary cookies/CSRF to Gateway over verified TLS with a server-held client certificate, and leaves user/Host, Tool ACL and root grant admission intact. It has bounded bodies/time, filtered headers and no automatic retries. HTTP/2 split Cookie fields are supported without accepting duplicate session/CSRF cookie names. Runtime credentials and private key URLs are denied by Vite.
Live success: signed-in initialization and catalog discovery now display the
isolated phase1_native_binding_intake published schema in the Portal dialog.
No Tool invocation or grant enrollment was submitted. Negative live checks:
anonymous direct MCP and forged-session ingress 401; foreign Origin, supplied
app identity, duplicate CSRF cookies and credential-file URL 403. Seven new
ingress tests and 20 existing client/dialog/fetch tests pass, as do focused
client ESLint and whitespace checks. This does not qualify native dispatch or
the complete authorization matrix.
The opt-in environment file and credentials are ignored by Git. Preparation,
activation and guarded rollback instructions are in
portal-config-loc/all-in-lt/workflow-actions/README.md. The app credential
expires 2026-09-15 21:37:22 UTC. Production/installer packaging and managed
credential rotation are not implemented by this local Vite plugin. Grant
enrollment and authenticated status/cancel controls remain prerequisites for
the native qualification run; Phase 1 is not complete.
Portal lifecycle controls continuation (2026-09-14)
The invocation dialog now includes a collapsed Manage existing workflow run
panel for explicit status/result reads and confirmed cancellation by exact
workflowInstanceId. It discovers the existing Gateway lifecycle tools before
calling them, uses the signed-in MCP ingress, and does not start work or read
status automatically. A cancellation attempt disables further cancellation and
ID edits while preserving status reads for reconciliation. It does not equate
an accepted cancellation with completed cleanup or released VM ownership.
The focused controls/dialog/client suite passes 20 tests. Live read-only
qualification with random nonexistent ID
78bc469f-459a-48f6-82fa-3f5f5e9d0834 reached Gateway and was rejected:
Gateway RPC -32001: Access denied: no access control rule defined for workflow_get_status@call.
No cancellation was sent and no run was created. Activation of scoped rules for
workflow_get_status@call, workflow_get_result@call, and workflow_cancel@call
requires explicit approval; no ACL changes were made in this continuation.
At that checkpoint, grant enrollment was pending. Source inspection then found
the broker /workflow/credentials/enroll, /complete, and /revoke routes and
issuer consent/PKCE flow in the ordinary Workflow router. Those routes and the
broker grant flow were later retired. The actual root binding pins then included profile
workflow-action-v1, workflow definition ID, and definition/policy/response
policy digests. Native execution was blocked at that checkpoint pending enrollment
and lifecycle authorization/transport qualification.
At that checkpoint, root admission required an owner-authorized renewable grant; the UI did not fabricate or auto-enroll one. This grant prerequisite was later replaced by the Workflow Invoke user-token and LONG registration contracts.
Isolated qualification fixture continuation (2026-09-14)
Imported an isolated two-turn, inspect-only intake definition, asynchronous Tool
binding and its parameters in light-portal-event/workflow/20260914-phase1-native-qualification.
The initial definition exposed Java/Rust digest disagreement for explicit nulls;
corrective events published version 1.0.1 without optional null fields. Both
digest implementations now agree, and the definition, Tool and parameter
projections are verified. Original event history and three failed legacy
projection records remain intact. This is a fixture correction, not a general
fix to cross-language digest canonicalization.
Added read-only examples validate_development_qualification and
validate_qualification_owner_rule. The first checks typed workflow parsing,
round budget scopes and pristine intake. The second passed six strict CEL
owner/Host/missing-claim cases against the exact imported qualification rule.
The current generic JWT rule does not test users, so the fixture uses an explicit
owner-and-Host rule rather than relying on an Allowed users field alone.
With operator approval, the owner-only Tool is now published on portal-bff-loc
(com.networknt.portal.gateway-1.0.0, environment loc), publication
01a0a1ac-fa40-797c-9e8f-c5399a3ce4b4, current snapshot
01a0a1ad-28b6-707c-b612-f6ee61cd2767. Preview preserved all seven existing Tools
and all 873 existing endpoint rules; snapshot comparison changed only tools,
endpointRules and ruleBodies. The local Gateway restarted and registered with
Controller at 20:48:36 UTC, loading access-control revision
809fcf46c220fa12d0647036dd87349fe2b59bf079e1343240bdf655cb66fe63 with
default-deny still enabled.
An initial activation mistakenly targeted the catalog’s Rust AI Gateway in
dev, not the running local Gateway. Its previous snapshot
2e1268cf-2e37-40b3-bcb4-5352aed96a7a was restored with operator approval.
Staged desired properties remain on that unused instance and must be reviewed
before any future snapshot capture/activation; the event fixture README records
the full rollback and publication identities.
No native job/grant has been created. Configuration activation is not live execution qualification. Native results, cancellation/release, snapshot transfer and installer application qualification remain outstanding. Both native runners were CONNECTED at the start of the earlier fixture continuation.
Native policy activation continuation (2026-09-14)
With explicit operator approval, imported four native coding-profile authoring
updates and four corrective contract-digest events. The first preview rejected
the unchanged digest after worker pins changed; correction retained the original
import history and the normal guarded append path. All eight imports succeeded.
Artifacts and publication IDs are recorded in
light-portal-event/genai/20260914-phase1-native-policy/README.md.
All four policies were published through candidate-checked Portal activation, with explicit action-time confirmation for the interactive policies. Their full workspace bindings match authorization revision 3, retaining the same owner and inspect/implement/review intents while adding the existing Workflow service IDs. Native binaries, runner bindings/configuration and combined Controller admission were backed up and replaced together; both Workflow-Agent origins are admitted. Controller and all four Agents were recreated, preserving interactive image tags.
Live startup exposed a Rustls provider-selection panic in the runner transport
task. Explicit provider initialization at the runner entry point fixed it. The
corrected binary and matching admission pins are deployed. Both native runners
return HTTP 200 ready: true from /readyz; Controller reports both CONNECTED
with the expected binary and effective configuration digests. Exact digests and
publication IDs are in the event README. Transport identity gates passed, but
only exercised empty authenticated polls, not native jobs or grants. These checks
do not establish loaded Agent policy or end-to-end Phase 1 qualification.
Configserver was backed up to /tmp/phase1-native-policy.jbANDG/configserver-before.dump.
No database was wiped. The isolated development workflow fixture and native,
snapshot, cancellation and installer application qualification remain pending.
Atomic snapshot publication continuation (2026-09-14)
Snapshot publication now uses the same database transaction as the accepted Agent report, without acquiring a second pool connection. A rollback leaves no committed artifact metadata or accepted result. Content-addressed bytes can remain unreferenced after rollback; the immutable report retry repeats promotion with the same artifact ID. Replays check execution/process/task identity, content, policy and retention/deletion fences, not only artifact ID and digest.
The fresh PostgreSQL runtime-role gate passed with a one-connection publication
pool: rollback leaves zero metadata rows; retry after reopening filesystem storage
and committed replay leave exactly one; changed content/process/execution and
deletion-pending evidence are rejected. All 20 contract tests and 86 Workflow
library tests passed, along with the snapshot assembly/recovery test and
cargo check --locked -p light-workflow --all-targets. These are component and
database checks, not the remaining authenticated native/installer exit gates.
Deployed only Workflow as networknt/light-workflow:phase1-snapshot-tx-20260914
(manifest list sha256:cc6513e232db7f5df06957d7a94dc44b54c290a183880be5a9065d817bea7dca)
using the existing local qualification override. It is healthy; both Workflow
Agent mTLS gates passed again (authorized empty poll 200, wrong Host and missing
scope 403). The release pin and configserver data remain unchanged.
No new policy events were imported and no native runner policy was changed in this continuation. Additional policy/qualification imports and coordinated runner admission changes require the requested operator approval. Preflight confirmed all four native profiles still pin the old worker binary digests.
Live deployment and cancellation continuation (2026-09-14)
The ordinary invocation cancellation route now enters development feature cancellation before taking the invocation lock. It preserves the configured cancellation policy and effect fence, checks ownership, and cannot let an old stage cancel a newer active claim. Cooperative cancellation fences new dispatch; it does not manufacture compensation or native cleanup confirmation. VM release still requires terminal cleanup evidence for the exact generation.
Agent job and development cancellation reconciliation now starts independently of Workflow’s direct runner switch. The local deployment keeps direct runner execution disabled while enabling this reconciliation for Agent-owned jobs.
Fresh private PostgreSQL custom-format backups of configserver and all three
operational databases are retained under /tmp/phase1-deploy.9vy24c (temporary
local evidence, not a durable backup location). Installed and digest-checked
Workflow migrations 0008/0009 and Agent migration 0004 in operations,
operations_networknt, and operations_taiji. Existing migrations were retained.
The initial migrator-role attempt failed a REFERENCES privilege check and rolled
back; the unchanged canonical SQL then succeeded under the existing PostgreSQL
table owner. No databases were wiped and configserver data was not changed.
Targeted local Compose overrides now run these qualification images:
| Service | Image | Manifest digest |
|---|---|---|
| Workflow | networknt/light-workflow:phase1-intake-20260914 | sha256:ec6fa5ea701088a45f346577a52faa84b1ab2dd978be0dbbd5a42085e50d83d1 |
| Controller | networknt/controller-rs:phase1-qualification-20260914 | sha256:db4d3df50a7adf5c048e75a5da0e86ba29a15c666c1effb543edc35eb9290d57 |
| Both Workflow-specific Agents | networknt/light-agent:phase1-intake-20260914 | sha256:a8a237a73b1d8b45cc0302c7f186424a98fe6ecb34aab3bb90109b1226970132 |
The release pin remains 2.3.5-dev.20260909.2338; interactive Agent images and
native runner binaries/configuration remain unchanged. The local override and
targeted recreation script are in /tmp/phase1-deploy.9vy24c; a normal deployment
without that override does not select these qualification images.
Verification: the fresh PostgreSQL runtime-role gate passed, including disabled
cancellation, effect-fenced cancellation, cooperative release-pending, ownership,
post-cancel dispatch rejection, terminal cleanup and replacement-generation
protection. All 20 contract tests, 86 Workflow library tests and 9 Workflow binary
tests passed. Workflow and Controller are healthy; both Workflow Agents registered
and both existing native runners reconnected. The new read-only local gate
bash scripts/run-development-workflow-transport-gate.sh passed for both Agent
mTLS identities: empty poll 200, foreign Host 403 and missing scope 403. This is
transport qualification, not native dispatch, snapshot or cancellation proof.
Live startup logs also explain the earlier legacy UI smoke: the legacy event
consumer is intentionally disabled (local_event_source_unavailable); direct
grant-bound invocation admission remains active.
Next qualification prerequisite: refresh the published native worker policy and
matching runner admission together. The rebuilt Codex worker reports capability
digest sha256:9793da1ba3b167e5172adbadfd40bced171986b65615b0d0b0056e93689e5f51,
which differs from the currently pinned policy. Its binaries are staged, not
installed; existing policy hashes were not relaxed. Then publish an isolated
development definition/Tool binding and prove the authenticated native run,
snapshot acceptance, cancellation cleanup and installer application path.
Operator controls and fixed-action integration also remain outstanding.
Gateway admission and initial intake continuation (2026-09-14)
Browser access and the signed-in Portal session were verified. The existing
Workflow Admin Start form sends the legacy startWorkflow event command; it
does not call the renewable-grant invocation boundary. A read-only
workflow-mcp-smoke start returned instance
01a0a15b-21bf-712c-bf0e-73ad08852783, but no matching runtime process or
quarantine row was found during the check. This is command acceptance only,
not a successful execution or native qualification. No development-stage
definitions were present among the ten active definitions shown in the Portal.
This historical note predates the single-start-path change. Current root
Workflow launches enter through Gateway MCP workflow_start; the public
POST /v1/workflow-invocations root start route has been removed. The separate
/development-stage route remains for claimed development-stage execution and
is not a general workflow launch path. The old event command and the old root
invocation route must not be used to start workflows.
For first intake, the published definition must pin:
document:
metadata:
developmentWorkflowStage: {kind: intake, phaseId: null}
developmentWorkflowIntake:
vmId: <pilot-vm>
workspaceBinding: {id: <published-binding-id>, digest: "sha256:<digest>"}
maximumTurns: 12
maximumRemediationRounds: 3
maximumDurationSeconds: 3600
The published input schema must allow stageClaim and featureIntake, where
featureIntake contains only an issue with repository, number, and its
matching GitHub issue URL. The initial claim uses canonical non-nil UUID feature
and transition IDs, predecessor version 1, intake stage, empty accepted inputs,
the exact published definition/binding and an absolute deadline no later than
the incoming Gateway authority. Workflow derives pristine state and budget
limits from the pinned policy; input cannot supply owners, VM generations,
historical results or larger limits. Issue references are validated syntactically;
this does not establish issue contents, frozen requirements or human acceptance.
Intake creation, VM acquisition, stage claim, invocation and initial task commit together with the existing run-authority admission transaction. Replays compare the immutable creation fingerprint even after VM release, without reacquiring a replacement owner’s slot. Gateway attempt-relative deadlines are narrowed to the immutable claim deadline before persistence, keeping identical retries stable without widening their authority.
Verification: all 20 contract tests and 86 Workflow library tests passed. The fresh PostgreSQL runtime-role gate passed initial intake rollback with no orphan feature/VM/process, concurrent identical intake/claim starts, changed-input and owner rejection, deadline narrowing, and replay after a replacement VM holder. These are component/database tests, not authenticated HTTP or native execution.
No application image was deployed or live schema changed in this continuation. Next: publish an isolated development intake/stage workflow and its exact Tool binding, deploy the coordinated qualification binaries/migrations, then execute through the real grant boundary. Operator controls, native execution/recovery, fixed-action integration and installer application gates remain outstanding.
Snapshot and cancellation integration continuation (2026-09-14)
Implemented fixed WorkspaceExecutionSpec.managerSnapshot reads, admitted only
for a Workflow principal, published runner workspace binding, existing task,
read-only intent and pinned checkpoint. Interactive Chat never creates this
field. The worker does not launch a native model for this operation. Results
carry a bounded chunk and manifest through the existing Controller execution
channel and authenticated Agent job report. Workflow binds them to the original
request, assembles persisted reports, verifies the package digest and Git trees,
then stages/promotes the artifact with Workflow-owned storage access. Public
task results expose transfer progress and the completed artifact reference, not
all accumulated bytes. Snapshot task names must be pinned in definition metadata
developmentWorkflowSnapshotTasks; these fixed reads do not charge model turns.
Added service-authenticated Controller endpoint
POST /internal/execution/requests/{request_id}/cancel. It fences scheduling,
returns exact lease cancellation messages for delivered work, and reports
confirmation only after all issued attempts have confirmed cleanup. Missing
requests, disconnected runners, and timeouts do not produce confirmation.
Agent terminal reports now include Controller cleanup evidence or a local
never-dispatched receipt. Workflow cancellation revokes native and fixed-action
admission, cancels stage execution, and retains the VM while tasks, effects or
native reports remain unresolved. The background reconciler performs an exact
owner/generation CAS before releasing the slot. controller-rs is now part of
the uncommitted implementation set.
Passed in this continuation:
- Fixed manager read: correct scope, exact replay, interactive denial and digest mismatch rejection.
- Multi-chunk assembly: interrupted/incomplete and out-of-order transfer, identical duplicates, corruption/identity/bounds rejection and verification after removal of runner-local storage.
- Disposable PostgreSQL Workflow gate: cancellation holds before a report, releases after a verified fence, and late replay cannot release a replacement.
- Disposable PostgreSQL Controller runner gate: another service cannot cancel; delivered work remains unconfirmed; a matching confirmed terminal result settles cancellation. This is a real database test, not live native execution.
- Workflow library: 83 tests passed before the additional cleanup-proof unit test. Workspace protocol/library checks and all-target compilation passed.
- Debian qualification builds succeeded for Workflow, Agent, both native worker
binaries, native runner and Controller. Workflow/Agent qualification image
tags are
phase1-completion-20260914; release pins remain unchanged.
Live qualification is not complete. Read-only inspection found zero active
Workflow run grants, zero live invocations and zero active execution attempts.
An authenticated Portal session and scoped run are required; no grant was
fabricated and no authentication check was bypassed. The user was asked to sign
in at https://localhost:3000. New binaries/migrations from this continuation
have not replaced live services. Native end-to-end/restart/cleanup and installer
application-level qualification remain pending. Previous installer component
and database evidence below remains separate.
scripts/run-development-workflow-completion-gates.sh runs the component and
database checks with explicitly supplied disposable Workflow and Controller
databases. It does not claim to run the authenticated live gates.
Native dispatch continuation (2026-09-14; not live-qualified)
The working tree now uses Workflow-owned workflow_agent_job_t dispatch intent
instead of querying/writing Agent-owned catalog and queue tables through the
Workflow pool. Registered Workflow Agents pull immutable, bounded jobs and
report terminal results over the existing authenticated mTLS job boundary.
Service-mode calls require an exact registered Agent definition UUID and live
run authority; inline calls retain their separate catalog path. Published
developmentWorkflowTurns[taskName] metadata pins the turn kind and budget scope.
Reservation and dispatch intent commit together; task identity is the stable job
identity and changed retries are rejected.
Feature cancellation revokes native-job admission atomically. Agent polling requests cancellation when authorization is revoked. This does not establish confirmed cleanup or release a held VM. The terminal Agent report/fence bridge is implemented but still needs authenticated transport/replay and live native qualification, including cancellation races. Do not deploy this as qualified.
New canonical migrations are workflow-store/0009_workflow_agent_dispatch and
agent-store/0004_workflow_job_transport. Both deployment bundles contain all
26 migrations. Neither new migration was applied to the user’s live databases
in this continuation, and no Agent/runner service was rebuilt or restarted.
Verification in this continuation:
- Fresh disposable PostgreSQL development-store gate passed, including native intent replay, changed-budget/Agent rejection, Host isolation, one turn charge, and cancellation preventing further reservation.
- Workflow library: 83 passed. Agent library: 21 passed, 4 PostgreSQL tests ignored (not claimed as live qualification). Typed transport bounds: passed.
- Installer Python tests: 32 passed. Its isolated operational-database runtime gate passed fresh bootstrap, repeated bootstrap, three-Host ledger parity, runtime-role isolation, and swapped-credential rejection using the new bundle.
- Strict Clippy was blocked by existing
collapsible_iffindings incrates/config-loader/src/lib.rs:568and:586; those files were not changed.
Still outstanding in the four requested areas: actual Codex/Claude dispatch and result recovery; fixed runner-to-Controller snapshot chunk transfer and durable export recovery; confirmed cancellation/cleanup receipts and VM-generation CAS release; installer application/image/artifact restart qualification. The installer database gate alone does not establish those application guarantees.
Overall Phase 1 backlog
- Qualify migration/image/settings and recovery in the installer stack variant;
local
all-in-ltstorage qualification is recorded above. - Qualify the pinned-policy intake creation path and wire Workflow Admin controls.
- Connect the acceptance/replan routes to Workflow Admin and qualify HTTP auth; add human signoff/disposition producers and changed-input supersession.
- Connect turn reservations to real Codex/Claude stage rounds and schema repair; qualify the exact published and runner-local workflow-agent identities.
- Transfer snapshot chunks through runner → Controller → Workflow, export them with Workflow-owned credentials, and exercise interrupted transfer/recovery.
- Connect cancellation to confirmed Controller execution/effect fencing and release after cancelled/failed stages; no timeout-based release. Successful terminal acceptance now releases the slot in the database gate.
- Implement fixed validation, issue/comment effects and immutable document publication with reconciliation, then run all parent-design live exit gates.
Do not mark Phase 1 complete or deploy this as a qualified pilot based on the store gate alone.
Local lifecycle ACL qualification continuation (2026-09-14)
Reconciled three Gateway-generated ConfigInstance streams from version 1 to 2
using exact current values and guarded event import, preserving configserver
data and historical events. Fixtures and audit notes are in
light-portal-event/config/20260914-gateway-publication-reconciliation/.
The approved owner-and-Host lifecycle rules then applied through Portal at
endpointRules version 3. Snapshot 01a0a222-a6dc-7288-9845-f709ead64d8e is current;
only endpointRules differs from predecessor
01a0a1c0-e753-7bea-b5b0-add24ed27ffa, which remains available for rollback.
After Gateway restart, access-control remained enabled with default deny and
revision 724263f9fb59a7b4a2c9650c97cc31efe75c0df37c7bbea8b473a3aa6b984562.
A signed-in Portal status request for a nonexistent qualification UUID passed
Gateway authorization. It failed downstream with a non-JSON lifecycle response.
The configured ordinary HTTP Workflow listener returns an empty HTTP 403 for
this route; enforce_action_receiver requires authenticated dual identity and
the mTLS peer when action authorization is enabled. The Gateway lifecycle client
still uses the ordinary invocation URL/client. This is not a successful
Workflow lifecycle qualification, and no native run or cancellation was sent.
The publication versioning source patch adds atomic ConfigInstance companion events plus a projection/history consistency guard. Six command tests and ten existing publication persistence tests pass. These counts do not establish live publication, replay, or deployment qualification; the source patch is not yet deployed. Phase 1 remains incomplete.
Gateway-to-Workflow mTLS continuation (2026-09-14)
The downstream transport blocker above is resolved locally. When the A2
workflow-actions.yml profile is present, MCP Workflow dispatch uses its
authorization.control endpoint, client identity, CA, and service token. This
fixed HTTPS transport takes precedence over the legacy
mcp-router.workflow.invocationUrl and scope-token environment settings. With no
A2 profile the legacy transport is unchanged. Invalid A2 transport configuration
rejects startup/reload; it does not fall back to the ordinary HTTP listener.
Only Workflow start, wait, ambiguous-start recovery, status, result, and cancellation use this dedicated client. General Tool HTTP clients do not receive the Workflow identity. The client trusts only the supplied CA, verifies hostname and peer certificates, disables proxy use, redirects, and automatic retries, and preserves the inbound user Authorization alongside the A2 app token.
Verification: cargo test -p light-pingora --lib --locked --quiet passed 450
tests, with five ignored. The two added tests cover unsafe endpoint and missing
credential rejection, a real mutual-TLS handshake, user/app header preservation,
redirect refusal, and plain-HTTP refusal. The musl release Gateway build passed.
The binary was installed into the existing local light-gateway container and
restarted at 23:33 UTC. SHA-256:
656c934ca6ebf0fa302c32a3694841218b111addae123bded68bd674f81eff41.
The release image pin remains 2.3.5-dev.20260909.2338; this container-layer
qualification replacement is lost when the container is recreated. The image
must be rebuilt through the normal release workflow before durable deployment.
The old binary is retained at
/tmp/phase1-gateway-mtls.o7RopZ/light-gateway.previous, SHA-256
98b83d592bff270e48bf9e51c688eead06722f2a219830cd06a1b2d580111e2a.
Local rollback: stop this container, copy that saved binary to
light-gateway:/app/light-gateway, and start it again. No configuration, database,
credential, or policy changes are needed for this binary rollback.
Signed-in Portal status and result reads for nonexistent UUID
78bc469f-459a-48f6-82fa-3f5f5e9d0834 both passed Gateway authorization and
Workflow’s authenticated receiver, returning the expected JSON HTTP 404
workflow invocation is unavailable. Gateway default-deny and ACL revision
724263f9fb59a7b4a2c9650c97cc31efe75c0df37c7bbea8b473a3aa6b984562 were unchanged.
No native invocation or cancellation was submitted. This establishes the local
read transport, not grant enrollment, native execution, cleanup, or installer
qualification. Phase 1 remains incomplete.
Enrollment UI and authenticated broker routing (2026-09-14)
Added the workflow_authorize MCP lifecycle action and Portal’s explicit
“Authorize workflow” panel. The action requires its own Gateway ACL and a
second authorization check on the selected published Workflow Tool. Gateway
derives the immutable workflow-action-v1 binding from that Tool; caller inputs
contain only Tool name, scope and a 60–3600 second duration. It rejects nested
action/delegation enrollment and cleartext dispatch. The issuer remains the
authority for scope ceilings and final user consent. The response is bounded to
16 KiB and exposes an HTTPS consent URL and references, not reusable credentials.
The Workflow mTLS listener now mounts enroll/complete/revoke routes. Those routes check the app/TLS-peer pair as well as the existing allowed-caller and user-token checks. The public PKCE callback remains on its separate listener. Portal never accepts passwords or tokens, never automatically opens consent, and fences an uncertain enrollment submission. The selected Tool binding cannot be supplied by the browser. No existing endpoint ACL was modified.
Verification: 25 focused Portal tests and ESLint passed; all 86 Workflow library tests passed. MCP’s full pre-new-contract-test suite passed 450 tests with five ignored; two additional enrollment contract tests passed afterward. Initial socket-suite failures were sandbox socket denials; the full suites passed with local socket access. Gateway and Workflow musl release builds passed.
Local container binaries were replaced and restarted at 23:42 UTC, preserving
the existing configuration and release image references. These replacements are
container-layer qualification only and do not survive container recreation.
New Gateway SHA-256:
830476a8c62e3d67bb3224023b25f8d69308e7d4d2046c648f0e9e6c4d456aef.
New Workflow SHA-256:
ae9f59e4f45ea9e5e6c7cdb27af0594ac55109143f66a79641eb6b3cb8bbd208.
Previous binaries are in /tmp/phase1-enrollment.p7tKSl/ as
light-gateway.previous and light-workflow.previous. Roll back by stopping the
corresponding container, copying its saved binary to /app/light-gateway or
/app/light-workflow, and starting it again; do not restore database snapshots.
After restart, the signed-in status check still returned the expected Workflow
JSON 404 for the nonexistent qualification run. An enrollment request over
valid client TLS but without app/user credentials returned 403. ACL revision
724263f9fb59a7b4a2c9650c97cc31efe75c0df37c7bbea8b473a3aa6b984562 and default deny
were unchanged. workflow_authorize@call is deliberately absent from the
version-3 endpointRules map pending explicit activation approval.
The Portal form is prepared with proposed portal.r portal.w scopes and a
900-second duration, but its consent checkbox is unchecked. No enrollment,
grant, native invocation or cancellation was submitted. Next: approve the same
owner/Host ACL for this enrollment action, begin issuer consent for the isolated
qualification Tool, have the owner complete issuer authentication/consent, then
qualify the native execution and cleanup. Phase 1 remains incomplete.
Superseding scope-consent decision and rollback (2026-09-14)
The user rejected additional workflow-specific consent: only the existing Portal
portal.r / portal.w scope consent is required. The enrollment UI and generated
workflow_authorize MCP tool have been removed from source. The running local
Gateway was restored to the pre-enrollment binary, SHA-256
656c934ca6ebf0fa302c32a3694841218b111addae123bded68bd674f81eff41, preserving
the mTLS status/result/cancel transport. Snapshot
01a0a222-a6dc-7288-9845-f709ead64d8e is current again; only endpointRules differed
from the deactivated enrollment snapshot. The editable instance property still
contains the enrollment assignment; reconcile it before capturing another snapshot.
The earlier enrollment request was submitted once, but no user approval or native invocation was submitted by the agent. Do not reuse its expired browser link. The older issuer/broker interactive contract has not yet been replaced with backend reuse of existing scope authorization. Removing the Portal prompt alone does not complete that integration. A1’s scheduled and hours-long renewal tests are waived by the user, not passed; Phase 1 remains incomplete.
Verification: Portal dialog/client/run-controls tests 20 passed; MCP module tests 175 passed and 3 ignored; lifecycle catalog regression 1 passed; focused ESLint and diff whitespace checks passed. Live Portal no longer shows the enrollment prompt; Gateway is running with exit code zero after restoration. No database wipe, commit or push was performed.
Historical backend acquisition correction (2026-09-14)
This supersedes the remaining acquisition gap in the preceding rollback note. At that checkpoint, root HTTPS Gateway invocation without a supplied grant called Workflow’s enrollment API internally. The API acquired and redeemed a one-time issuer code over mTLS using existing Portal scope authorization and returned only a grant ID. There was no new consent Tool, browser redirect, or additional login. This acquisition path was later retired by Workflow Invoke.
The former acquisition path validated the exact source access token against its issuance audit, active authorization-code session, client scope and user authority. It recorded token fingerprints rather than bearer tokens. Older access tokens needed an ordinary Portal refresh before that path could use them.
The short isolated real-mTLS test covers initial and refreshed-token acquisition, renewal, scope expansion, wrong Host, missing provenance and source revocation. OAuth unit tests passed 22/22 executed; Workflow 86/86; MCP 175/175 executed (2 OAuth and 3 MCP ignored tests are not counted as passes). No hours-long or scheduled qualification was run, per user waiver. The production Portal/native invocation sequence was not run by this correction; Phase 1 is still incomplete.
Local container-layer binaries installed:
- OAuth
/app/service:1b975116e94a344f86366e7fd8f4585abe10326cceda49afaf3c0fb130891c0c - Workflow
/app/light-workflow:1db52645044f907127f946f8bf206c040279767d78caf7ecf54fd8e4a8bc2b5d - Gateway
/app/light-gateway:88c304cc0929ad26e018d956d3a509d2bf92d9445422ed142f5d5134d64ef639
All three were running with zero restarts after deployment. The unauthenticated
live Workflow enrollment probe returned 403. Backups are under
/tmp/phase1-backend-acquisition.4FBLYy/; no application database reset occurred.
At the time, these replacements had not been incorporated in normal images.
No commit, push or Phase 1 completion comment was made at that checkpoint.
Native qualification and policy upgrade (2026-09-15 UTC)
Native attempt 01a0a2ce-b655-7c12-89b2-57c1110a633e reached the worker,
but failed before a model turn with Codex qualification evidence digest mismatch.
Execution 01a0a2ce-c261-73d1-a317-20c84750186b reported FAILED with
CONFIRMED cleanup. This is dispatch/preflight evidence, not native success.
The Codex Workflow authoring profile was corrected by importing event
01a0a2d1-4a7f-72f0-bcbf-9bee159d8f23 from
light-portal-event/genai/20260915-codex-workflow-evidence/events.json.
Only its qualification evidence digest changed, to the checked-in contract digest
sha256:5bea40c988edd30a30aa7cd25e0be4ccc69be54fa66229f940769515fd39787b.
Publication version 3 produced snapshot a732ced2-8bb4-3f2f-a5ed-792256a3d4de.
Although the UI reported synchronous projection failure, the event committed and
the asynchronous projection completed; preview then reported CURRENT. The
publication was not blindly retried.
Restart exposed missing accepted-version evidence for Workflow-only Agents.
initialize_workflow_policy now pins AGENT_POLICY reference evidence against
agent_policy_snapshot_t, and monotonic runtime-scope upgrades recognize that
source as well as interactive sessions. The real disposable-PostgreSQL regression
passed, including repeat startup, forward upgrade and downgrade rejection;
cargo check --locked -p light-agent --lib and the musl release build passed.
Local Codex Workflow Agent uses image
networknt/light-agent:phase1-policy-upgrade-20260915; its scoped recreation
overlay/script is under /tmp/phase1-policy-upgrade.HdE1ir/. The retained prior
snapshot was briefly selected to let the fixed runtime record its genuine accepted
version-2 evidence, then the corrected version-3 snapshot was restored. After
restart the Agent was running, and both accepted publication versions were verified
in PostgreSQL. No fabricated reference evidence or database reset was used.
Attempt 01a0a2d5-ec58-73e1-848b-6fb855e9c9ba, accepted while the Agent
was down, was cancelled through Portal. Its feature
21342de0-074b-4f92-b705-0ed4dd16aa91 still owns personal VM generation 4.
A subsequent Portal lifecycle read returned HTTP 401: current subject authorization
no longer matches the accepted disclosure ceiling. Do not release this fence by
hand or start overlapping work. Session authorization and normal cancellation
reconciliation need qualification next. Native success, snapshot transfer,
in-flight cancellation/restart, GitHub publication and installer end-to-end
qualification remain incomplete. No Phase 1 completion comment was posted.
Offline Agent cancellation delivery (2026-09-15 UTC)
Signing out and in again did not restore lifecycle disclosure for historical run
01a0a2d5-ec58-73e1-848b-6fb855e9c9ba; the exact accepted-claims digest guard
remains unchanged. A separate cancellation-delivery bug was proven: Workflow’s
poll excluded cancelled/expired PENDING jobs, so an Agent offline at acceptance
could never learn of them and acknowledge cleanup.
The pull message now carries optional cancellationRequested (default false).
The existing mTLS/app/Host/Agent checks still select the matching Agent. Cleanup-only
delivery skips execution authorization, not peer authentication; Agent validates
immutable job identity and records cancellation atomically with admission. Its
turn-creation query excludes cancellation requests. Previously admitted turns
retain their Controller identities and require existing positive cleanup proof.
Workflow never treats its own PENDING state as non-dispatch evidence.
Validation: three transport tests, the real disposable-PostgreSQL Agent regression
(early/expired cleanup, replay and existing-turn identity), and the exact cleanup
receipt test passed. Both musl application builds passed. Local images are
networknt/light-agent:phase1-cancel-delivery-20260915 and
networknt/light-workflow:phase1-cancel-delivery-20260915; recreation overlay is
/tmp/phase1-cancel-delivery.C3LNhy/compose.yml.
After deploying Workflow and both dedicated Workflow Agents, the historical job
reported FAILED (deadline exceeded) with not-dispatched cleanup. Normal
reconciliation released personal VM generation 4 and advanced it to free generation
5, without administrative database writes. This qualifies the offline-before-
admission cancellation path, not in-flight native cancellation.
Fresh inspect-only run 01a0a2e7-3720-7872-ba49-1e68dc775505 was accepted through
Portal and its lifecycle status read succeeded using the refreshed session.
Feature b6ce8b8a-c72b-4fc8-a3fb-741f853512fc reached native execution
01a0a2e7-3eee-76c2-b004-fd5c74973473. Both Codex job
01a0a2e7-3766-7191-bd13-c61d3b50d2b3 and Claude job
efe47edd-823d-4c2f-b5fe-03bf2724d331 succeeded. Both Controller attempts
reported SUCCEEDED / CONFIRMED cleanup, the workflow completed, and Portal
displayed Claude’s acknowledgement. Normal post-qualification cancellation
released the development feature and advanced the VM to free generation 6.
This qualifies successful native inspection for both engines, not editing or
publication effects.
Portal refresh disclosure correction (2026-09-15 UTC)
The fresh run reproduced lifecycle HTTP 401 after its initial successful status
read. Portal’s RefreshCoordinator::renew generates a new csrf value at every
refresh, but the stable subject digest retained this nonce. The correction excludes
only that nonce from disclosure normalization; HTTP CSRF validation is unchanged.
Historical accepted claim objects are checked against their original digest before
being normalized for comparison. No stored authorization evidence is rewritten.
Identity, Host, client, role and scope changes still fail closed.
Ten contract/fixture tests, 18 Workflow API tests and 10 Gateway workflow tests
passed, plus both musl builds. Local Gateway/Workflow images use tag
phase1-csrf-disclosure-20260915; the overlay is
/tmp/phase1-csrf-disclosure.muCAoO/compose.yml. The existing completed run became
readable through Portal after deployment with no additional login or consent.
Two snapshot component tests (chunk recovery/corruption and manager scoping/replay)
also passed, but native snapshot-transfer/restart qualification remains separate.
In-flight native cancellation (2026-09-15 UTC)
Run 01a0a2ef-8227-7561-b78c-f4901fbd5bcd, feature
9968a37f-f222-49c3-bebd-188942639abb, was accepted for the same bounded
inspect-only definition. Controller execution
01a0a2ef-999b-79e1-b063-e10370a30af6 was verified STARTED before Portal
cancellation. Portal returned CANCELLED at 2026-09-15T02:40:41.185780Z.
Controller then reported CANCELLED / CONFIRMED cleanup; personal VM advanced to
free generation 7. Workflow was restarted immediately after the cancellation
request, and the released state survived. Because cleanup completed quickly, this
is not evidence of crashing a runner while cleanup is still pending.
Remaining live gates include native snapshot transfer/restart, interruption during unfinished runner cleanup, GitHub publication, and installer application end-to-end qualification. Phase 1 remains incomplete; no completion comment, commit or push was made. Configuration data was preserved throughout.
Snapshot qualification preparation (2026-09-15 UTC)
Snapshot requests previously required the caller to know the server-generated
stage execution ID. Native enqueue now fills an omitted managerSnapshot.stageId
from the accepted stage receipt before computing the immutable job digest.
Explicit mismatches fail; existing pinned-task, active-feature and VM-owner checks
remain. The focused native-job tests pass, including exact replay and malformed or
conflicting stage rejection. This source change has not yet been deployed.
Read-only sizing of the successful native task found 128 repository checkouts and
381,228,402 bytes of tracked Git blobs. Snapshot packages retain complete file
contents and enforce a 33,554,432-byte serialized bound, so the existing personal
task is unsuitable for this bounded transfer qualification. Do not raise the cap
or omit repositories silently. A separate qualification workspace restricted to
the approved networknt/light-agent repository is proposed; its registration and
matching policy/runner binding require operator approval. No new events were
imported and no GitHub effects were performed in this continuation.
Native snapshot transfer qualified (2026-09-15 UTC)
The subsequent owner-approved isolated phase1-qualification workspace contains
only networknt/light-agent. Definition v1.0.5 and its Gateway binding are active.
Run 01a0a32e-d304-7112-b51e-896b8e0b1b19 completed a three-chunk, 323591-byte
snapshot transfer. All worker reports were SUCCEEDED / CONFIRMED cleanup.
Portal returned transferComplete: true; retained artifact
d9c2397c-95e0-78cc-976c-32c4ad898eca is VERIFIED / BOUND / RETAINED. Independently
hashed stored bytes match receipt SHA-256
9562fcf4727f80baafd309ee3f8bb420ac8852e6913a5681ab5817fb16f945f7.
Fixes preserve owner-bound workspace authorization and strict snapshot-only
request/result shapes. Normal model-output requirements no longer apply to a
fixed manager read. Workflow compares typed requests so absent/null optional
first-chunk package digests do not prevent reconciliation. The saved first
report was accepted after a Workflow restart without rerunning that worker.
Both later chunks completed, and qualification cleanup released VM personal
at generation 10. The live transfer used the Codex runner; Claude compatibility
is covered by the shared code change, not a separate live transfer claim.
Verification: 41 runner unit tests, one runner WebSocket integration test,
optional-digest regression, and snapshot chunk corruption/recovery test passed.
Detailed event, deployment, failure, and passing evidence is in sibling repo
light-portal-event/workflow/20260915-snapshot-qualification/README.md.
Daily automation, interruption during unfinished runner cleanup, GitHub
publication, and installer application end-to-end gates remain separate.
Phase 1 remains incomplete; no completion comment or commit was made.
Unfinished cleanup recovery correction (2026-09-15)
Runner Supervisor::recover discarded backend cleanup errors and then assigned
CONFIRMED solely because a backend operation ID existed. The new regression
failed against that behavior. Recovery now persists CLEANUP_REQUIRED and returns
an error when bounded cleanup retries fail, without publishing a terminal result
that could falsely authorize release. A later recovery retries cleanup, not the
original execution; only positive cleanup permits an UNKNOWN terminal outcome
with CONFIRMED cleanup.
Regression coverage includes reopening the SQLite journal after failed cleanup, successful later cleanup, and killing a child process after it durably records unfinished cleanup. Reopening that killed process’s journal with unavailable backend evidence leaves cleanup pending and emits no confirmation. The child test uses an injected backend, not a live Controller/VM or model worker.
After the user’s full local rebuild/restart, both dedicated personal Workflow
runners were verified idle and updated to binary SHA-256
3eb58ccfdd948a8a66c5c5efc89e2313e6b11534aeb9e52844e1e8a0e7ecbf69.
Only the two runner binary admission hashes changed; policy/configuration hashes
were preserved. Controller restarted to load admission, and both runners
reconnected healthy. Previous binaries/admission are retained under
/tmp/phase1-cleanup-runner.wQTyT1 for this local deployment’s rollback.
The automatic-login daily snapshot gate passed against the rebuilt Compose stack
and updated runners: invocation 01a0a51c-7da5-7d81-bd04-ce0d034f1019, report
light-portal-test/reports/snapshot-qualification/daily-dbc66af1-6755-4d24-bba0-2a6be0117ef0.json.
The report confirms artifact verification and VM cleanup. The isolated child
process-kill regression also passed after the build. Full live unfinished-native-
cleanup/Controller-fencing qualification remains pending; neither this regression,
normal snapshot smoke nor the earlier in-flight cancellation qualifies it.
Native cleanup interruption prerequisites (2026-09-15)
Tracing the live native path uncovered a second cleanup-proof bug:
result_accepted treated a Controller acknowledgement as local cleanup evidence,
cleared the terminal payload and marked CLEANUP_CONFIRMED even for a result with
FAILED cleanup. The new regression failed before the fix. Acknowledgement now
rejects unconfirmed cleanup without changing its durable terminal payload/backlog;
staged-file cleanup also runs before clearing that payload. This additional
source change is not yet deployed.
The full native interruption gate is still blocked on durable native-process
recovery: native executions currently persist an agent-worker:<execution-id>
operation identifier, while restart recovery delegates that identifier to the
configured backend. It is not evidence that the native process group has stopped.
Before deliberately interrupting the live worker, implement durable process or
supervisor-owned containment identity and positive cleanup proof, including
PID-reuse/unknown-evidence rejection and replay-safe Controller cleanup delivery.
Do not force-clear a reservation or present acknowledgement, missing backend
state, process timeout, or the isolated injected-backend test as that proof.
Native launch now captures Linux PID/start ticks, boot ID, PID namespace and cgroup membership, then persists them in a separate SQLite native-process journal before sending execution input. The identity is immutable and tied to the exact execution/lease/fencing token. A failed capture or journal write terminates the just-spawned worker instead of delivering execution input. Tests cover process stat parsing, live identity capture, durable reopen, exact replay, conflicting process identity, stale fencing, and missing historical evidence.
This is identity evidence only: a missing PID or empty process group is not proof that descendants which created new sessions have stopped. Positive containment cleanup/recovery and Controller cleanup receipt delivery still need implementation before the live interruption gate. These changes are not deployed yet, and no live native worker was interrupted for this test.
Delegated native containment deployment (2026-09-15)
With explicit owner approval, only light-workflow-runner-personal and
light-workflow-runner-claude-personal received Delegate=yes drop-ins named
zz-native-cleanup.conf; existing KillMode=control-group remains. Opt-in
agentWorker.nativeCgroup: true creates a separate execution cgroup and attaches
the worker before execution input. Its recorded containment identity binds the
execution UUID, parent cgroup, inode, boot and cgroup namespace. Cleanup targets
only that child, uses cgroup.kill, and requires populated 0. Normal native
completion and nonterminal restart recovery use this evidence, not mock-backend
absence. Historical missing containment evidence fails closed.
The isolated qualify-native-containment example passed in a disposable delegated
systemd service: a worker and setsid descendant were removed; repeated cleanup
passed. All 50 runner unit tests and the WebSocket integration test passed.
This component test is not live Controller/VM interruption proof.
Deployed runner SHA-256:
58af49a73582fed11d6351c72096b67d9fc991c96c6bf7b0426709d9d02e0b13.
The two binary/config admission hashes were regenerated, Controller restarted,
and both dedicated runners reconnected ready. Backups of prior runner/configs/
admission are under /tmp/phase1-cgroup-recovery.QPhrDE.
The ordinary live snapshot gate passed with containment enabled; report:
light-portal-test/reports/snapshot-qualification/daily-e259bce5-514d-47c1-a984-0e538d3d3631.json.
All three execution cgroups reported unpopulated after cleanup. Deliberate live
interruption during unfinished cleanup and Controller-fenced VM release remain
unqualified; do not infer those results from the normal snapshot gate.
Operator-authorized retirement of historical UNKNOWN (2026-09-15)
With explicit owner approval, reset-native-scope --confirm-fence fenced only
the dedicated Claude user service. It checked the exact configured process,
unit invocation, boot, cgroup namespace and inode, stopped the control-group
unit, and required an inactive unit and empty/absent kernel cgroup. A separate
durable operator confirmation remains bound to the original execution, lease,
fence and terminal digest; this is not an automatic model retry or a database
override. Controller authenticates the matching enrollment and current session
before accepting operator-unit-fence:<sha256> cleanup evidence.
Runner tests: 55 passed, with additional acknowledged-proof/reopen/stale-lease
regression passing. The real PostgreSQL Controller fencing test passed in
disposable database phase1_operator_controller_1789499514520.
Live failed feature 5a961638-261b-4b1f-9c64-7797c09d11f1 was retired through
the normal owner cancellation flow: state cancelled, VM generation 31 released.
The runner reconnected with zero cleanup backlog. No uncertain turn was replayed
and no Configserver database was wiped. This qualifies historical operator
retirement, not deliberate interruption of a newly running native worker.
The fresh publication cycle and full installer application qualification remain pending; Phase 1 is not complete.
Enterprise Development Workflow Orchestration
Status: proposed corporate deployment design, separated September 13, 2026. The coding-harness gateway/broker contracts provide a foundation; this document does not claim complete enterprise workflow, token-service, accounting, or deployment qualification.
This document covers codex-enterprise workers inside a corporate network,
using llm-gateway for model access. Developer-owned Codex/Claude subscriptions
and local VM workflows are covered by
Personal Development Workflow Orchestration.
Decision And Shared Lifecycle
Use light-workflow for durable feature stages, review closure, budgets,
scheduling, approvals, and GitHub coordination. Use light-agent as the
authorized Agent/session service, and runner-managed light-agent-worker
instances for repository work. The trusted launcher configures the Codex harness
with an admitted custom Responses provider pointing to llm-gateway.
codex-enterprise describes this execution configuration. Actual Agent
definitions, model aliases, runner pools, policies, and credentials must be
published and admitted; the name alone does not prove a deployment is ready.
Reuse the personal document’s versioned business contracts:
- existing-issue or requirement-dialogue intake and an immutable requirement;
- design review and optional exact-digest human sign-off;
- optional plan and phase-by-phase implementation/review;
- durable feature ownership and atomic stage handoffs, even for separately started workflows;
- diffable candidate snapshots and retained before/after evidence for resumed reviews;
- workflow-owned finding IDs and reviewer-verified closure;
- initial full final reviews, resumed fix verification, and explicit coverage escalation instead of restarting all reviewers on every edit;
- fixed document revision publication, implementation commit/push/PR actions, per-repository recovery receipts, and explicit completion targets.
The personal document’s filesystem backend and one-feature-per-VM reservation are deployment choices for its pilot. Enterprise deployment qualifies its own artifact backend and pool admission/release policy, preserving durable snapshot export, verified recovery, and confirmed fencing before capacity is reused.
Author, phase reviewer, primary final reviewer, and independent final reviewer
are separate logical roles. They may use separately configured
codex-enterprise executions with different admitted model routes; they never
share implementer-private conversations. The enterprise design does not require
claude-personal or a personal login. A second model/harness route is eligible
only after the enterprise protocol and isolation gates pass. Independent roles
remain required even if the selected models share a vendor.
The same lifecycle can run as top-level stages connected by accepted artifacts. A parent/child coordinator is an optional later runtime feature, not a reason to fork the business state machine. Profile changes preserve common transitions and artifacts while adding enterprise authorization and budget gates. Switching profile mid-stage requires new admitted execution/session bindings; it cannot reuse a personal conversation as an enterprise session.
Architecture And Authority
flowchart TD
U[Corporate user] --> P[Portal and API ingress]
P --> W[light-workflow]
W --> A[light-agent: enterprise roles]
A --> C[Controller and corporate runner pool]
C --> B[Attempt broker and trusted sandbox launcher]
B --> X[Codex enterprise worker]
X --> L[llm-gateway Responses endpoint]
L --> M[Admitted model provider]
W --> T[Fixed validation and Git publication services]
| Component | Authority |
|---|---|
| Corporate identity/authorization service | Verify human identity, workload delegation, revocable grants, and approver scope |
light-workflow | Common lifecycle, feature-ready queue/fairness, review ledger, aggregate budgets, approvals, and publication intents |
light-agent | Immutable Agent policies, turn admission, durable jobs/results, and role/model selection |
| Controller/runner | Pool admission, capacity, placement, leases, cancellation, and sandbox lifecycle |
| Attempt broker/launcher | Bind a single execution to allowed routes and short-lived credentials |
llm-gateway | Authenticate model requests, resolve logical aliases, protect provider keys, enforce limits, and issue usage evidence |
| Fixed test/publication services | Validate exact candidate contents and execute authorized external effects |
| Artifact/audit stores | Retain immutable work packages, review coverage, usage/publication receipts, and decisions |
light-gateway is the Portal/API ingress and policy edge; llm-gateway owns
model routing and provider access. They are different authorities even if a
deployment packages them together. Neither a model nor a repository config
selects the acting identity, billing subject, provider credential, or publication
target.
Corporate Network And Worker Isolation
Run pooled or dedicated workers in admitted per-attempt sandboxes. Corporate
placement alone is not isolation. Pin the worker binary, capability digest,
launcher/profile digest, repository inputs, model route, and allowed tools.
The enterprise pool requires its qualified sandbox-launch contract and model
egress restricted to the configured llm-gateway; direct vendor fallback is
not an implicit recovery path.
Use the qualified repository materialization contract. A pooled execution must
not inherit another user’s writable checkout, Git administration, build scratch,
credential store, or model transcript. Review reads the exact candidate through
an isolated read-only materialization with explicitly excluded build scratch.
The current personal shared-workspace configuration rejects an enterprise broker
and sandbox launcher, so its local workspaceConfig cannot simply be enabled
on an enterprise pool. A managed enterprise workspace path needs its own
admission and qualification.
Keep reusable provider/Git credentials in trusted services. The Codex process
may receive only its admitted short-lived attempt credential through the
qualified broker/launcher path; shell/tool/MCP children must not inherit it.
Do not mount personal Codex/Claude login stores or route subscription credentials
through llm-gateway.
Corporate Git hosting, registries, package mirrors, test services, and artifact
storage follow explicitly admitted endpoints. The existing local workspace
GitHub helper recognizes github.com clone sources; support for a corporate
GitHub Enterprise host requires validated host/repository handling and its own
qualification, not an assumption that local gh setup already covers it.
Model Routing Through llm-gateway
The trusted launcher configures Codex’s custom Responses provider. The harness
sees logical aliases such as coding-implementer and coding-reviewer;
llm-gateway resolves them to eligible provider deployments. A requested model
or repository-local config cannot replace the provider, endpoint, credential
source, or route policy.
Use the attempt broker contract in Coding Harness Integration. Bindings cover Host/user, workload actor, workflow/stage, Agent session/turn, execution attempt, model route, authorization and billing policy. Credentials have an audience, generation, expiry, and revocation state. The broker validates and re-reads the current envelope when serving the attempt.
Qualify each admitted route for the pinned Codex/Responses contract: buffered and streaming output, tool calls/results, completion ordering, cancellation, errors, and usage. A provider being reachable is not evidence that its model or protocol transformation can operate the coding harness. Ineligible routes fail admission; no fallback to personal billing is allowed.
Native session continuity still follows the common new/resume/close contract. A model, policy, sandbox, or scope change that invalidates a binding requires explicit replacement from accepted artifacts. An uncertain execution is fenced, not retried under another provider or identity.
Identity, Delegation, And Reauthorization
Use the shared User, Application, And Workflow Authorization
contract as the issue #374 prerequisite. Preserve the verified initiating human
and the immediate calling workload separately. Forward valid user credentials
where the issuer permits; unattended execution obtains fresh issuer-backed
credentials under a revocable grant, with the caller’s app token in
X-Scope-Token.
The enterprise authorization envelope additionally binds the following evidence; execution/accounting fields need not be embedded in the user identity token:
- issuer, audience, token ID, issue/expiry times, and tenant/
hostId; - verified human subject and acting workload/principal;
- feature/run, stage, Agent/session/turn, and action/attempt correlation;
- policy, data-boundary, route, and caller-claims digests;
- server-selected
billingSubjectandbudgetPolicyId.
Store verified identity metadata, grant references, and digests with workflow
state, not reusable bearer tokens. A revocable continuation grant authorizes
fresh credentials when the original user JWT expires. Expiry/revocation enters
REAUTHORIZATION_REQUIRED and stops new model/tool spending until authorized.
Possessing a workflow ID or an old session checkpoint is insufficient authority.
Workflow Admin and Chat starts must produce equivalent verified attribution. A service acting for a user records both actors; it does not replace every user with the Agent service account. Scheduled continuations retain the originating grant or an explicitly authorized service budget and actor.
Usage, Reservations, And Cost Attribution
llm-gateway is the target authority for normalized provider usage and cost.
Before dispatch, reserve the allowed request budget under the verified user or
cost center. Reconcile input/output/cached/reasoning tokens when supplied by the
provider, charged cost, route/deployment, request/attempt identity, and receipt
completeness. Do not treat missing usage as zero or accept model-reported usage
as accounting evidence.
The durable ledger and signed receipts bind usage to the same user, actor, workflow, session, turn, route, and attempt used for admission. Retries and fallbacks retain distinct provider-attempt records while preventing duplicate charges for replay of a completed attempt. Incomplete evidence follows a stated conservative reservation or reconciliation policy.
Agent and workflow budgets may be narrower than gateway limits, but reconcile
against trusted receipts rather than overriding them. A denied enterprise
reservation enters BUDGET_EXHAUSTED; a later authorization may resume work.
This is an enterprise-only state. Personal subscription runs retain bounded
turn/round/time limits and do not synthesize a trusted token/cost ledger.
Receipt validation in a worker contract is not a production ledger service. Token exchange, durable reservation/usage storage, receipt persistence/emission, and long-running reauthorization must be separately demonstrated in deployment.
Scheduling, Approvals, And External Effects
light-workflow owns fairness among its feature-ready turns, using the same
persisted queue ownership as the personal design. Agent jobs carry admitted
execution/results; Controller and runner enforce actual pool capacity. Add
tenant/user admission limits and deadlines without creating a competing
feature-fairness queue in each downstream service. Gateway request limits
remain a separate model-resource constraint.
Enterprise authorization may require design, security, publication, merge, signing, or deployment decisions. Bind each decision to the exact action intent, candidate/review-coverage digest, allowed response, approver scope, expiry, and idempotency identity. Honor existing applicable grants; normal progress comments do not create a new approval requirement at every turn.
Worklist owns workflow human-task assignment, claim/release, expiry, and completion. If Chat renders the same task, it submits to the same authoritative record. An Agent must not invent a second approval or advance a task from a natural-language “yes.” Separate direct-agent interactions retain their own Agent authority and dedicated decision record.
Fixed publication services use scoped enterprise Git credentials and the common commit/push/PR receipt contract. They validate current review coverage, including accepted delta transitions, before committing the exact approved files. Candidate drift, unexpected remote refs, or incomplete multi-repository delivery blocks completion. Keep document revisions and supersession explicit; the initial immutable revision-task strategy needs no force push.
Current Boundary And Qualification Gaps
The coding-harness documentation records implemented gateway/attempt-binding,
role-isolation, cancellation, and broker contracts. Source locations include
apps/light-workflow-runner/src/broker.rs, worker_process.rs,
configuration.rs, and the coding runtime. These are integration foundations,
not proof that every corporate service named here is deployed.
Before enabling an enterprise feature workflow, verify:
- Workflow-to-Agent catalog/job visibility and complete durable turn/result propagation in the actual operational database topology;
- publication of enterprise Agent definitions, qualified model routes, runner pools, delegated identity bindings, and protected network endpoints;
- real token exchange, revocation, continuation grants, and normalized usage ledger/receipt services, not only contract fixtures;
- clean enterprise repository materialization and isolation across users and concurrent attempts;
- common stage/finding/review-coverage/comment/publication contracts, including uncertain execution and partial external-effect recovery;
- corporate Git host and artifact/test infrastructure compatibility.
The broader shared Portal UI/CLI and distribution work described below is rollout work. Its presence here does not make it a prerequisite for the personal pilot.
Enterprise Delivery Plan
Phase E0: Reuse And Freeze The Business Contracts
Adopt the common standalone-stage inputs/results, finding IDs, review coverage, budgets, and fixed publication contracts. Publish separate author/reviewer roles and their admitted enterprise model routes.
Exit gate: identical requirement/phase/finding fixtures produce the same common lifecycle decisions in personal and enterprise profiles; only documented authorization/accounting gates differ.
Phase E1: Identity And Usage Services
Implement/qualify token exchange, continuation grants, enterprise reservations, normalized usage storage, signed receipt persistence/emission, and reauthorization. This phase owns those services for the coding-harness enterprise integration.
Exit gate: verified end-user/workload attribution survives start, retry, restart, and grant renewal. Cross-user/tenant, expired/revoked grant, quota denial, duplicate attempt, and incomplete-usage cases enforce the intended outcome without exposing reusable credentials.
Phase E2: Corporate Runner And Gateway Qualification
Admit the enterprise sandbox, broker, restricted model egress, repository inputs, and corporate Git/test/artifact endpoints. Qualify each model route against the pinned coding-harness protocol and actual accounting services.
Exit gate: buffered/streaming/tool/cancellation/error runs produce bound receipts; parallel attempts cannot see each other’s files, credentials, or conversations. Missing isolation or an unqualified route fails admission.
Phase E3: Feature Delivery And Operational Recovery
Run the shared design/plan/phase/final lifecycle with enterprise approvals, fixed commit/push/PR services, and the selected completion target.
Exit gate: real model calls and Git effects complete with correct user/cost attribution. Worker/service restarts, revoked grants, pool contention, and partial publication reconcile without duplicated edits/effects or lost usage evidence.
Shared Portal Rollout Extensions
The removed broad UI/CLI and installer requirements remain follow-up work, separate from either profile’s first functional loop:
- Workflow Admin supplies authoring, start/status, pause/resume, and cancellation.
- Worklist supplies durable human-task ownership and schema-bound decisions.
- Future Chat cards may render structured choices and authoritative task links; they do not change decision ownership.
- A future thin CLI uses those same public APIs and decision contracts rather than database access or an independent approval mechanism.
portal-config-loc/all-in-ltandlight-portal-installshould eventually pass the same personal-profile lifecycle/API conformance suite. Packaging differences must not alter review or publication semantics.- Corporate deployments qualify their own identity, gateway, approval, and operational topology in addition to the common lifecycle suite.
These extensions need separate rollout acceptance and do not imply that Chat
cards, a CLI, or both local installers must be built to run feature-design.
References
- Personal Development Workflow Orchestration
- Coding Harness Integration
- Light-Agent Execution
- Workflow Coding Thread Lifecycle
- Agent LLM Dual-Token Authorization
- LLM Gateway API
Shared Task Workspaces
Status: proposed design, September 10, 2026. This document specifies new workspace execution behavior; it does not claim that the current bundle-based coding request, runner, or Claude adapter implements it.
Implementation has started in crates/task-workspace and apps/light-workspace.
The local workspace service guide describes
the runnable CLI/MCP service, qualification, and remaining integration gaps.
The Chat and Workflow Integration design
defines the next implementation: authenticated workspace jobs, interactive task
selection, durable agent handoff, and deployment gates.
This design extends Development Workflow Orchestration and Coding Harness Integration. The workflow remains the lifecycle authority. The runner owns local workspace materialization, process access, and exclusive mutation. Agents perform implementation and review through the existing adapter boundary.
Decision
Register a persistent workspace on a runner host containing a collection of Git repositories. Granting an agent access to this workspace grants access to all repositories in it. There is no per-repository permission matrix. Choosing which repositories a task changes is planning information, not an authorization boundary.
For each task, create a task workspace containing a Git worktree for each
registered repository. Each worktree checks out that repository’s task branch,
created from its recorded develop revision. A branch is a Git ref; a worktree
is the directory where that branch is checked out. They are related but are not
the same object. Git manages worktrees independently for each repository; the
runner groups them into a single multi-repository task workspace.
Codex and Claude can use the same task workspace and see the same files, including uncommitted changes. Only one execution may write to a task workspace at a time. Independent tasks use different worktrees and branches and can run in parallel. Branches belong to tasks, not permanently to an agent or model.
The normal delivery sequence is:
Existing or new issue
-> task branches from develop
-> implementation across repositories
-> stable review of uncommitted changes by another agent
-> remediation and renewed review
-> commit and push task branches
-> related PRs targeting develop
-> integration testing on develop
-> release PR from develop to master
Workspace layout and identity
Example managed layout; actual roots are deployment configuration:
/workspaces/portal/
repositories/ # runner-managed source repositories
light-fabric/
light-portal/
portal-view/
tasks/
task-384/
light-fabric/ # branch agent/task-384 in light-fabric
light-portal/ # branch agent/task-384 in light-portal
portal-view/ # branch agent/task-384 in portal-view
task-391/
light-fabric/ # branch agent/task-391 in light-fabric
light-portal/ # branch agent/task-391 in light-portal
portal-view/ # branch agent/task-391 in portal-view
The branch names may match across repositories but their histories are separate. Use a stable task identifier, not an issue number alone: issue numbers can collide across repositories. Record issue identity as repository plus issue number. The example task identifiers are illustrative.
A workspace registration records its owner/Host, runner binding, canonical root, repository identities and remotes, default integration/release branches, and indexing configuration. Repository membership is versioned. A task pins that membership revision so adding a repository does not silently alter an active review. An explicit refresh may add newly registered repositories to the task; it invalidates review evidence. All repositories in the pinned membership are available to every agent granted access to that task’s parent workspace.
The runner creates a worktree and task branch per repository without switching branches in the source checkout. Unchanged repositories need no commits or PRs. Creation records the fetched base commit independently for each repository; there is no global Git revision shared across repositories. Provisioning is idempotent and partial creation is journaled. A retry verifies existing worktrees and ownership instead of overwriting directories or moving branches. Existing user checkouts and their dirty changes are never adopted implicitly.
Ownership and proposed records
Names below describe proposed contracts, not existing API fields or tables.
| Record | Essential state |
|---|---|
| Workspace | Workspace identity, Host/owner, runner, managed root, membership revision, repository catalog, agent grants |
| Task workspace | Workflow/task identity, parent workspace and membership revision, directory, lifecycle state |
| Repository checkout | Repository identity, worktree path, task branch, base ref and commit, current HEAD, remote tracking state |
| Write lease | Task workspace, execution owner, fencing generation, expiration and renewal state |
| Review checkpoint | Repository membership, all HEADs, index and working-tree content manifest, untracked files, aggregate digest |
| Review result | Checkpoint digest, reviewer identity/adapter, findings, disposition and test evidence |
| Delivery record | Per-repository commits, pushed refs, issue/PR identifiers, expected remote revisions, effect idempotency keys |
light-workflow records durable progress and dispatches bounded jobs. The runner
resolves workspace identifiers to registered paths; browser-supplied absolute
paths do not grant access. The worker receives its task workspace, operation
mode, and lease/checkpoint identity. Authentication and adapter choice remain
separate from workspace identity, so an authorized claude-personal can review
a task implemented by codex-personal.
Shared implementation and review
- The implementer obtains the task workspace write lease and edits/tests any repository in the workspace. Other agents may inspect it, but observations made during writes are provisional and cannot constitute review approval.
- Before review, the runner stops or waits for all mutating processes, including background tools, and captures a checkpoint across every repository.
- The reviewer receives a fresh execution context, task requirements and the checkpoint. It reads the same worktrees, including staged, unstaged, deleted, binary and untracked files. Hidden conversation state is not transferred.
- Review mounts/access are read-only. Tests that generate files run in a disposable copy of the checkpoint with separate build output. A reviewer asking to edit must enter a remediation stage and obtain the write lease.
- Findings are bound to the checkpoint digest. Remediation releases the review freeze, grants a new writer lease, and invalidates the previous approval.
- Once review and required tests pass, the trusted commit action verifies the checkpoint again and stages/commits exactly its intended file manifest. Untracked intended files are included; local caches and generated index data are excluded. Any unexpected mutation blocks finalization and requires review.
A checkpoint cannot be represented solely by git diff: it must cover HEAD,
the index, working-tree content and untracked intended files in each repository.
Ignored files are not implicitly deliverable; promotion of an ignored file must
be explicit and appear in a new checkpoint. Submodule changes need explicit
repository registration and coordinated checkout handling; a parent worktree
does not automatically materialize or authorize an external repository.
The write lease is enforced by runner-managed process lifetime and filesystem access, not by advisory text in an agent prompt. Lease expiration alone does not permit a new writer: the runner must fence or terminate the previous execution and confirm quiescence. If that cannot be confirmed, the task stays blocked for recovery. During review, even the prior implementer must lose write access.
Worktrees share Git metadata and objects with their source repository. Serialize operations that mutate shared repository administration, such as fetch, worktree creation/removal and maintenance, with a repository-level lock. Ordinary edits in independent task worktrees remain concurrent. Git operations that can affect other tasks must use the runner’s authorized tooling; filesystem write access to all shared Git administrative state must not be handed to arbitrary workers.
Git and GitHub operations
Agents can request clone/fetch/checkout, branch, commit, push, issue creation or updates, and PR operations through authorized tools. Host-owned credential helpers or GitHub integrations execute these actions without placing reusable credentials in prompts, repository files, or task artifacts. Workspace access covers every repository; operation-level policy still distinguishes reading, editing, publishing, merging and release actions.
The existing orchestration design’s trusted fixed actions remain the mechanism for externally visible mutations. This supports agent-driven GitHub work without requiring arbitrary credential-bearing shell execution. Standing workflow policy can authorize routine actions; an additional user confirmation is needed only when the action exceeds that authorization or an explicit gate requires it.
For a new task, create an issue if requested by the workflow; otherwise link the
existing issue. Review occurs before committing locally. After approval, commit
and push each changed repository’s task branch, then create or update one PR per
repository targeting develop. Link related PRs, their issue(s), test evidence
and cross-repository merge dependencies. Persist returned GitHub identifiers so
retries reconcile completed effects instead of duplicating issues or PRs.
Changes after review, including conflict resolution, rebasing, merging a newer
develop, or hook-generated edits, require a new checkpoint and the applicable
review/test gates. Validate final commit content against the reviewed manifest.
Do not automatically force-push over unknown remote changes. Reconcile partial
multi-repository commits/pushes from the delivery journal; do not reset completed
repositories to simulate an atomic operation.
PRs merge to develop under the project’s normal checks. Cross-repository
merges are not atomic: use backward-compatible changes, a documented merge
order and integration tests against the recorded combination of revisions.
A release workflow promotes a qualified develop revision to master through
a release PR. Task completion does not implicitly authorize a release merge.
Shared indexing and tools
Provide host-managed GitNexus and codebase-memory-mcp integrations to all
workspace-authorized agents. Their configuration and lifecycle belong to the
workspace service rather than a particular Codex or Claude conversation.
Repository instructions such as AGENTS.md remain visible in every worktree.
Index identity includes workspace, task worktree, repository, HEAD and content checkpoint/generation. A source checkout’s index must not silently answer as though it describes a task branch or its uncommitted edits. Index queries return freshness metadata. Review uses indexes matching its checkpoint or reports stale coverage and falls back to direct source inspection. Cross-repository queries operate over the task’s pinned repository membership and identify the contributing repository/revision for each result.
Shared indexing processes must not write into frozen source trees. Keep index storage and lock files outside reviewed worktrees where the tool permits it; otherwise index a disposable checkpoint copy. Serialize incompatible index updates. Qualification must establish how each tool handles worktrees, repository identity, uncommitted changes and cross-repository relationships. This design does not assume both tools already provide those capabilities or that their results are interchangeable.
Chat and execution contract
Keep turn type, repository input mode, and adapter routing separate:
- Turn type expresses the interaction supported by the agent’s published policy. A workspace does not itself grant an otherwise forbidden turn type.
- Input mode selects an immutable bundle or a registered task workspace.
- Adapter/authentication profile selects the coding harness and credentials;
a question about code must not imply an automatic switch to
llm-gateway.
For workspace execution, Chat selects a registered workspace and starts or resumes a task. It displays repositories, branches, implementation/review stage, active writer, indexing freshness, findings and related PRs. It does not ask for bundle URI/hash/size. Repository selection may help scope a prompt but is not an access control. The coding-only agent can support code-understanding jobs through its coding adapter with read-only execution authority; this requires a defined job contract rather than pretending every question is a patch implementation.
Extend the versioned request protocol with a discriminated input contract, such
as bundle versus taskWorkspace. Workspace jobs bind workspace/task identity,
expected membership/checkpoint, execution role and lease generation. Keep bundle
validation intact; do not relax its single-repository digest requirements to
smuggle in mutable directories. Unsupported runtimes reject the new contract.
Recovery and lifecycle
Task states include provisioning, implementing, review-frozen, remediating, ready-to-publish, publishing, PR-open, integrated, paused and archived. A failure records the interrupted stage and evidence rather than erasing working changes. On restart, reconcile the journal, actual worktrees, active processes, local Git refs and remote effect records before resuming. Never infer review approval from an agent’s completion message alone.
A paused task retains its worktrees and branch state. Cancellation terminates executions and retains recoverable changes until the retention policy permits cleanup. Cleanup requires no active leases and explicit checks for dirty files, unpushed commits and delivery status. Preserve referenced review/test evidence. Remove worktrees with Git’s worktree management and never recursively delete a path supplied by the client. Removing an agent grant stops new access and fences active access according to the workspace revocation policy.
Delivery phases and acceptance criteria
Phase 1: Persistent multi-repository tasks
Implement workspace registration, membership snapshots, per-task worktrees, versioned input contracts and runner lease/journal handling. Demonstrate two parallel tasks over the same three repositories with different task branches; changes and branch operations in one task must not modify the other. Existing bundle requests continue to pass their admission and staging tests.
Phase 2: Shared review and indexing
Implement checkpoints, enforced review freeze, handoff, disposable test copies and indexing integrations. A second authorized agent must see the implementer’s uncommitted changes in all repositories. Attempts to mutate a frozen workspace must fail. An edit after review must invalidate approval. Verify index freshness for different branches and untracked files, plus recovery from a lost writer.
Phase 3: GitHub delivery and release integration
Implement authorized Git/GitHub tools, idempotent delivery records and linked
PRs targeting develop. Exercise a partial push failure and retry without
lost commits or duplicate PRs. Exercise a changed remote branch and conflict
resolution requiring renewed review. Test the recorded multi-repository revision
set before release promotion to master.
Phase 4: Adapter interoperability
Qualify Codex implementation followed by Claude review and the reverse against the same workspace contracts. Adapter qualification, subscription authentication and tool access remain independent gates. Switching agents must preserve files, checkpoint identity and workflow state without reusing private model context.
Shared Task Workspace Runtime
The first implementation is an owner-local CLI and stdio MCP service named
light-workspace, backed by the task-workspace Rust crate. It provides workspace
registration, multi-repository task worktrees, shared file tools, review freezes,
reviewed commits, GitHub task-branch delivery and checkpoint-scoped indexing.
The operational guide is maintained with the executable:
Portal Chat can dispatch standalone inspect/implement tasks through the personal
native runner when a matching codingProfile.workspaceBindings policy is
published. The CLI/MCP registration alone does not enable that integration.
See Chat and Workflow integration
for the implemented path and the remaining workflow milestone.
For a guided walkthrough, see the shared task workspace tutorial.
Build and register
Run these commands in a Linux terminal as the user that will launch the agents.
The examples use Steve’s host paths. Prerequisites are Rust/Cargo, Git, jq,
bubblewrap (/usr/bin/bwrap) for execute, and authenticated gh for GitHub
operations. Git must be able to read each registered origin without interactive
credential prompts. Index providers are optional and must be configured before
registration if you intend to use them.
cd /home/steve/workspace/light-fabric
cargo build --release -p light-workspace
umask 077
# Replace HOST_ID with the Portal Host ID that owns this workspace.
# Discovery reads origins and local refs without changing source checkouts.
target/release/light-workspace discover /home/steve/workspace personal HOST_ID \
com.networknt.agent.codex-personal-1.0.0 \
com.networknt.agent.claude-personal-1.0.0 > discovery.json
# Set grants BEFORE the first registration. Do not redirect onto discovery.json.
jq '.workspace | .operations = ["edit", "execute", "review", "commit", "push", "issue", "pull-request"]' \
discovery.json > workspace.json
jq '{id, hostId, agents, operations, indexers, repositoryCount: (.repositories | length)}' workspace.json
If you already registered successfully, skip discovery and registration and go to Connect a local agent. Use the persisted registration when checking IDs and grants; changing the input file does not update it.
The result has workspace, integrationBranchNotKnownLocally, and
skippedRepositories with directory/reason diagnostics for unsuitable names or
unreadable origins. A skipped child does not stop discovery. Extract the
workspace object into a private registration file. Select its operation grants
explicitly; discovery grants none. Each agent grant covers all repositories.
Example (replace paths and identities):
{
"schemaVersion": 1,
"id": "personal",
"hostId": "HOST_ID",
"agents": ["com.networknt.agent.codex-personal-1.0.0", "com.networknt.agent.claude-personal-1.0.0"],
"operations": ["edit", "execute", "review", "commit", "push", "issue", "pull-request"],
"repositories": [
{
"name": "service",
"source": "[email protected]:OWNER/SERVICE.git",
"integrationBranch": "develop",
"releaseBranch": "master"
}
],
"indexers": {
"gitnexus": {
"executable": "/absolute/path/to/gitnexus",
"args": ["analyze", "--force"],
"timeoutSeconds": 300
},
"codebase-memory-mcp": {
"executable": "/absolute/path/to/codebase-memory-mcp",
"args": [],
"timeoutSeconds": 300
}
}
}
Register the extracted workspace.json, not the discovery wrapper. Keep the
persistent store outside the source tree:
install -d -m 700 /home/steve/.local/share/light-workspace
target/release/light-workspace \
/home/steve/.local/share/light-workspace register workspace.json
Expected output: {"registered":"personal"}. The private directory stores managed
repositories, task worktrees, checkpoints, and indexes. Both agents use this path.
Registration is metadata-only and idempotent for identical input. It rejects changes to an existing registration. The current implementation does not yet provide membership migration or live grant revocation. Do not edit registration files while tasks or clients are active. Credentials stay in host-managed Git and GitHub authentication, not this JSON or the model’s tool arguments.
Connect a local agent
Here, local agent means the Codex CLI or Claude Code application running as
steve on this Linux host. Open an ordinary terminal to run the commands below.
These steps do not configure the codex-personal service in Portal Chat or require
a Compose restart. The MCP client starts light-workspace as a child process and
exchanges JSON over stdin/stdout; there is no URL or port to enter, and you do not
start serve manually in another terminal.
Check the registered identities
jq '{id, agents, operations}' \
/home/steve/.local/share/light-workspace/personal/workspace.json
The examples below assume the full agent IDs produced by the discovery command.
Every launch identity must match an entry in agents exactly. The MCP server name
personal-workspace is only a client-side label; the workspace ID is personal.
Both clients use the same private store, with a different agent identity. This is
an owner-local launch setting, not network authentication or isolation between
host administrators. Do not run the clients as different Linux users against this
0700 store without designing an authenticated shared service first.
Add the server to Codex CLI
Run this once in your Linux terminal, using your existing Codex installation and
login. It updates your local Codex configuration, not workspace.json:
codex mcp add personal-workspace -- \
/home/steve/workspace/light-fabric/target/release/light-workspace \
/home/steve/.local/share/light-workspace serve personal \
com.networknt.agent.codex-personal-1.0.0
codex mcp list
The list should contain personal-workspace and the command above. Start a new
Codex CLI session by running codex; in its interactive prompt enter /mcp.
Check that personal-workspace connects and exposes task_workspace.
An existing session may need to be restarted to pick up the new configuration.
Codex stores this entry in ~/.codex/config.toml. Do not paste an mcpServers
JSON object into that TOML file. For long workspace operations you can add
tool_timeout_sec = 600 inside the existing [mcp_servers.personal-workspace]
table. That is a client timeout, not a guarantee that indexing all repositories
will finish in ten minutes. Prefer the direct CLI for the initial clone and large
index builds. See the official Codex MCP documentation.
Add the server to Claude Code (optional second agent)
In your Linux terminal, with Claude Code installed and signed in:
claude mcp add --transport stdio --scope user personal-workspace -- \
/home/steve/workspace/light-fabric/target/release/light-workspace \
/home/steve/.local/share/light-workspace serve personal \
com.networknt.agent.claude-personal-1.0.0
claude mcp get personal-workspace
Start a new session with claude, then use /mcp to inspect the connection.
--scope user keeps this setting in your user configuration instead of creating
a repository .mcp.json. See the Claude Code MCP documentation.
You can connect Codex first and add Claude later; two separate processes share
persistent tasks through the same store and task locks.
Create the first task outside the model session
Run this in your Linux terminal after checking Git access to the registered
origins. The first creation clones every registered repository and fetches its
integration branch. For a 129-repository workspace this can take time and disk
space; registration itself did not download these repositories. The managed copies
come from the recorded origins, so uncommitted edits in /home/steve/workspace
are not imported.
printf '%s\n' '{"operation":"create","task":"workspace-smoke-1"}' | \
/home/steve/workspace/light-fabric/target/release/light-workspace \
/home/steve/.local/share/light-workspace call personal \
com.networknt.agent.codex-personal-1.0.0
Expected output includes "state": "ready" and checkouts with one managed path
per repository, each on agent/workspace-smoke-1. Retrying the same task ID resumes
provisioning or returns the existing task. Do not invent a new task ID on every
retry. Missing develop branches must be resolved explicitly; there is no fallback
to master. Task creation makes no commits, pushes, issues, or PRs.
Ask the connected agent to use it
In the Codex conversation, enter:
Use the personal-workspace MCP server’s task_workspace tool. Call status for task workspace-smoke-1 and report its state and repository names. Use only the workspace tools for this task; do not access its files through host shell or native file tools. Do not edit, commit, push, or create GitHub resources.
The tool arguments for that first call are:
{"operation":"status","task":"workspace-smoke-1"}
Then ask it to read a specific file, for example:
For workspace-smoke-1, use task_workspace read to read README.md in repository light-fabric and summarize the project. Do not change files.
Use a repository name returned by status if light-fabric is not registered.
This confirms the model can use the tool rather than merely seeing its name.
The model client handles its own model login; the workspace server does not log
in to Codex/Claude on your behalf.
The manager must remain the only agent write path. Instructions to use MCP are
workflow guidance, not a security boundary: an agent with unrestricted host file
or shell access can bypass the freeze. This tutorial does not configure a hardened
native-agent sandbox. Workspace execute does sandbox the commands it launches.
Implement, review and deliver
createa task.filesreturns the repository contents and revision evidence;readreturns UTF-8 file content and its digest (up to 1 MiB).edittakes a nestededitobject:{operation: "edit", task, edit: {repository, path, content, expectedDigest}}. For a new file useexpectedDigest: null; for deletion usecontent: null`. Digest preconditions prevent overwriting another agent’s intervening changes.- For local commands,
executetakesaccess: "implement", an absolute executable under/usr, an argument array, andtimeoutSeconds(1–300). Linux bubblewrap provides an isolated environment with writable task files, read-only Git administration, no host home/credentials, and no network. This command is for inspection/build tools; it does not launch an authenticated native Codex or Claude model session. Network-dependent builds need a future explicitly authorized dependency mechanism. freezecaptures the cross-repository HEAD, staging state, file content, executable modes, deletions and untracked files. All further edits are rejected.- The second agent reads the same files or runs
executewithaccess: "review". Review execution mounts task files read-only. Build/test commands that need source-tree writes must use an appropriate disposable environment; a general writable review-test-copy command is not implemented yet. - Submit
reviewwith the exact checkpoint digest,approvedandfindings. Findings are limited to 64 KiB. An agent that contributed edits cannot approve its own task changes.remediateinvalidates approval and reopens editing. Any content or staging change makes review stale. commitaccepts a message after approval. It stages the reviewed files, rejects clean-filter transformations, disables Git hooks, and commits changed repositories. Configure the host’s committer identity beforehand. Repeating the same action reconciles already-created commits; it does not duplicate them.pushcreates remote task refs only if absent or already at the reviewed commit. An absence lease prevents a racing push from overwriting a newly created ref. Unknown remote changes require reconciliation. It never pushes integration or release refs.githubwithaction: "issue"creates an issue;action: "pull-request"creates a PR after a successful push. Supplyrepository,titleandbody; use the body to link existing issues and related repository PRs. PRs target the configured integration branch (developby default), and their head is checked against the reviewed commit. GitHub.com SSH/HTTPS origins are supported. The host needsghauthenticated for those repositories. No GitHub Enterprise URL adapter is supplied.
Task worktrees share a task-wide kernel lock. Separate tasks run concurrently.
Running commands journal their generation before launch. A timeout or lost process
leaves an interrupted task; no new writer is automatically admitted. A failure to
spawn the sandbox restores the previous state. After confirming that the old
execution and its descendants have terminated, an operator can recover using the
generation returned by status:
target/debug/light-workspace /absolute/private/store recover personal TASK_ID codex-personal --fenced-generation GENERATION
The command rejects stale generations and active task locks. The flag records the operator’s fencing assertion; it does not terminate processes itself. Unchanged read-only review executions retain their checkpoint and approval. Writer recovery invalidates them. Legacy interrupted records without execution context also invalidate them. There is no model-facing recovery override. Automatic crash reconciliation and task lifecycle retention still require runner integration.
GitHub effects persist intent before execution. On uncertain issue retries, the service enumerates paginated issue bodies through the GitHub API and matches the exact recorded marker, excluding pull requests. It does not rely on full-text search indexing of HTML comments. PR retries match the marker on the task branch. Missing or ambiguous matches remain uncertain and never cause duplicate creation. The current service creates issues/PRs but does not edit existing issue content, merge PRs, run release automation, or track required CI checks.
Indexing
index takes a configured provider. It requires a frozen checkpoint and makes
an independent clone per repository with the reviewed working-tree contents.
Generated AGENTS.md, .gitnexus, and .codebase-memory files stay in those
copies. Index caches and registry HOME are isolated per task/index generation.
Host-owned indexing commands are trusted administrative tools, not arbitrary
model-provided executable paths. Repeated indexing of the same checkpoint reuses
the active index. A replacement is published only after every repository succeeds;
a failed build preserves the previous receipt and removes the failed copy. Under
the task lock, the service reclaims obsolete generation directories, retaining the
active generation plus at most one build in progress. Provider processes inherit
the lock so restart cleanup cannot remove a still-running provider’s generation.
index-status reports the input checkpoint digest and freshness. Freshness means
that source identity matches; it does not prove complete parsing or complete
cross-repository relationships. Per-repository provider output/coverage is retained
in private *-output.json files under the returned index root.
index-query requires provider, repository and query. GitNexus receives a
concept query; codebase-memory-mcp receives a symbol-name regex through
search_graph. Queries reject stale input and use the correct task’s private
provider cache. Automatic cross-repository graph linking and provider-native
architecture/impact tools are not yet exposed by this facade.
The codebase-memory adapter uses the upstream one-shot CLI contract, with no installer or watcher activation. Qualified locally with GitNexus’s installed CLI and codebase-memory-mcp 0.10.8 on a small Python repository. Both indexed and found its function without changing the frozen source.
Current limits and qualification
This first implementation rejects checkpoint symlinks and submodules rather than following paths outside managed worktrees. File tools handle UTF-8 text; binary changes can be inspected through file digests and made by sandboxed commands. The manager must be the sole agent write path. Giving a native agent unrestricted host shell access alongside these tools would bypass the review freeze.
Run:
cargo test -p task-workspace -p light-workspace
# Explicit host qualification, including the normally ignored namespace tests:
cargo test -p task-workspace --test workspaces -- --include-ignored
cargo clippy -p task-workspace -p light-workspace --all-targets -- -D warnings
Tests use disposable repositories and a mock GitHub CLI. They cover three-repository parallel tasks, shared edits, digest preconditions, stale/self-review rejection, read-only review, host-file isolation, timeout recovery, index copies/freshness, reviewed multi-repository commits, local-remote pushes, PR base/head verification, uncertain-effect retry, and the stdio MCP transport. They do not perform billable model calls or create real GitHub issues/PRs.
Setup troubleshooting
| Symptom | Check or next step |
|---|---|
No such file or directory from registration | Run from the directory containing workspace.json, or give its absolute path. Extract .workspace from discovery first. |
| Shell says the binary does not exist | Build with --release and use target/release/light-workspace consistently. |
workspace already registered with different membership or grants | Registrations are immutable. Before any tasks or active clients exist, back up the unused workspace directory and register the corrected input. With existing tasks, retain the registration and use a new workspace ID/store until migration is implemented. Never overwrite stored metadata to bypass the guard. |
agent has no workspace grant | Match the full agent ID in the stored agents array; a shortened name is a different identity. |
serve appears to hang in a terminal | It is waiting for MCP JSON. Let the client launch it; use call for a direct JSON operation. |
/mcp does not show the server | Check the client registration and absolute binary path, then start a new client session. |
| Tool timeout during first task creation | Use the terminal call command above and retry the same task ID; check Git authentication and integration branches. |
index provider is not configured by host | Discovery leaves indexers empty. Add provider configuration before registration; merely installing an indexer does not register it. Existing registrations have no update command. |
| Review command cannot write | Review mounts source files read-only. Writable build/test copies are not yet implemented. |
| Portal Chat still requests a repository bundle | Expected: this local MCP setup is not wired into Portal Chat yet. |
Shared Task Workspaces: Chat and Workflow Integration
Status: shared Codex/Claude worker and Chat extension, September 12, 2026. See Shared native coding sessions for the current session contract and deployment prerequisites. Higher-level development workflow orchestration below remains a design; its cross-database job bridge is not yet qualified in the local stack.
This document extends Shared Task Workspaces and Development Workflow Orchestration.
Implemented standalone Chat path
- Portal publishes
codingProfile.workspaceBindingswith Host, environment, runner, membership revision, authorization revision, human subjects, agents and allowed intents. Java/Rust fixtures verify signed serialization. - An authenticated coding Chat session advertises only its allowed workspaces. The browser sends a workspace ID, new/existing task selection, intent, instruction and optional expected checkpoint. Host paths and authority are not accepted from the browser.
- The Agent schedules through the existing durable execution outbox and Controller transport, pinned to the bound personal runner. The runner checks its private workspace configuration and the local registration again.
- Native Codex and Claude use isolated per-conversation native state with only their corresponding personal login. Host repositories, configuration, plugins and unrelated MCP servers are not mounted. Its only workspace tool is a scoped file facade: repository listing, paginated file listing, reads and digest-conditional edits.
- A task lock spans the model turn. Unfinished writers are fenced. A durable
execution receipt prevents an uncertain model turn from being automatically
replayed; completed retries return their saved result. The job lock covers setup,
but uncertainty is persisted only immediately before
turn/start. Setup and authentication failures remain retryable. Existing tasks reopen from their pinned local revisions; only new tasks fetch integration branches. - Chat displays the final explanation, task ID and checkpoint, and selects that existing task for follow-up. Result notifications are deduplicated.
The adapters support inspect, implement, and read-only review on personal native runners. Review is tied to an exact workspace checkpoint. Running commands/tests, indexing, trusted approval, and GitHub delivery remain separate manager operations. Enterprise workspace adapters and workflow stages require separate qualification. Registering an MCP server alone does not enable Chat: publish a matching binding and deploy the workspace-enabled runner.
The native adapter has passed live read and edit tests across three disposable
repositories, including verification of the edited file and durable checkpoint.
The browser-to-Controller-to-runner gate passed on September 11 against the local
personal workspace and existing workspace-smoke-1 task (128 repositories).
Chat returned both README titles, then completed a digest-conditional file edit
and returned a new checkpoint. The file was verified in the task worktree, with
the task ready and the original checkout untouched. The deployment record is in
portal-config-loc/all-in-lt/light-workflow-runner-personal/workspace-chat.md.
First-time provisioning of all 128 repositories from Chat remains unqualified;
use an existing task for this milestone. Workflow orchestration remains next.
Decision and user experience
Support both interactive Chat jobs and durable development workflows through one workspace execution contract. Chat is a user interface, not the owner of a worker process or write lease. The runner resolves registered workspace IDs to host paths; no browser or model supplies a filesystem path as authorization.
An agent grant covers every repository in the workspace. Repository names in a prompt or task plan express intent, not additional access controls. Task branches belong to a task, not to a particular agent. Codex and Claude use the same task identity to inspect the same staged, unstaged, and untracked changes.
The normal interaction in /app/genai/chat is:
- Connect to an agent. Retain its published turn policy: a coding-only agent remains coding-only, including for code-understanding questions.
- For coding turns, choose Repository bundle or Workspace, provided the selected agent and runner advertise support. Existing bundle behavior remains.
- Select a registered workspace such as
personal. Select an existing task or choose New task with a short description. Display its repository count, integration branch policy, and provisioning readiness. - Choose Inspect code, Implement, or Implement and review. These are job intents, separate from Chat/coding turn types and provider routing.
- Submit the instruction. The backend creates a stable task identity and starts provisioning automatically when needed. Show progress and allow reconnection.
- Show task stage, current execution, checkpoint, findings, and delivery links. A later conversation or another authorized agent can resume the same task.
Workspace mode hides bundle URI/hash/size inputs. Neither changing the selected agent nor refreshing the browser silently creates a replacement task. A new-task choice remains available when context preselects an existing task. Agents without workspace support show the reason that this mode is unavailable.
Existing implementation and proposed responsibilities
| Component | Current responsibility | Integration work |
|---|---|---|
portal-view/src/pages/genai/Chat.tsx | Agent session and coding input UI | Workspace/task selection, job intent, progress and approval views |
light-agent | Session, turn authorization and coding dispatch | Admit versioned workspace inputs; retain published turn policy |
| Portal publisher/config server | Published agent configuration | Workspace catalog/grants, runner binding and revision publication |
light-workflow | Durable workflow execution | Development stages, waits, retries, budgets and user decisions |
| Controller and runner | Authenticated execution dispatch | Workspace-aware job admission, host routing, fencing and reconciliation |
crates/task-workspace / apps/light-workspace | Local tasks, locks, file tools, checkpoints, indexing and GitHub receipts | Runner-owned scoped tool facade and registration adoption |
| Coding adapters | Bounded provider-specific execution | Scoped workspace tools and inspection/implementation/review qualification |
No browser-to-host HTTP wrapper is added around the current owner-local MCP process. Reuse the authenticated execution transport. The host runner launches or embeds the workspace manager locally; any future network service needs a separate authenticated protocol and is outside this first deployment.
flowchart TD
Chat[GenAI Chat] --> Admit[Authenticated job admission]
Workflow[Development workflow] --> Admit
Chat -->|Start or resume development process| Workflow
Admit --> Controller[Controller execution dispatch]
Controller --> Runner[Host runner]
Runner --> Manager[Workspace manager]
Runner --> Adapter[Selected coding adapter]
Adapter -->|Scoped tools| Manager
Manager --> Trees[Task worktrees and checkpoint copies]
Manager --> Effects[Trusted Git and GitHub actions]
Runner --> Events[Durable results and progress]
Events --> Chat
Events --> Workflow
Durable job and workflow ownership
Interactive inspection and implementation are bounded jobs with durable execution records. They do not require a full development workflow. A workflow owns the multi-stage lifecycle once Implement and review is selected or a workflow is started directly. Both paths use the same task manager and admission rules.
A task has at most one active workflow owner. While owned, independent Chat mutations become commands to that workflow (pause, revise instruction, resume), not a second execution path that races its stages. Read-only status is always subject to access checks. Independent inspection may run only against a stable checkpoint; requests during active writes wait or report that the task is busy.
The workflow record owns stage and delivery intent. The workspace journal owns actual local files, execution generation and checkpoint state. Neither can invent progress for the other: a workflow records a completed stage only after consuming a matching durable runner receipt. Reconciliation compares both after restart.
Proposed request and capability contracts
Extend the versioned coding envelope with a discriminated repository input. Keep the existing bundle schema intact. Illustrative workspace input:
{
"schemaVersion": 2,
"requestId": "opaque-idempotency-key",
"intent": "inspect",
"repositoryInput": {
"kind": "taskWorkspace",
"workspaceId": "personal",
"taskId": "opaque-task-id",
"expectedMembershipRevision": "opaque-revision",
"expectedCheckpointDigest": null
},
"instruction": "Explain how the config server publishes agent policy"
}
These field names are proposals. Final schemas require cross-language fixtures and version negotiation before publication. New-task creation is a separate admission operation: allocate the task ID once, persist it against the request ID, then enqueue provisioning. Do not let a retry allocate a different branch name.
The server derives Host, user, agent, environment, runner binding, operation grants, policy digest, deadline and execution identity from authenticated context and published policy. Runner-issued execution generations are never accepted as client-created fencing evidence. Checkpoint and membership expectations supplied by the client are concurrency preconditions, not authorization.
Capabilities must distinguish supported input modes and intents. Admission takes the intersection of agent policy, workspace grants, runner capabilities and adapter qualification. Capability advertisement alone never grants an operation. An old runtime rejects the new input kind before execution rather than treating it as a bundle or silently switching to an LLM gateway path.
The workspace catalog exposes display names, repository summaries, readiness and permitted actions to authorized users. It does not expose host credentials or unnecessary absolute paths. A human user’s permission and the selected agent’s grant must both authorize the workspace; possession of an agent ID is insufficient.
Registration, publication and preflight
Use config-server publication for the workspace identity, Host/environment, repository catalog, operation/agent grants, runner binding, and revision. Keep the physical store path, provider executables and credential handles in host-owned runner configuration. Do not put private keys or GitHub tokens in Chat inputs, event payloads, or model context.
Adopt an existing local registration only after comparing its canonical catalog and grants with the published revision. The initial integration must support the existing private store without deleting worktrees or rewriting task evidence. Local immutable registrations are not silently made mutable by adding a publisher.
Separate repository membership revision from authorization revision in the new contract. Tasks pin membership. Current grants are rechecked on every dispatch and tool operation; revocation denies new access and fences affected active executions. Changes to membership require explicit migration and new review evidence. The current local digest covers more than repository membership, so schema migration must be explicit and legacy tasks must retain their original digest semantics.
Before declaring a workspace ready:
- Verify remote access and the integration branch of each repository; report each missing branch, archived/unwritable remote, or credential failure by repository.
- Fetch the required integration ref before checking its commit in an existing managed bare clone. A newly created remote branch must repair stale local refs.
- Validate store ownership, space, Git identity for delivery, required tools and configured indexers. Index readiness is separate from basic file-tool readiness.
- Record per-repository provisioning progress. Resume completed clones on retry; inspect interrupted clones and repair only the task’s own stale registrations.
Do not create missing remote branches automatically during a coding turn. That is an explicit administrative action. Error messages include safe repository identity, operation and branch, while excluding credential-bearing URLs and raw Git stderr. Large-workspace provisioning is asynchronous, with a durable operation ID, rather than a browser request held open until all repositories finish.
Runner and adapter enforcement
The runner owns the task lease and process lifecycle. The adapter receives a
short-lived tool binding scoped to task, execution, role, generation and policy
revision. Tool calls validate that binding and task state. Do not expose the
current unrestricted local call/serve identity selection to a model.
| Intent | Allowed task access | Required outcome |
|---|---|---|
| Inspect | Stable read-only checkpoint; no approval authority | Answer plus checkpoint identity and referenced files |
| Implement | Exclusive write access on a writable task | Recorded changes and test results; explicit freeze for handoff |
| Review | Read-only frozen checkpoint; independent reviewer | Structured findings and disposition tied to exact digest |
| Deliver | Trusted fixed actions, separately authorized | Per-repository effect receipts and reconciliation state |
The local runtime currently supports read-only execution only in Frozen/Approved states. Add an inspection snapshot/lease contract rather than disguising inspection as implementation or recording an approval for a code-understanding question. An inspection snapshot must not reopen a committed task or invalidate approval when unchanged. Its branch/content/index identity must be recorded.
Native adapters must not also receive unrestricted host shell or file access to the task store. Prompts saying “use MCP only” are not enforcement. Qualify a restricted tool set or an OS boundary that confines native file access and keeps Git administration and host credentials inaccessible. If an adapter cannot meet that boundary, do not advertise workspace mutation/review support for it.
Model network access belongs to the adapter’s narrowly authorized provider path. Workspace commands remain offline by default. Dependency downloads require a separate policy-controlled mechanism. Review tests needing writes run in disposable checkpoint copies with separate outputs; source review worktrees remain read-only. These writable test copies are new work, not a current local-runtime capability.
Development workflow
The first template implements the following sequence with explicit stage receipts:
Admit request -> provision/resume task -> implement -> freeze checkpoint
-> optional checkpoint indexing/tests -> independent review
-> if changes requested: remediate -> implement -> freeze -> review
-> delivery authorization -> commit -> push -> open related PRs to develop
-> PR-open outcome
Inputs identify workspace/task, requirement or issue references, implementer and reviewer agent deployments, budgets, retry limits and delivery policy. An issue reference includes repository and issue number. All role agents need workspace grants. Enforce a different non-contributing reviewer; do not trust a display name or reuse hidden implementer context as the reviewer’s context.
Handoff includes requirements, accepted plan, checkpoint, changed-file summary, test evidence and earlier findings. Agent completion text cannot approve a task. Any change after freeze invalidates approval and requires another review. Bound remediation rounds and total cost/time; exhausting a budget pauses for a decision.
Delivery follows the published approval policy. A human approval, when required, shows the checkpoint, per-repository change/commit plan, destination branches and intended GitHub effects. Bind the decision to that exact intent and digest. A retry with identical evidence does not demand another approval; changed evidence does.
Multi-repository delivery is not atomic. Persist each commit/push/PR receipt and
show partial completion. Reconcile uncertain writes before retrying; never force
push or duplicate a GitHub resource to hide partial failure. Stop on conflicting
remote refs and require explicit remediation and renewed review. PRs target
develop; integration testing, merge and release promotion to master are separate
workflow stages/extensions, not automatic consequences of opening PRs.
Progress, cancellation and recovery
Return accepted operation IDs promptly. Persist progress with monotonic sequence numbers and task/execution correlation; stream it through the existing session transport. On reconnect, retrieve a snapshot and events after the last sequence. Browser disconnect does not cancel work or release a write lease.
Duplicate submission with the same request ID returns the same task/execution. Reusing that ID with different input is a conflict. Stage retries carry stable idempotency keys and expected checkpoint/generation. Task state changes use compare-and-swap semantics; concurrent Chat sessions cannot both acquire a writer.
Cancellation requests termination, waits for process-tree fencing, records the last durable state, and retains files. A timed-out writer remains interrupted until fencing is confirmed. Do not expose operator recovery as a model tool. Read-only recovery preserves approval only after recapturing and matching the checkpoint. Lease expiration alone is not proof that an old process cannot write.
Recovery must work while the workflow, controller or runner restarts independently. If the host is offline, show pending/unavailable state rather than scheduling the same mutable task on an unrelated host. Cleanup requires no active execution and an explicit retention decision that accounts for dirty files and unpushed commits.
Local deployment
For Steve’s setup, source discovery uses /home/steve/workspace; the persistent
store is /home/steve/.local/share/light-workspace. Run the personal execution
runner and its workspace manager on that host as the authorized store owner.
Do not mount the entire host home into the Portal agent containers.
portal-config-loc/all-in-lt should supply host-runner preparation/startup scripts,
private runtime configuration templates, workspace publication/adoption checks,
and a deployment smoke test. Compose continues to host Portal, controller,
workflow and agent services; configure their authenticated runner connectivity
using the established runner transport. A standalone stdio container exposed on
a new port is not the integration. Mirror supported installation assets in
light-portal-install after the local path is qualified.
No global runner enablement without admission, execution bindings/database and credentials. Configuration examples must identify which values come from config server and which are host secrets. Preserve bootstrap-only agent configuration.
Implementation phases and acceptance gates
| Phase | Deliverable | Required evidence |
|---|---|---|
| 1. Contracts and registration | Versioned inputs/capabilities, published workspace grants, existing-store adoption and branch preflight | Java/Rust/TS fixtures agree; old runtimes reject unsupported modes; grants/Host mismatches rejected; existing local tasks preserved |
| 2. Runner integration | Async provisioning, scoped tools, job records, inspection snapshots, enforced adapter access | Three repositories, two concurrent tasks, same-task writer conflict; native bypass attempts rejected; cancellation/restart fencing |
| 3. Chat | Workspace/task selectors, inspect/implement jobs, progress and reconnect | Browser creates/resumes a task without CLI or bundle; duplicate send creates one task; coding-only inspection stays on coding adapter; bundle and Tech Support regression tests |
| 4. Workflow handoff | Durable implement/freeze/review/remediate stages, budgets and approvals | Codex edits and Claude sees exact uncommitted changes; self-review/stale approval rejected; restart each stage; bounded remediation loop |
| 5. Delivery and deployment | Trusted effects, partial delivery reconciliation, local startup assets and tutorial | Local-remote/mock GitHub gates plus explicitly authorized live qualification; no duplicate PR; fresh deployment and restart pass end to end |
Track implementation across light-fabric, light-portal, portal-view,
portal-config-loc, light-portal-install, light-portal-doc, and
light-portal-test. Backend/API contracts precede UI wiring; deployment examples
must use a qualified binary and compatible published revisions. A phase is not
complete merely because a local CLI test passes.
The completion gate is a browser-driven task on the local deployment: create or
resume personal, implement across multiple repositories, independently review
the shared changes, approve delivery under policy, and open PRs to develop.
Repeat via a directly started workflow and prove both entry points produce the
same authorization, checkpoint and effect receipts. Document any remaining
adapter, indexing, dependency-build or release limitations explicitly.
Deploy Native
This page describes the recommended VM deployment model for the Rust
light-agent native binary.
Use this model when a customer wants to run an agent service on a VM and expose
the chat UI/WebSocket endpoint outside Kubernetes. The agent serves the local
chat UI, connects to an LLM provider, calls MCP tools through light-gateway,
stores conversation memory in Postgres, and registers with controller.
Recommended Model
Deliver a versioned install bundle, not an ad hoc runtime script.
The bundle should contain:
light-agentnative binary.public/static assets for the chat UI.- Minimal bootstrap config files.
- A
systemdunit. - An install script for filesystem setup.
- A root-owned environment file for secrets.
Use systemd to run the service:
- It restarts the process on failure.
- It keeps logs in the host journal.
- It avoids shell-history and process-list leakage from command-line secrets.
- It gives the customer a standard operational surface:
start,stop,restart,status, andjournalctl.
Do not use a long-running shell wrapper to pass the bootstrap token, database URL, or model configuration. Use config files and an environment file instead.
Runtime Layout
light-agent uses relative runtime paths:
configpublic
The systemd service should therefore set WorkingDirectory to the installed
application directory.
Recommended VM layout:
/opt/light-agent/
light-agent -> releases/2.2.1/light-agent
releases/
2.2.1/
light-agent
config -> /etc/light-agent
public/
index.html
/etc/light-agent/
startup.yml
server.yml
portal-registry.yml
client.yml
mcp-client.yml
ollama.yml
values.yml
ca.pem
light-agent.env
/var/lib/light-agent/
config-cache/
The local config directory contains bootstrap and agent-specific config.
Runtime config downloaded from config-server should be written to
/var/lib/light-agent/config-cache by setting externalConfigDir in
startup.yml.
Keep /etc/light-agent readable by the service user. Keep
/var/lib/light-agent/config-cache writable by the service user.
Build Artifact
Build a release binary from light-fabric:
cargo build --release -p light-agent
The artifact is:
target/release/light-agent
For a static Linux build that matches the Docker build target:
rustup target add x86_64-unknown-linux-musl
cargo build --release -p light-agent --target x86_64-unknown-linux-musl
The static artifact is:
target/x86_64-unknown-linux-musl/release/light-agent
Build on a compatible Linux distribution for the customer VM. If the customer
fleet has mixed Linux versions, prefer a static or target-compatible build so
the binary does not fail on an older glibc.
Package with a versioned filename:
light-agent-<version>-linux-amd64.tar.gz
Include the static assets from:
apps/light-agent/public/
Runtime Dependencies
The VM must be able to reach:
- Controller, through
portalRegistry.portalUrl. - Config-server, through
startup.configServerUri. light-gateway, throughmcp-client.gatewayUrlandmcp-client.path.- The model provider, currently Ollama by default.
- Postgres, through
DATABASE_URL.
The Postgres database must contain the Hindsight memory tables used by
light-agent, including:
agent_memory_bank_tagent_memory_unit_tagent_session_history_t
LIGHT_AGENT_HOST_ID must be a valid host UUID for the target tenant/host. The
agent stores memory and session history under this host id.
Agent Roles
The same binary can run different logical agents. Use a different service id,
port, install directory, and systemd unit for each concurrently running role.
Common service ids are:
com.networknt.agent.account-1.0.0
com.networknt.agent.advisor-1.0.0
com.networknt.agent.tech-support-1.0.0
For a single account agent, keep the service name light-agent. For multiple
agents on the same VM, use names such as:
light-agent-account
light-agent-advisor
light-agent-tech-support
Each role needs a unique listener port if they run on the same VM.
Bootstrap Config
The local bootstrap config needs enough information to reach config-server,
controller, light-gateway, Ollama, and Postgres.
Example values.yml for an account agent:
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
startup.externalConfigDir: /var/lib/light-agent/config-cache
light-config-server-uri: https://config-server.customer.example.com:8435
server.serviceId: com.networknt.agent.account-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: agent-account-01.customer.example.com
server.httpPort: 8083
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.customer.example.com:8438
client.verifyHostname: true
mcp-client.gatewayUrl: https://mcp-gateway.customer.example.com
mcp-client.path: /mcp
mcp-client.timeoutMs: 5000
ollama.ollamaUrl: http://ollama.customer.example.com:11434
ollama.model: llama3.1:8b
server.advertisedAddress must be a stable address that controller and clients
can use to reach the VM agent. Do not advertise 127.0.0.1 or 0.0.0.0.
Example startup.yml:
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
externalConfigDir: ${startup.externalConfigDir:/var/lib/light-agent/config-cache}
Example server.yml:
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8083}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
Example portal-registry.yml:
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
Example client.yml:
tls:
verifyHostname: ${client.verifyHostname:true}
Example mcp-client.yml:
gatewayUrl: ${mcp-client.gatewayUrl:https://mcp-gateway.customer.example.com}
path: ${mcp-client.path:/mcp}
timeoutMs: ${mcp-client.timeoutMs:5000}
Example ollama.yml:
ollamaUrl: ${ollama.ollamaUrl:http://localhost:11434}
model: ${ollama.model:llama3.1:8b}
For the current light-agent implementation, keep ollama.yml and
mcp-client.yml in the local bootstrap config. They are read during process
initialization before the runtime completes remote config bootstrap.
Secrets
Keep secrets in a root-owned environment file or in the customer’s secret manager. Do not pass secrets in command-line arguments.
Example /etc/light-agent/light-agent.env:
LIGHT_PORTAL_AUTHORIZATION=Bearer <token>
light_4j_config_password=<config-password-if-needed>
LIGHT_AGENT_HOST_ID=<host-uuid>
DATABASE_URL=postgres://agent_user:<password>@postgres.customer.example.com:5432/configserver
RUST_LOG=info
AGENT_LOG_ANSI=false
Permissions:
chown root:light-agent /etc/light-agent/light-agent.env
chmod 0640 /etc/light-agent/light-agent.env
LIGHT_PORTAL_AUTHORIZATION is used for config-server bootstrap and controller
registration. It is not the end-user chat token. If downstream MCP tools require
caller identity, the browser or BFF should send the user’s Authorization
header to the agent WebSocket endpoint so the agent can forward it to
light-gateway.
Systemd Unit
Example /etc/systemd/system/light-agent.service:
[Unit]
Description=Light Agent
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=light-agent
Group=light-agent
WorkingDirectory=/opt/light-agent
EnvironmentFile=/etc/light-agent/light-agent.env
ExecStart=/opt/light-agent/light-agent
Restart=on-failure
RestartSec=5
LimitNOFILE=65535
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=full
ProtectHome=true
ReadWritePaths=/var/lib/light-agent/config-cache
[Install]
WantedBy=multi-user.target
Install and start:
systemctl daemon-reload
systemctl enable light-agent
systemctl start light-agent
systemctl status light-agent
View logs:
journalctl -u light-agent -f
Install Script Scope
An install script is useful, but keep it deterministic and small.
It should:
- Create the
light-agentuser and group. - Create
/opt/light-agent,/etc/light-agent, and/var/lib/light-agent/config-cache. - Install the binary with executable permissions.
- Install the
public/static assets. - Install bootstrap config files.
- Install or update the
systemdunit. - Set file ownership and permissions.
- Print the next operator steps for adding secrets and starting the service.
It should not:
- Embed bearer tokens.
- Pass tokens to
ExecStart. - Rewrite customer config-server state.
- Start the process before secrets, CA files, and database access are ready.
Startup Flow
The expected runtime flow is:
systemd
-> /opt/light-agent/light-agent
-> read local config/values.yml, ollama.yml, and mcp-client.yml
-> connect to Postgres with DATABASE_URL
-> build the MCP client for light-gateway
-> call config-server with LIGHT_PORTAL_AUTHORIZATION
-> write downloaded runtime config into /var/lib/light-agent/config-cache
-> start the Axum HTTP/WebSocket server
-> register the agent with controller using portalRegistry.portalUrl
-> serve the chat UI from public/
-> forward tool discovery and tool calls to light-gateway
When startup.yml configures config-server, the runtime tries to download the
latest values.yml before starting. If that download fails for any reason, the
runtime continues startup with the available local and cached config, including
config-cache/values.yml when present.
Endpoints
The native service exposes:
GET /health
GET /
GET /chat
/chat upgrades to WebSocket. The static chat UI is served from public/.
For local testing on the VM:
curl -i http://127.0.0.1:8083/health
Upgrade And Rollback
Use versioned binary releases:
/opt/light-agent/releases/2.2.1/light-agent
/opt/light-agent/releases/2.2.2/light-agent
/opt/light-agent/light-agent -> releases/2.2.2/light-agent
Upgrade:
systemctl stop light-agent
ln -sfn /opt/light-agent/releases/2.2.2/light-agent /opt/light-agent/light-agent
systemctl start light-agent
Rollback:
systemctl stop light-agent
ln -sfn /opt/light-agent/releases/2.2.1/light-agent /opt/light-agent/light-agent
systemctl start light-agent
Do not delete config-cache during a normal binary rollback. It is the local
cache of the config-server-delivered runtime state.
Validation Checklist
Before handing the VM to the customer:
systemctl status light-agentis active.journalctl -u light-agentshows successful config-server bootstrap.journalctl -u light-agentshows successful controller registration.- The controller shows the agent registered with the expected service id, environment, address, and port.
curl http://127.0.0.1:8083/healthreturns200 OK.- The chat UI loads from the VM address.
- The chat WebSocket connects to
/chat. - Logs show that the agent can connect to Postgres.
- Logs do not show MCP
tools/listfailures fromlight-gateway. - A chat request can discover and call a tool through
light-gateway. - Restarting the VM starts the agent automatically.
Security Checklist
- Store bearer tokens, config passwords, and database passwords outside the install bundle.
- Use a customer CA file instead of disabling TLS verification in production.
- Use a stable DNS name for
server.advertisedAddress. - Restrict inbound VM firewall rules to the required agent port.
- Restrict outbound VM firewall rules to config-server, controller,
light-gateway, Ollama, and Postgres. - Run as the dedicated
light-agentuser. - Keep
/etc/light-agent/light-agent.envreadable only by root and the service group. - Keep
/etc/light-agentwritable only by administrators. - Keep only
/var/lib/light-agent/config-cachewritable by the service. - Rotate
LIGHT_PORTAL_AUTHORIZATIONthrough the customer secret process.
Deploy Kubernetes
This page describes the recommended Kubernetes deployment model for the Rust
light-agent image from light-fabric/apps/light-agent.
Use this model when an agent service runs in a cluster and exposes the chat
UI/WebSocket endpoint through a Kubernetes Service, Ingress, or Gateway API. The
agent serves the local chat UI, connects to an LLM provider, calls MCP tools
through light-gateway, stores conversation memory in Postgres, and registers
with controller.
Recommended Model
Deploy the agent as a normal single-container Kubernetes workload:
Deploymentfor the agent pod.Servicefor stable in-cluster access.ConfigMapfor bootstrap config and non-secret values.Secretfor bearer tokens, config passwords, host id, and database URL.emptyDirorPersistentVolumeClaimforconfig-cache.ConfigMapor custom image layer forpublic/chat UI assets.- Optional
Ingress,Gateway API,NodePort, orLoadBalancerfor external browser access.
Keep runtime policy and shared platform configuration in config-server. The
Kubernetes bootstrap config should only contain enough information for startup,
trust, model/provider selection, light-gateway access, database access, and
controller registration.
Image
Build the image from the workspace root:
./apps/light-agent/build.sh 2.2.1
For local testing without pushing:
./apps/light-agent/build.sh 2.2.1 --local
Use immutable tags in Kubernetes. Avoid latest for customer deployments.
The current runtime image uses:
/app/light-agent
/app/config -> /config
The process runs as the image user agent. Mount /config for bootstrap
config and make /app/config-cache writable.
The current Dockerfile does not copy apps/light-agent/public/ into the runtime
image. For Kubernetes, either mount the public/ files from a ConfigMap or
build a custom image that includes them under /app/public.
Runtime Paths
Recommended container layout:
/config/
startup.yml
server.yml
portal-registry.yml
client.yml
mcp-client.yml
ollama.yml
values.yml
ca.pem
/app/config-cache/
values.yml
downloaded certs and files
/app/public/
index.html
Use a read-only projected volume for /config. Use a writable volume for
/app/config-cache.
For most deployments, use emptyDir for config-cache. This gives each pod a
fresh cache and avoids accidentally keeping stale config across pod replacement.
Use a PersistentVolumeClaim only when the customer explicitly wants the agent
to restart from the last downloaded config during a config-server outage. A
persistent cache improves outage tolerance but can also preserve stale runtime
state.
Runtime Dependencies
The pod must be able to reach:
- Controller, through
portalRegistry.portalUrl. - Config-server, through
startup.configServerUri. light-gateway, throughmcp-client.gatewayUrlandmcp-client.path.- The model provider, currently Ollama by default.
- Postgres, through
DATABASE_URL.
The Postgres database must contain the Hindsight memory tables used by
light-agent, including:
agent_memory_bank_tagent_memory_unit_tagent_session_history_t
LIGHT_AGENT_HOST_ID must be a valid host UUID for the target tenant/host. The
agent stores memory and session history under this host id.
Agent Roles
The same image can run different logical agents. Use a different service id, deployment name, Service name, and port for each concurrently running role.
Common service ids are:
com.networknt.agent.account-1.0.0
com.networknt.agent.advisor-1.0.0
com.networknt.agent.tech-support-1.0.0
For a single account agent, a conventional Kubernetes name is
light-agent-account. For multiple agents in the same namespace, use names such
as:
light-agent-account
light-agent-advisor
light-agent-tech-support
Each role needs a unique Service name. If they share one namespace and expose through one Ingress host, route each role by host or path.
Registration Address
In Kubernetes, do not register the pod IP. Pod IPs are ephemeral.
If controller and callers are inside the same cluster, advertise the Service DNS name:
server.advertisedAddress: light-agent-account.light-agent
The pattern is:
<service-name>.<namespace>
The port is still registered separately from the host/address.
If controller or callers are outside the cluster, advertise the externally reachable DNS name instead, such as the Ingress or LoadBalancer hostname:
server.advertisedAddress: account-agent.customer.example.com
Bootstrap Config
Example values.yml for an in-cluster controller, config-server, gateway,
Ollama, and Postgres:
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
startup.externalConfigDir: /app/config-cache
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.networknt.agent.account-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: light-agent-account.light-agent
server.httpPort: 8083
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
client.verifyHostname: true
mcp-client.gatewayUrl: https://ai-microgateway.light-gateway:8443
mcp-client.path: /mcp
mcp-client.timeoutMs: 5000
ollama.ollamaUrl: http://ollama.ai.svc.cluster.local:11434
ollama.model: llama3.1:8b
Example startup.yml:
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
externalConfigDir: ${startup.externalConfigDir:/app/config-cache}
Example server.yml:
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8083}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
Example portal-registry.yml:
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
Example client.yml:
tls:
verifyHostname: ${client.verifyHostname:true}
Example mcp-client.yml:
gatewayUrl: ${mcp-client.gatewayUrl:https://ai-microgateway.light-gateway:8443}
path: ${mcp-client.path:/mcp}
timeoutMs: ${mcp-client.timeoutMs:5000}
Example ollama.yml:
ollamaUrl: ${ollama.ollamaUrl:http://ollama.ai.svc.cluster.local:11434}
model: ${ollama.model:llama3.1:8b}
For the current light-agent implementation, keep ollama.yml and
mcp-client.yml in the local bootstrap config. They are read during process
initialization before the runtime completes remote config bootstrap.
Use the customer CA in ca.pem. Do not disable hostname verification in
production to work around certificate SAN problems.
Secrets
Store the portal bearer token, optional config password, host id, and database
URL in a Kubernetes Secret.
Example:
apiVersion: v1
kind: Secret
metadata:
name: light-agent-account-secret
namespace: light-agent
type: Opaque
stringData:
LIGHT_PORTAL_AUTHORIZATION: "Bearer <token>"
light_4j_config_password: "<config-password-if-needed>"
LIGHT_AGENT_HOST_ID: "<host-uuid>"
DATABASE_URL: "postgres://agent_user:<password>@postgres.lightapi.svc.cluster.local:5432/configserver"
data:
ca.pem: <base64-ca-pem>
LIGHT_PORTAL_AUTHORIZATION is used for config-server bootstrap and controller
registration. It is not the end-user chat token. If downstream MCP tools require
caller identity, the browser or BFF should send the user’s Authorization
header to the agent WebSocket endpoint so the agent can forward it to
light-gateway.
Do not store real bearer tokens, database passwords, or customer CA material in Git, ConfigMaps, Helm values committed to the repo, or rendered deployment examples.
Example Manifests
Example ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
name: light-agent-account-config
namespace: light-agent
labels:
app.kubernetes.io/name: light-agent-account
app.kubernetes.io/component: agent
data:
values.yml: |
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
startup.externalConfigDir: /app/config-cache
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.networknt.agent.account-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: light-agent-account.light-agent
server.httpPort: 8083
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
client.verifyHostname: true
mcp-client.gatewayUrl: https://ai-microgateway.light-gateway:8443
mcp-client.path: /mcp
mcp-client.timeoutMs: 5000
ollama.ollamaUrl: http://ollama.ai.svc.cluster.local:11434
ollama.model: llama3.1:8b
startup.yml: |
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
externalConfigDir: ${startup.externalConfigDir:/app/config-cache}
server.yml: |
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8083}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
serviceId: ${server.serviceId:com.networknt.agent.account-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
portal-registry.yml: |
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
client.yml: |
tls:
verifyHostname: ${client.verifyHostname:true}
mcp-client.yml: |
gatewayUrl: ${mcp-client.gatewayUrl:https://ai-microgateway.light-gateway:8443}
path: ${mcp-client.path:/mcp}
timeoutMs: ${mcp-client.timeoutMs:5000}
ollama.yml: |
ollamaUrl: ${ollama.ollamaUrl:http://ollama.ai.svc.cluster.local:11434}
model: ${ollama.model:llama3.1:8b}
Create the public/ ConfigMap from the repo asset:
kubectl -n light-agent create configmap light-agent-account-public \
--from-file=index.html=apps/light-agent/public/index.html \
--dry-run=client -o yaml
Example Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: light-agent-account
namespace: light-agent
labels:
app.kubernetes.io/name: light-agent-account
app.kubernetes.io/component: agent
app.kubernetes.io/part-of: lightapi
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/name: light-agent-account
template:
metadata:
labels:
app.kubernetes.io/name: light-agent-account
app.kubernetes.io/component: agent
app.kubernetes.io/part-of: lightapi
spec:
securityContext:
fsGroup: 999
fsGroupChangePolicy: OnRootMismatch
containers:
- name: light-agent
image: networknt/light-agent:2.2.1
imagePullPolicy: IfNotPresent
env:
- name: LIGHT_PORTAL_AUTHORIZATION
valueFrom:
secretKeyRef:
name: light-agent-account-secret
key: LIGHT_PORTAL_AUTHORIZATION
- name: light_4j_config_password
valueFrom:
secretKeyRef:
name: light-agent-account-secret
key: light_4j_config_password
optional: true
- name: LIGHT_AGENT_HOST_ID
valueFrom:
secretKeyRef:
name: light-agent-account-secret
key: LIGHT_AGENT_HOST_ID
- name: DATABASE_URL
valueFrom:
secretKeyRef:
name: light-agent-account-secret
key: DATABASE_URL
- name: RUST_LOG
value: info
- name: AGENT_LOG_ANSI
value: "false"
ports:
- name: http
containerPort: 8083
protocol: TCP
- name: https
containerPort: 8443
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 30
periodSeconds: 30
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 1000m
memory: 768Mi
volumeMounts:
- name: bootstrap-config
mountPath: /config
readOnly: true
- name: config-cache
mountPath: /app/config-cache
- name: public
mountPath: /app/public
readOnly: true
volumes:
- name: bootstrap-config
projected:
sources:
- configMap:
name: light-agent-account-config
- secret:
name: light-agent-account-secret
items:
- key: ca.pem
path: ca.pem
- name: config-cache
emptyDir: {}
- name: public
configMap:
name: light-agent-account-public
Example Service:
apiVersion: v1
kind: Service
metadata:
name: light-agent-account
namespace: light-agent
labels:
app.kubernetes.io/name: light-agent-account
app.kubernetes.io/component: agent
spec:
type: ClusterIP
selector:
app.kubernetes.io/name: light-agent-account
ports:
- name: http
port: 8083
targetPort: http
protocol: TCP
- name: https
port: 8443
targetPort: https
protocol: TCP
External Access
For local testing with a ClusterIP Service:
kubectl -n light-agent port-forward svc/light-agent-account 8083:8083
Health check:
curl -i http://127.0.0.1:8083/health
If exposing through Ingress, make sure WebSocket upgrade is supported and idle timeouts are long enough for chat sessions.
Example NGINX Ingress annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
nginx.ingress.kubernetes.io/backend-protocol: "HTTP"
If downstream MCP tools require caller identity, put the agent behind a BFF or
authenticated reverse proxy that forwards the user’s Authorization header to
the WebSocket request. A browser-created WebSocket from the embedded static UI
does not directly set arbitrary authorization headers.
Deploy Through Light-Deployer
The repo template lives at:
apps/light-agent/k8s/light-agent
Use the same template rules as light-gateway.
When light-deployer runs outside the cluster and has
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR set, repoUrl: "local" can point to local
templates.
When light-deployer runs inside Kubernetes, use a real Git URL:
{
"template": {
"repoUrl": "https://github.com/networknt/light-fabric.git",
"ref": "main",
"path": "apps/light-agent/k8s/light-agent"
}
}
Do not use repoUrl: "local" for an in-cluster deployer unless the template
repo is mounted into the deployer container and
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR points to it.
Keep Namespace out of templates rendered by light-deployer if the deployer
policy blocks cluster-scoped resources. Create the namespace separately:
kubectl create namespace light-agent
Config-Server Requirements
Before deploying the agent pod, config-server should already have config for the tuple used by startup:
host = startup.host
serviceId = server.serviceId
envTag = server.environment
At minimum, config-server should return runtime config for:
values.ymlserver.ymlwhen listener or registration settings are centrally managed.portal-registry.ymlwhen controller URLs or registry settings are centrally managed.client.ymlwhen TLS verification behavior is centrally managed.
For the current light-agent, keep mcp-client.yml and ollama.yml in the
local bootstrap ConfigMap even if other runtime config comes from config-server.
They are loaded before remote bootstrap completes.
Startup Flow
Expected runtime flow:
Kubernetes starts pod
-> /app/light-agent
-> read local /config/values.yml, ollama.yml, and mcp-client.yml
-> connect to Postgres with DATABASE_URL
-> build the MCP client for light-gateway
-> call config-server with LIGHT_PORTAL_AUTHORIZATION
-> write downloaded runtime config into /app/config-cache
-> start the Axum HTTP/WebSocket server
-> register the agent with controller using portalRegistry.portalUrl
-> serve the chat UI from /app/public
-> forward tool discovery and tool calls to light-gateway
When startup.yml configures config-server, the runtime tries to download the
latest values.yml before starting. If that download fails for any reason, the
runtime continues startup with the available local and cached config, including
/app/config-cache/values.yml when present.
Upgrade And Rollback
Use Kubernetes rolling updates with immutable image tags:
kubectl -n light-agent set image deploy/light-agent-account \
light-agent=networknt/light-agent:2.2.2
kubectl -n light-agent rollout status deploy/light-agent-account
Rollback:
kubectl -n light-agent rollout undo deploy/light-agent-account
For production, prefer changing only one variable at a time: either image tag or config-server runtime config, not both in the same rollout.
Validation Checklist
After deployment:
kubectl -n light-agent rollout status deploy/light-agent-accountsucceeds.- Pods are ready and restart count is stable.
- Logs show successful Postgres connection.
- Logs show successful config-server bootstrap.
- Logs show successful controller registration.
- Controller shows the agent registered with the expected service id, environment, host, and port.
curl http://127.0.0.1:8083/healthsucceeds through port-forward or Ingress.- The chat UI loads.
- The chat WebSocket connects to
/chat. - MCP
tools/listreacheslight-gateway. - MCP
tools/callreaches the backend MCP server throughlight-gateway. - A pod restart still starts cleanly with the selected cache policy.
Security Checklist
- Keep bearer tokens, config passwords, database passwords, and host ids in
Kubernetes
Secret, notConfigMap. - Use customer CA trust and keep
client.verifyHostname: truein production. - Use immutable image tags and image pull credentials from Kubernetes secrets when the registry is private.
- Run as the non-root image user.
- Make
/configread-only. - Make only
/app/config-cachewritable. - Restrict ingress traffic to required agent ports.
- Restrict egress traffic to config-server, controller,
light-gateway, Ollama, and Postgres. - Rotate
LIGHT_PORTAL_AUTHORIZATIONthrough the customer secret process.
Light-Deployer
light-deployer is the cluster-local Kubernetes deployment executor for Light
Portal.
It renders Kubernetes templates, validates manifests, applies resources through
kube-rs, reports rollout status, and exposes deployment tools through an MCP
JSON-RPC endpoint for local and MicroK8s testing.
Key Capabilities
- MCP JSON-RPC endpoint at
POST /mcp - AST-based YAML template rendering
- Git template fetching with
gix - Kubernetes dry-run, apply, delete, status, and prune
- redacted manifest summaries and diffs
- SSE deployment events
Runtime
light-deployer uses light-runtime, light-axum, config-loader, and
portal-registry so it follows the same service boot model as light-agent.
Testing Path
Use these pages in order when testing locally:
Start with standalone noop mode to validate template rendering. Then move to
MicroK8s real mode once the render request and target templates are correct.
For MCP clients, Light Portal, and AI agents, use POST /mcp with JSON-RPC
methods such as tools/list and tools/call. The /mcp/tools/* routes are
kept only as local debugging conveniences.
Build Local
This page builds the light-deployer binary and container image from the
Light Fabric workspace.
Run all commands from the repository root:
cd ~/workspace/light-fabric
Rust Build
Use cargo check first for a quick compile validation:
cargo check -p light-deployer
Run the deployer tests:
cargo test -p light-deployer
Build a debug binary:
cargo build -p light-deployer
Build a release binary:
cargo build --release -p light-deployer
The release binary is written to:
target/release/light-deployer
Docker Image
Build the local image:
./apps/light-deployer/build.sh latest
The default image name is:
networknt/light-deployer:latest
To override the image name:
IMAGE=localhost:32000/light-deployer:latest ./apps/light-deployer/build.sh latest
Verify the image exists:
docker image inspect networknt/light-deployer:latest
What The Image Contains
The Dockerfile copies:
/usr/local/bin/light-deployer/app/config
The container runs from /app, so the default runtime config directory is:
/app/config
The default HTTP port is 7088, configured in:
apps/light-deployer/config/server.yml
Expected Result
Before moving on, these commands should pass:
cargo check -p light-deployer
cargo test -p light-deployer
./apps/light-deployer/build.sh latest
docker image inspect networknt/light-deployer:latest
Prepare Config
light-deployer uses two kinds of configuration:
- runtime config loaded by
light-runtime - deployment request data sent through MCP
tools/callatPOST /mcp
Runtime Config Files
Default config lives in:
apps/light-deployer/config
Files:
server.yml: HTTP/HTTPS bind settings and service identitydeployer.yml: local deployer policyportal-registry.yml: future portal/controller registry settings
When running from the workspace root, the deployer automatically uses:
apps/light-deployer/config
When running inside the Docker image, it uses:
/app/config
Override the config directory with:
LIGHT_DEPLOYER_CONFIG_DIR=/path/to/config
Server Config
The default server config listens on HTTP port 7088:
ip: ${server.ip:0.0.0.0}
httpPort: ${server.httpPort:7088}
enableHttp: ${server.enableHttp:true}
enableHttps: ${server.enableHttps:false}
serviceId: ${server.serviceId:com.networknt.light-deployer-0.1.0}
enableRegistry: ${server.enableRegistry:false}
To change the port without editing the file, provide values through the normal runtime values mechanism, or use a copied config directory for local testing.
Deployer Policy
The default policy is permissive enough for local testing:
deployerId: ${deployer.deployerId:local-light-deployer}
clusterId: ${deployer.clusterId:local}
allowedNamespaces: []
allowedRepoHosts: []
allowedRepoPrefixes: []
allowedImageRegistries: []
devInsecure: ${deployer.devInsecure:false}
Empty allow lists mean the policy does not restrict that dimension. For production, configure explicit values.
Example tighter policy:
deployerId: petstore-microk8s
clusterId: microk8s-local
allowedNamespaces:
- petstore-dev
allowedRepoHosts:
- github.com
allowedRepoPrefixes:
- https://github.com/networknt/
allowedImageRegistries:
- networknt
devInsecure: false
prune:
enabled: true
maxDeletePercent: 30
sensitiveKinds:
- PersistentVolumeClaim
overrideRequired: true
Git Access
Public repositories do not need credentials.
For private HTTPS repositories, set:
LIGHT_DEPLOYER_GIT_TOKEN=...
Defaults:
- GitHub username:
x-access-token - Bitbucket Cloud username:
x-token-auth
For Bitbucket app passwords or other Git servers:
LIGHT_DEPLOYER_GIT_USERNAME=my-user
LIGHT_DEPLOYER_GIT_TOKEN=my-token-or-app-password
Only HTTPS token auth is supported in Phase 1. SSH auth is deferred.
Template Repository Requirements
The target application repository should contain a k8s/ directory with YAML
templates. The deployer reads all .yaml and .yml files under the requested
template path.
Example template reference:
{
"template": {
"repoUrl": "https://github.com/networknt/openapi-petstore.git",
"ref": "master",
"path": "k8s"
}
}
For local testing without Git clone, set:
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR=/home/steve/workspace/openapi-petstore
Then use:
{
"template": {
"repoUrl": "local",
"ref": "master",
"path": "k8s"
}
}
Request Values
The request values object supplies placeholder values for templates.
Example for openapi-petstore:
{
"name": "openapi-petstore",
"image": {
"repository": "networknt/openapi-petstore",
"tag": "latest",
"pullPolicy": "IfNotPresent"
},
"service": {
"name": "openapi-petstore",
"type": "ClusterIP"
},
"resources": {
"requests": {
"memory": "64Mi",
"cpu": "250m"
},
"limits": {
"memory": "256Mi",
"cpu": "500m"
}
}
}
The current renderer replaces placeholders inside YAML string scalar values. Avoid placeholders in Kubernetes fields that must be numeric unless the template keeps those fields as fixed numbers.
Run Standalone
Standalone mode is the fastest way to test light-deployer before using a
real Kubernetes cluster.
Use noop mode first. It validates config, HTTP endpoints, template loading,
rendering, resource summaries, and response shape without mutating Kubernetes.
Run all commands from:
cd /home/steve/workspace/light-fabric
Start With Built-In Sample
Start the deployer with the sample template directory:
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR=apps/light-deployer/examples/petstore \
LIGHT_DEPLOYER_KUBE_MODE=noop \
cargo run -p light-deployer
The service listens on:
http://127.0.0.1:7088
Check health from another terminal:
curl -fsSL http://127.0.0.1:7088/health
Expected output:
ok
List Tools With MCP JSON-RPC
The MCP endpoint is JSON-RPC 2.0 over HTTP at:
POST /mcp
List all deployment tools:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "tools-list-1",
"method": "tools/list",
"params": {}
}'
Call a tool through MCP:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "render-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
For local debugging, the deployer also exposes REST-style convenience endpoints:
curl -fsSL http://127.0.0.1:7088/mcp/tools/list
curl -fsSL http://127.0.0.1:7088/mcp/tools
curl -fsSL http://127.0.0.1:7088/mcp/tools/deployment.render
Use POST /mcp for MCP clients and AI agents.
Render The Built-In Sample
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "render-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
Expected response shape:
{
"jsonrpc": "2.0",
"result": {
"isError": false,
"structuredContent": {
"action": "render",
"status": "rendered",
"deployerId": "local-light-deployer",
"clusterId": "local",
"resources": [
{
"kind": "Deployment",
"name": "petstore"
},
{
"kind": "Service",
"name": "petstore"
}
]
}
}
}
The exact requestId and manifestHash will differ.
Render openapi-petstore Locally
If /home/steve/workspace/openapi-petstore is available and has a k8s/
folder, run:
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR=/home/steve/workspace/openapi-petstore \
LIGHT_DEPLOYER_KUBE_MODE=noop \
cargo run -p light-deployer
Render request:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "render-openapi-petstore-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "openapi-petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "petstore-dev",
"values": {
"name": "openapi-petstore",
"image": {
"repository": "networknt/openapi-petstore",
"tag": "latest",
"pullPolicy": "IfNotPresent"
},
"service": {
"name": "openapi-petstore",
"type": "ClusterIP"
},
"resources": {
"requests": {
"memory": "64Mi",
"cpu": "250m"
},
"limits": {
"memory": "256Mi",
"cpu": "500m"
}
}
},
"template": {
"repoUrl": "local",
"ref": "master",
"path": "k8s"
}
}
}
}'
Expected resources:
Deployment/openapi-petstoreService/openapi-petstore
Test Git Fetch
Stop the local-template run and restart without LIGHT_DEPLOYER_TEMPLATE_BASE_DIR:
LIGHT_DEPLOYER_KUBE_MODE=noop \
cargo run -p light-deployer
Render from GitHub:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "render-git-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "openapi-petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "petstore-dev",
"values": {
"name": "openapi-petstore",
"image": {
"repository": "networknt/openapi-petstore",
"tag": "latest"
}
},
"template": {
"repoUrl": "https://github.com/networknt/openapi-petstore.git",
"ref": "master",
"path": "k8s"
}
}
}
}'
For a private repository:
LIGHT_DEPLOYER_GIT_TOKEN=... \
LIGHT_DEPLOYER_KUBE_MODE=noop \
cargo run -p light-deployer
For Bitbucket app-password style auth:
LIGHT_DEPLOYER_GIT_USERNAME=my-user \
LIGHT_DEPLOYER_GIT_TOKEN=my-app-password \
LIGHT_DEPLOYER_KUBE_MODE=noop \
cargo run -p light-deployer
Dry Run And Diff In Noop Mode
Noop mode can also exercise the request path for these tools:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "dry-run-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.dryRun",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "diff-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.diff",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
These calls do not validate against Kubernetes unless real mode is enabled.
Stop The Service
Press Ctrl-C in the terminal running cargo run.
Run Kubernetes
This page runs light-deployer inside MicroK8s and uses the in-cluster
ServiceAccount with kube-rs.
Prerequisites
MicroK8s should be running and microk8s kubectl should work:
microk8s status --wait-ready
microk8s kubectl get nodes
Build the image first:
cd /home/steve/workspace/light-fabric
./apps/light-deployer/build.sh latest
Import Image Into MicroK8s
docker save networknt/light-deployer:latest | microk8s ctr image import -
If your MicroK8s install requires elevated permissions:
docker save networknt/light-deployer:latest | sudo microk8s ctr image import -
Verify the image is available:
microk8s ctr images ls | grep light-deployer
Install Deployer
Apply the included manifests:
microk8s kubectl apply -f apps/light-deployer/k8s/namespace.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/rbac.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/deployment.yaml
microk8s kubectl apply -f apps/light-deployer/k8s/service.yaml
Wait for the pod:
microk8s kubectl -n light-deployer rollout status deploy/light-deployer
microk8s kubectl -n light-deployer get pods
Check logs:
microk8s kubectl -n light-deployer logs deploy/light-deployer
The deployment sets:
LIGHT_DEPLOYER_KUBE_MODE=real
So the service uses real Kubernetes API calls from inside the cluster.
Port Forward
microk8s kubectl -n light-deployer port-forward svc/light-deployer 7088:7088
In another terminal:
curl -fsSL http://127.0.0.1:7088/health
Expected:
ok
List Tools
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "tools-list-1",
"method": "tools/list",
"params": {}
}'
The response contains the deployer’s tool names, descriptions, input schemas, and invocation metadata. Light Portal can use this JSON-RPC response to populate MCP tools for the API details view.
Render In Kubernetes
Rendering does not mutate the cluster:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "render-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.render",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
Dry Run In Kubernetes
Dry-run renders the manifest and asks the Kubernetes API to validate it without persisting resources:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "dry-run-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.dryRun",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
Expected status:
{
"jsonrpc": "2.0",
"result": {
"isError": false,
"structuredContent": {
"status": "validated"
}
}
}
Deploy Sample
The sample request deploys into the light-deployer namespace so it matches
the included namespace-scoped RBAC.
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "apply-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.apply",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "petstore",
"replicas": 1,
"image": {
"repository": "nginx",
"tag": "1.27"
},
"containerPort": 80,
"service": {
"port": 80
}
},
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
The response should return quickly with an accepted/applying-style status. The operation continues in the deployer.
Watch Kubernetes resources:
microk8s kubectl -n light-deployer get deploy,svc,pods
Stream Events
Use the requestId from the deployment response:
curl -N "http://127.0.0.1:7088/events?request_id=<requestId>"
The event stream reports deployment progress and failures for that request.
Check Status
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "status-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.status",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
Undeploy Sample
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "delete-sample-1",
"method": "tools/call",
"params": {
"name": "deployment.delete",
"arguments": {
"hostId": "local-host",
"instanceId": "petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"template": {
"repoUrl": "local",
"ref": "main",
"path": "k8s"
}
}
}
}'
Then verify resources:
microk8s kubectl -n light-deployer get deploy,svc,pods
Deploy openapi-petstore From Git
After the openapi-petstore repository has a k8s/ folder committed, use a
request like this:
curl -fsSL http://127.0.0.1:7088/mcp \
-H 'content-type: application/json' \
-d '{
"jsonrpc": "2.0",
"id": "apply-openapi-petstore-1",
"method": "tools/call",
"params": {
"name": "deployment.apply",
"arguments": {
"hostId": "local-host",
"instanceId": "openapi-petstore-dev",
"environment": "dev",
"clusterId": "microk8s-local",
"namespace": "light-deployer",
"values": {
"name": "openapi-petstore",
"image": {
"repository": "networknt/openapi-petstore",
"tag": "latest",
"pullPolicy": "IfNotPresent"
},
"service": {
"name": "openapi-petstore",
"type": "ClusterIP"
}
},
"template": {
"repoUrl": "https://github.com/networknt/openapi-petstore.git",
"ref": "master",
"path": "k8s"
}
}
}
}'
For private Git access, set LIGHT_DEPLOYER_GIT_TOKEN on the deployer pod.
In Kubernetes this should be injected from a Secret, not written directly into
the deployment manifest.
Update The Deployer Image
After rebuilding locally:
./apps/light-deployer/build.sh latest
docker save networknt/light-deployer:latest | microk8s ctr image import -
microk8s kubectl -n light-deployer rollout restart deploy/light-deployer
microk8s kubectl -n light-deployer rollout status deploy/light-deployer
Remove The Deployer
microk8s kubectl delete -f apps/light-deployer/k8s/service.yaml
microk8s kubectl delete -f apps/light-deployer/k8s/deployment.yaml
microk8s kubectl delete -f apps/light-deployer/k8s/rbac.yaml
microk8s kubectl delete -f apps/light-deployer/k8s/namespace.yaml
Light-Gateway
light-gateway is the Pingora-based gateway product in Light Fabric.
It is intended to host gateway behavior such as routing, proxying, and eventually AI/MCP gateway integrations while using the shared runtime and config model.
Key Dependencies
light-runtimelight-pingoraconfig-loader
Runtime
The gateway uses light-pingora as its transport framework and
light-runtime for lifecycle, bootstrap, and service configuration.
Endpoint Identity
Status
- Decision state: Accepted for implementation
- Owner: Light Gateway maintainers
- Decision date: 2026-08-06
- Revised: 2026-08-06
- Tracking issue: networknt/light-fabric#297
Purpose
Normal HTTP requests need the method in their endpoint identity. A path alone
cannot distinguish operations such as GET /v1/models and
POST /v1/models.
The generated access-control snapshot already uses method-qualified keys:
/v1/models@get
/v1/chat/completions@post
The gateway previously sent only /v1/models to access control, so the
generated /v1/models@get rule could not match. This design fixes that
mismatch. It does not introduce schema versions, capability negotiation,
legacy modes, dual rule formats, or changes to generated rules.
Terms
| Value | Example | Used for |
|---|---|---|
| Request path | /v1/accounts/123 | Routing, URI rewriting, rate limiting, and path-prefix checks |
| Path template | /v1/accounts/{accountId} | Stable endpoint and metrics dimensions |
| HTTP method | GET | Transport behavior and endpoint qualification |
| HTTP endpoint | /v1/accounts/{accountId}@get | Access control, response filtering, logs, audit, and endpoint metrics |
Paths and endpoints are different values. Code that routes or rewrites a URI uses a path. Code that identifies an operation uses an endpoint.
Identity Rules
Normal HTTP
A normal HTTP endpoint is:
{matched-path-template-or-request-path}@{lowercase-method}
Examples:
/v1/models@get
/v1/chat/completions@post
/v1/accounts/{accountId}@patch
The matched handler template is preferred because it avoids concrete IDs in policies, logs, and metric dimensions. The query string is never part of the endpoint.
HTTP methods are distinct. A GET rule must not authorize POST, PUT, PATCH, DELETE, HEAD, or OPTIONS on the same path.
Portal Hybrid Requests
Portal query and command requests multiplex operations over shared transport paths. Their access-control identity remains the generated logical operation ID already derived from the request envelope:
lightapi.net/service/getApi/0.1.0
The Portal server accepts GET and POST transports with the same semantics, so the transport method is not added to this logical ID. Existing Portal rules remain unchanged and match exactly.
MCP and OpenAPI Tools
Existing tool endpoint rules remain unchanged:
- Native MCP operations use their configured
@callidentity, such asweather@call. - OpenAPI-backed tools use the proxied HTTP identity, such as
/offers@get.
This change does not make MCP catalog fields mandatory and does not couple access-control rule loading to catalog loading.
WebSocket
WebSocket connection authorization uses path@connect, including the existing
controller endpoint:
/ctrl/mcp@connect
The controller identity is anchored to the concrete /ctrl/mcp request path,
not to a matched handler template. This preserves the controller route’s
fail-closed behavior even when handler configuration uses a template or
wildcard that also matches the controller path.
WebSocket routing still uses the upgrade path. It must not use an HTTP
@get endpoint as its connection-policy identity.
Access-Control Matching
Access control compares the operation as well as the selector:
- Exact endpoint match.
- Template or parent-path match only when both identities have the same operation suffix.
For example:
| Rule | Request endpoint | Match |
|---|---|---|
/v1/models@get | /v1/models@get | Yes |
/v1/models@get | /v1/models@post | No |
/v1/accounts/{id}@get | /v1/accounts/123@get | Yes |
/v1/accounts/{id}@get | /v1/accounts/123@delete | No |
Methodless logical IDs, such as Portal operation IDs, match exactly. They are
not implicitly converted to @call, and a qualified HTTP lookup never falls
back to a methodless path rule.
Endpoint parsing splits on the final @, allowing selectors such as
/users/[email protected]@get.
defaultDeny keeps its existing meaning after lookup:
true: an unmatched endpoint is denied;false: an unmatched endpoint is allowed.
The fix is to make the generated rule and runtime endpoint agree, not to alter that policy setting.
Consumer Boundaries
| Consumer | Input |
|---|---|
| Access control | Endpoint identity |
| Request and response filtering | Endpoint identity |
| Endpoint metrics | Endpoint identity plus existing method field |
| Logs and audit | Endpoint identity |
| Router selection and rewrites | Request path |
| Upstream URI construction | Request path and query |
| Rate limiting | Request path |
skipPathPrefixes | Request path |
Routing code must never append @method to an upstream URI. The current router
matches query-rewrite rules against the request path first and accepts endpoint
only as a secondary lookup key; it constructs the upstream URI exclusively
from the original and rewritten path. That existing fallback does not append
the endpoint to the path and does not need to change for this issue.
Metrics
The endpoint dimension includes the operation:
endpoint=/v1/accounts/{accountId}@get
method=GET
pathTemplate=/v1/accounts/{accountId}
The separate method field remains useful for method-wide aggregation.
pathTemplate provides path-oriented aggregation without parsing the endpoint.
When no template matches, use the bounded <unmatched> value rather than the
concrete request path.
Adding pathTemplate is an observability change only. It does not change rule
or routing configuration.
Request Flow
HTTP request
-> preserve request path and method
-> resolve handler and matched path template
-> render path-template@lowercase-method
-> authorize and filter with that endpoint
-> record endpoint metrics
-> route and build the upstream URI from path values
Portal, MCP, and WebSocket handlers replace or select the access-control identity at their existing protocol boundary as described above.
Development Cutover
There is one endpoint contract, with no transition mode:
- normal HTTP uses
path@method; - Portal logical IDs remain methodless;
- MCP uses existing
@callidentities; - WebSocket uses
@connect.
The current generated snapshot already follows this contract. No rule or configuration changes are required. Deploy the gateway code and restart the development environment together. If a development snapshot contains a methodless normal HTTP key, regenerate it instead of adding a runtime fallback.
Required Tests
| Scenario | Expected result |
|---|---|
GET /v1/models with /v1/models@get rule | Allowed when its rule permits access |
POST /v1/models with only a GET rule | Does not match the GET rule |
Template route /v1/accounts/{id} | Uses /v1/accounts/{id}@method |
| Response filtering | Uses the same endpoint as authorization |
| Portal GET and POST transports | Resolve to the same generated logical ID |
| Native MCP tool | Keeps its configured @call identity |
| OpenAPI-backed tool | Keeps its proxied HTTP identity |
| WebSocket upgrade | Uses path@connect for policy |
| Router and upstream URI | Never receive an @operation suffix |
| Metrics | Emit endpoint, method, and stable path template |
Tests cover both defaultDeny values so a method mismatch cannot be mistaken
for a successful rule match.
Decisions
- Only normal HTTP endpoint construction changes for issue #297.
- Existing generated rules and configuration are not changed.
- No endpoint schema version or gateway capability is introduced.
- No legacy identity mode or dual lookup is implemented.
- HTTP endpoint methods are lowercase.
- Access control and response filtering match the exact operation.
- Portal, MCP, and WebSocket retain their existing protocol identities.
- Routing, URI rewriting, rate limiting, and path-prefix behavior remain path-based.
SSE Passthrough Parity
Status
- Decision state: Accepted; Phases 0 through 3 implemented
- Owner: Light Gateway maintainers
- Created: 2026-09-03
- Reference: networknt/light-4j#2761
Purpose
The Java light-4j proxy and router support long-lived Server-Sent Events
(SSE) responses without applying the ordinary whole-request timeout. They can
identify an expected stream from the request Accept header or request path,
confirm a stream from the upstream Content-Type, protect upstream streaming
headers, and optionally close a stream after an idle period.
Before this implementation, light-fabric/apps/light-gateway forwarded
ordinary Pingora response body chunks as they arrived. It also had dedicated
streaming implementations for LLM and MCP traffic, but did not implement the
generic proxy and router configuration or lifecycle behavior introduced by
light-4j issue 2761.
This design closes that parity gap for ordinary proxy and router handler
chains. It does not replace the specialized LLM, MCP, A2A, or WebSocket
streaming implementations.
Baseline State
The pre-parity baseline had four materially different behaviors:
| Area | Current behavior |
|---|---|
| Ordinary proxy/router response | Pingora forwards each response body chunk without assembling the complete body. |
ProxyConfig and RouterConfig | They have maxRequestTime; the router also has pathPrefixMaxRequestTime. Neither value is currently enforced by the gateway request lifecycle. |
| Generic SSE recognition | There is no request Accept, request path, or upstream Content-Type classification. |
| Specialized streams | LLM and MCP use explicit streaming writers and their own policies. Model-provider sidecar and WebSocket paths also apply specialized timeout controls. |
Ordinary chunk forwarding means a simple upstream SSE response can appear to work today. That is not equivalent to the Java feature:
- there is no streaming-specific whole-exchange timeout;
- there is no configurable idle timeout between upstream bytes;
- upstream streaming headers are not protected from gateway header mutation;
- there is no response-side promotion when only the upstream response reveals that the exchange is a stream; and
- response-body transformations can buffer the stream until end-of-stream.
The last case is particularly important. Detokenization and access-control response filtering currently collect the complete response before emitting a transformed body. An unbounded SSE response cannot safely enter either path.
Goals
- Provide the same six operator-facing streaming properties in both
proxy.ymlandrouter.yml. - Preserve incremental forwarding for generic SSE responses over upstream and downstream HTTP/1.1 and HTTP/2 combinations.
- Select the streaming timeout before upstream response headers arrive for an
operator-configured stream path. Treat a client
Acceptmatch as provisional until the upstream response confirms streaming. - Promote an ordinary request to streaming behavior when the upstream
Content-Typeidentifies a stream. - Enforce an optional idle timeout that resets whenever upstream response bytes arrive.
- Keep timeout and stream state isolated to one exchange and one immutable configuration snapshot.
- Make incompatible response-body handlers fail closed instead of silently buffering an unbounded response or bypassing policy.
- Preserve existing deployments when the new properties are absent.
Non-Goals
- Do not merge generic SSE with the LLM, MCP, A2A, or WebSocket protocol implementations.
- Do not parse, reframe, validate, or synthesize SSE events in the generic proxy path. Event data remains opaque bytes.
- Do not add SSE replay,
Last-Event-IDstorage, event persistence, or delivery guarantees. - Do not make
X-Accel-Bufferingmandatory. That header is an optional deployment concern, not part of the Java parity contract. - Do not allow streaming classification to bypass authentication, authorization, admission control, rate limiting, request validation, or audit requirements.
- Do not apply configuration reloads retroactively to an exchange that is already streaming.
Decision
Add one shared generic streaming policy to light-pingora and embed it in both
ProxyConfig and RouterConfig. The selected route copies the effective
policy into GatewayRequestContext; later configuration reloads therefore
affect new requests only.
Classify the exchange in two stages:
- Expected stream: before proxying, match the request path or
Acceptheader. A path match selectsstreamMaxRequestTimeimmediately; anAcceptmatch retains the ordinary deadline until confirmed so an untrusted client cannot disable it. - Confirmed stream: after receiving upstream headers, match
Content-Type. Protect streaming headers, cancel or replace the ordinary exchange deadline, and enable the stream idle timeout.
The response body remains on Pingora’s normal chunked forwarding path. The gateway must not introduce a second buffering or event-decoding layer.
flowchart TD
A[Receive request] --> B[Select proxy or router config snapshot]
B --> C{Path or Accept identifies stream?}
C -- Trusted path --> D[Install stream exchange deadline]
C -- Accept or no match --> E[Install ordinary exchange deadline]
D --> F[Connect and send upstream request]
E --> F
F --> G[Receive upstream headers]
G --> H{Content-Type identifies stream?}
H -- Yes --> I[Confirm stream and update deadline]
H -- No --> J[Keep selected ordinary behavior]
I --> K{Buffered response handler active?}
K -- Yes --> L[Fail closed before response headers are committed]
K -- No --> M[Normalize streaming headers]
J --> N[Normal response processing]
M --> O[Forward each body chunk]
O --> P[Reset upstream idle deadline]
P --> O
Configuration Contract
The Rust names and defaults must match the Java configuration so the same Config Server properties can be projected into either runtime.
| Property | Rust type | Default | Meaning |
|---|---|---|---|
streamResponseContentTypes | list of strings | ["text/event-stream"] | Upstream response media types that confirm streaming behavior. |
streamRequestAcceptTypes | list of strings | ["text/event-stream"] | Request Accept media types that select streaming behavior before upstream headers. |
streamPathPrefixes | list of strings | [] | Request path prefixes that select streaming behavior before upstream headers. |
streamMaxRequestTime | unsigned milliseconds | 0 | Maximum whole-exchange duration for a stream; zero disables the whole-exchange deadline. |
streamIdleTimeout | unsigned milliseconds | 0 | Maximum silence between upstream response bytes; zero disables the idle deadline. |
streamResponseHeaderOverwrite | list of header names | Content-Type, Cache-Control, Connection, Transfer-Encoding, Content-Encoding, Content-Length | Headers for which the upstream streaming response must remain authoritative. |
Example proxy configuration:
maxRequestTime: ${proxy.maxRequestTime:0}
streamResponseContentTypes: ${proxy.streamResponseContentTypes:["text/event-stream"]}
streamRequestAcceptTypes: ${proxy.streamRequestAcceptTypes:["text/event-stream"]}
streamPathPrefixes: ${proxy.streamPathPrefixes:}
streamMaxRequestTime: ${proxy.streamMaxRequestTime:0}
streamIdleTimeout: ${proxy.streamIdleTimeout:0}
streamResponseHeaderOverwrite: ${proxy.streamResponseHeaderOverwrite:["Content-Type","Cache-Control","Connection","Transfer-Encoding","Content-Encoding","Content-Length"]}
The router uses the same field names under the router property namespace.
An absent property uses the documented default. An explicitly empty media-type or header list disables that matching or protection category. Empty path prefixes are ignored. Timeout values are non-negative; zero has the disabling meaning shown above.
Media-Type Matching
Matching must follow the Java behavior:
- compare case-insensitively;
- split comma-separated header values;
- ignore media-type parameters such as
charset=utf-8andq=1.0; - trim surrounding whitespace; and
- require equality after normalization rather than substring matching.
For example, all of the following identify the default SSE media type:
Accept: text/event-stream
Accept: application/json, text/event-stream
Accept: TEXT/EVENT-STREAM; q=1.0
Content-Type: text/event-stream; charset=utf-8
Path matching uses the request path already used by routing, excludes the
query string, and follows the Java startsWith prefix semantics. Empty
configured prefixes never match.
Per-Exchange State
Add a small immutable policy value in frameworks/light-pingora, shared by the
proxy and router configuration models. After route selection, store the
following request-local state in GatewayRequestContext:
streamPolicySnapshot
streamExpected
streamConfirmed
exchangeDeadline
streamIdleTimeout
lastUpstreamProgress
responseHeadersCommitted
The exact Rust representation is an implementation detail, but it must meet these invariants:
- no request mutates
ProxyConfig,RouterConfig, a route snapshot, or a gateway-global timeout; - retry attempts share one absolute whole-exchange deadline;
- an upstream response can promote the exchange from ordinary to streaming;
- once response headers or body bytes are committed downstream, the exchange is never retried; and
- configuration reload cannot change the limits of an active exchange.
Timeout Semantics
Whole-Exchange Deadline
maxRequestTime and pathPrefixMaxRequestTime are whole-exchange deadlines,
not socket-idle deadlines. The implementation must first make the existing
ordinary timeout fields effective, then select the stream deadline as follows:
if a configured request path identifies a stream:
effective deadline = streamMaxRequestTime
else if Accept identifies a possible stream:
effective deadline = ordinary timeout until Content-Type confirms streaming
else if router pathPrefixMaxRequestTime has a matching prefix:
effective deadline = matched prefix value
else:
effective deadline = maxRequestTime
If more than one timeout prefix matches, the longest prefix wins. This makes the most specific route policy authoritative and avoids depending on map iteration order.
Zero disables the selected deadline. An ordinary request or provisional
Accept match can later be confirmed as streaming by the upstream
Content-Type; at that point the ordinary deadline is cancelled and replaced
by the remaining streamMaxRequestTime policy. With the default value of zero,
it is cancelled. A provisional match receiving a non-streaming response clears
its streaming classification and retains the ordinary deadline.
The deadline must cover connection acquisition, connection establishment, request upload, response-header wait, every retry, response streaming, and downstream body completion. It must be based on the request start time, not reset for each retry or response chunk. Terminal logging, audit persistence, replay-reservation release, and connection cleanup run outside the deadline so expiry cannot cancel correctness-critical bookkeeping.
Pingora’s PeerOptions.read_timeout is a per-I/O progress timeout and cannot
implement this whole-exchange contract by itself. Add a cancellable
whole-exchange deadline at the patched pingora-proxy request driver boundary,
with a default-disabled trait hook so other ProxyHttp implementations retain
their current behavior. GatewayProxy supplies and updates the request-local
deadline after route selection and response classification.
Before downstream headers are committed, expiry returns HTTP 504. After headers are committed, the gateway closes the stream and records the timeout; it cannot replace an in-progress SSE response with a new HTTP error document.
Stream Idle Deadline
When streamIdleTimeout is positive, apply it as the upstream read timeout
after the response is confirmed as streaming. Each successful upstream body
read resets that timeout naturally. The timer covers silence between response
chunks; the whole-exchange deadline remains responsible for the request and
response-header phases.
Do not use PeerOptions.idle_timeout for this purpose. In Pingora that option
controls how long a released connection remains in the connection pool; it is
not the timeout between bytes of an active response.
Apply the idle policy to subsequent reads immediately after response classification. The Pingora driver needs a request-local update hook because the peer was constructed before the response headers arrived.
On idle expiry, close the upstream and downstream exchange, mark the upstream connection non-reusable, emit a distinct stream-idle outcome, and do not retry after any response has been committed.
Response Headers And Framing
The Java handler removes selected outbound headers before copying the upstream
streaming headers. Pingora already represents the upstream response as the
response being sent downstream, so blindly deleting the configured headers
would remove valid upstream Content-Type and Cache-Control values.
Implement equivalent outcome-based semantics:
- capture the configured authoritative upstream header values when the response is confirmed as streaming;
- apply normal gateway response-header handlers;
- restore the upstream values for headers in
streamResponseHeaderOverwrite; and - normalize hop-by-hop framing for the downstream HTTP version.
For a streaming response:
- remove
Content-Lengthunless the response has already ended with a known, complete body; - never forward contradictory
Content-LengthandTransfer-Encoding; - let Pingora generate HTTP/1.1 chunked framing when the body length is unknown;
- do not emit
Transfer-Encodingon HTTP/2; - preserve the upstream
Content-Type, including parameters; - preserve upstream
Cache-ControlandContent-Encodingunless a documented security policy rejects that encoding; and - strip or regenerate hop-by-hop
Connectionsemantics for the downstream protocol.
Header protection does not bypass mandatory correlation, CORS, rate-limit, or security headers when those headers are outside the configured overwrite set. Configuration validation must reject invalid header names.
Handler Compatibility
Streaming classification changes transport behavior only. All request-side security handlers still run before upstream selection.
Response handlers fall into two groups:
| Handler behavior | Streaming rule |
|---|---|
| Header-only, accounting, byte counting, logging, and metrics | Continue incrementally. |
| Complete-body transformation or inspection | Incompatible unless redesigned around an explicitly bounded streaming algorithm. |
The current detokenization and access-control response filters are complete-body handlers. The gateway must fail closed when either is active for an expected or confirmed stream:
- if an operator-configured path identifies streaming, reject it before opening the upstream exchange;
- treat an
Accept-only match as provisional and reject only if the upstream response confirms streaming; - if only the upstream response reveals streaming, reject before committing upstream response headers downstream; and
- never disable a configured security filter merely to permit streaming.
The implementation must allocate a stable gateway error code for this configuration/runtime incompatibility before release. The response should use HTTP 502 when an unexpected upstream streaming response conflicts with the configured handler chain. A deployment-time validator should also report known path-based conflicts so operators can correct them before traffic arrives.
Phase 2 assigns ERR13027 to this incompatibility. Known path-prefix
conflicts are reported during gateway construction/config validation, while
request-only and response-only classifications remain protected by the
runtime fail-closed checks.
Cache, Retry, And Connection Rules
- Generic SSE responses are not cached.
- An expected stream disables response caching before contacting upstream.
- A response-side stream confirmation disables any cache admission that has not already committed. Cache lookup must not serve a previously buffered object as a live SSE stream.
- Connection and pre-header retries remain allowed only while the request is replay-safe and the whole-exchange deadline has time remaining.
- No retry is allowed after downstream response headers or any body bytes have been written.
- An idle-timeout or malformed-framing connection is not returned to the upstream pool.
- Graceful shutdown follows the existing gateway drain policy. Active streams may continue only within the configured shutdown drain deadline.
Observability
Add bounded dimensions rather than raw paths or media-type values:
gateway_stream_kind = none | generic_sse | llm | mcp | a2a | websocket
gateway_stream_classification = request_accept | path_prefix | response_content_type
gateway_stream_outcome = completed | client_disconnect | upstream_error | exchange_timeout | idle_timeout | incompatible_handler | shutdown
Record counters for confirmed classifications and outcomes, plus stream
duration and bytes in each direction. An Accept preference followed by an
ordinary response is not a stream metric. Reuse the existing bounded endpoint
identity for the endpoint dimension. Logs should include the correlation ID,
endpoint, selected timeout values, classification source, per-exchange and
cumulative byte/duration totals, and outcome without logging SSE event payloads.
The access log must be emitted once, when the stream closes, rather than once per event. A long-lived stream must not retain unbounded per-event telemetry in memory.
Implementation Scope
frameworks/light-pingora
- Add the shared streaming policy and media-type/path matching helpers.
- Extend
ProxyConfigandRouterConfigwith the six parity properties and Java-compatible defaults. - Add focused configuration, normalization, and matching tests.
- Expose the selected policy through
ProxyRouteandRouterRoutewithout mutable global state.
apps/light-gateway
- Add request-local streaming classification and deadline state to
GatewayRequestContext. - Classify expected streams after the effective proxy/router route is selected.
- Apply the selected timeout policy in
upstream_peerand the request driver. - Confirm streams in
response_filterbefore response headers are committed. - Protect upstream streaming headers and normalize framing.
- Keep
response_body_filterincremental and update stream progress without copying or parsing event frames. - Reject incompatible complete-body response handlers.
- Add bounded metrics and terminal outcome logging.
Patched pingora-proxy
- Add a default-disabled, request-local whole-exchange deadline hook.
- Permit response classification to cancel or replace the active deadline.
- Permit response classification to update the active upstream read timeout.
- Ensure deadline expiry cancels both directions and prevents unsafe retry or connection reuse.
The Pingora patch should be kept narrow, covered by framework-level tests, and
structured for a possible upstream contribution. Generic SSE policy remains in
light-pingora and light-gateway; the patch provides transport lifecycle
primitives only.
Configuration And Documentation
- Add all six properties to the shipped
proxy.ymlandrouter.ymlfiles. - Publish the fields through the same runtime/configuration registration path as the existing proxy and router fields.
- Document units, defaults, zero-value behavior, media-type matching, and handler incompatibilities.
- Verify old configuration files deserialize to the compatibility defaults.
Delivery Plan
Phase 0: Configuration And Classification
Status: Complete (2026-09-03)
- Introduce the shared policy and six config properties.
- Add deterministic request and response classification helpers.
- Snapshot the selected policy into each request context.
- Prove old configurations retain their existing behavior.
Exit gate: configuration and matching unit tests pass, including missing, empty, mixed-case, parameterized, comma-separated, and invalid inputs.
Phase 1: Deadlines And Isolation
Status: Complete (2026-09-03)
- Make existing
maxRequestTimeandpathPrefixMaxRequestTimeeffective. - Add the generic request-driver deadline primitive.
- Select
streamMaxRequestTimeimmediately for operator-declared stream paths; keep the ordinary deadline for provisionalAcceptmatches until the upstream confirms a streaming response. - Support response-side cancellation or replacement of the ordinary deadline.
Exit gate: concurrent requests with different route timeouts prove that no request mutates shared timeout state; retry attempts consume one absolute deadline.
Phase 2: Streaming Response Safety
Status: Complete (2026-09-03)
- Apply the upstream read-idle timeout.
- Implement response-header authority and protocol-correct framing.
- Disable cache admission and post-commit retries.
- Reject incompatible response-body handlers.
- Add stream outcome metrics and logs.
Exit gate: live Pingora tests observe the first event before the upstream finishes, observe multiple separated events without coalescing the whole body, and prove idle closure, header behavior, and fail-closed handler interaction.
Phase 3: Protocol Matrix And Rollout
Status: Complete (2026-09-03)
- Exercise HTTP/1.1 and HTTP/2 on both sides of the gateway.
- Test direct proxy and service-router selection.
- Run soak and graceful-shutdown exercises with long-lived connections.
- Roll out first with the compatibility defaults and explicit stream path prefixes only where required.
Exit gate: the qualification matrix passes with bounded memory, stable file descriptor counts, no cross-request timeout interference, and no retry after response commitment.
Run the repeatable release gate from the repository root:
./scripts/run-sse-passthrough-phase3-gates.sh
The gate exercises all four downstream/upstream HTTP/1.1 and HTTP/2
combinations through a live Pingora listener, plus service-router selection.
It also runs 32 concurrent two-second responses with 100 ms heartbeats and
requires post-soak file descriptor growth of at most four and resident-memory
growth of at most 64 MiB. The same operational test proves that a response is
not retried after its headers and first event are committed, that shutdown
waits for an active stream to drain within the configured 500 ms graceful
period, and that a stream exceeding a 100 ms drain deadline is forcibly closed
and reported as a shutdown termination. Resource qualification is Linux-only
because it reads process metrics from /proc.
Phase 3 also makes http2Enabled effective for proxy and router upstreams.
When enabled, each selected upstream peer advertises HTTP/2 with HTTP/1.1
fallback through ALPN. The choice is captured in request context, so a config
reload or a concurrent request cannot change the protocol policy of an active
exchange.
Verification Matrix
| Test | Required evidence |
|---|---|
| Incremental passthrough | Client receives event 1 while the upstream connection remains open before event 2. |
| Accept detection | Default, mixed-case, parameterized, comma-separated, and multiple header values classify correctly. |
| Path detection | Configured prefixes select stream policy; empty or unrelated prefixes do not. |
| Response promotion | An ordinary request receiving text/event-stream cancels or replaces its ordinary deadline before streaming. |
| Ordinary timeout | Non-streaming requests still enforce maxRequestTime and longest matching pathPrefixMaxRequestTime. |
| Timeout isolation | Concurrent requests retain independent deadlines and config snapshots. |
| Idle timeout | Each chunk resets the idle timer; silence closes the stream; zero disables the timer. |
| Framing | No conflicting length/transfer headers across HTTP/1.1 and HTTP/2 combinations; HTTP/1.0 uses close-delimited framing without chunk markers. |
| Header authority | Configured upstream stream headers survive normal gateway header mutation. |
| Handler conflict | Detokenization and response filtering reject expected and response-discovered streams without bypass. |
| Retry boundary | Pre-header replay-safe failure may retry; post-commit failure never retries. |
| Cache boundary | Expected and confirmed SSE responses are not admitted to or served from cache. |
| Disconnect | Client disconnect cancels upstream work and releases permits and connections. |
| Reload | A config reload affects new requests and leaves an active stream on its captured policy. |
| Shutdown | Active streams drain only within the configured graceful-shutdown deadline. |
Compatibility And Rollout
The new fields are additive. Their Java-compatible defaults classify
text/event-stream, disable the stream whole-exchange and idle deadlines, and
protect the standard stream headers. For Rust rollout compatibility,
maxRequestTime defaults to zero so the newly effective ordinary
whole-exchange deadline is opt-in; an explicitly configured nonzero value and
pathPrefixMaxRequestTime remain enforced.
That correction can expose upstream calls that currently exceed configured timeouts. Before enabling enforcement in production:
- inventory configured proxy/router timeout values;
- compare them with observed non-stream request duration percentiles;
- add required path-specific exceptions;
- identify SSE endpoints by path where clients do not send an
Acceptheader; and - canary with timeout and stream outcome metrics enabled.
Use the following rollout sequence for each gateway deployment:
- leave the additive compatibility defaults unchanged;
- add
streamPathPrefixesonly for SSE routes whose clients omitAccept: text/event-stream; - run
run-sse-passthrough-phase3-gates.shagainst the release revision; - canary one instance and confirm stream completion, timeout, disconnect, and upstream-error outcomes while watching RSS and open-file trends;
- expand the canary only after ordinary-request timeout rates remain at the pre-rollout baseline; and
- enable upstream
http2Enabledindependently for proxy and router pools after their targets are confirmed to negotiate HTTP/2 correctly.
Rollback consists of reverting to the previous gateway image. Setting both stream timeouts to zero disables the new stream timers but does not disable classification, header correctness, cache safety, or handler-conflict checks.
Acceptance Criteria
The parity issue is complete when:
- proxy and router expose all six properties with Java-compatible defaults;
- request
Accept, request path, and responseContent-Typeclassification pass the verification matrix; - ordinary and streaming whole-exchange timeouts are enforced without shared state mutation;
- the stream idle timeout is based on active upstream read progress;
- SSE bytes reach the client incrementally with protocol-correct framing;
- response-body security transforms fail closed for unbounded streams;
- caching, retry, reload, disconnect, and shutdown boundaries are qualified;
- the dedicated LLM, MCP, A2A, and WebSocket tests remain green; and
- the Java and Rust configuration examples can use the same six field names and zero-value semantics.
Light Rule In Light-Gateway
light-gateway uses Light-Rule to enforce deterministic policy decisions in
the Pingora request path. Rules are written as inline
CEL expressions and evaluated entirely within the gateway
process — no external policy service is required.
The first production use is MCP tool authorization (req-acc) and response
filtering (res-fil) for the mcp handler.
This lets a gateway route agent MCP traffic to downstream MCP servers or API servers while enforcing fine-grained authorization locally from configuration delivered by config-server.
When It Runs
Light-Rule is invoked by light-gateway when all of these are true:
handler.ymlincludes themcphandler in the matched chain.mcp-router.ymlenables the MCP router and defines tools.access-control.ymland/orrule.ymlare available from local config or config-server.- A client sends
tools/callto the configured MCP endpoint, normally/mcp.
The dependency path is:
light-gateway
-> light-pingora
-> light-rule
light-gateway links light-pingora, and light-pingora links
light-rule. The rule engine is part of the gateway binary; there is no
dynamic plugin loading step.
Request Flow
For MCP traffic, the runtime flow is:
POST /mcp
-> handler.yml selects mcp
-> mcp-router parses JSON-RPC tools/call
-> access-control runtime builds rule context
-> light-rule evaluates req-acc CEL expressions
-> denied: return JSON-RPC error -32001
-> allowed: call downstream HTTP or MCP tool
-> light-rule evaluates optional res-fil CEL expressions
-> return JSON-RPC result
Authorization happens before the downstream call. Response filtering happens after the downstream response and before the MCP JSON-RPC response is returned to the agent.
Required Files
handler.yml
The mcp handler must be in the execution chain for the MCP path:
handlers:
- correlation
- security
- mcp
paths:
- path: /mcp
method: POST
exec:
- correlation
- security
- mcp
defaultHandlers: []
The security handler must run before mcp so that JWT claims are decoded
and available in the rule context when CEL expressions are evaluated.
mcp-router.yml
mcp-router.yml exposes the MCP endpoint and maps tools to downstream APIs or
downstream MCP servers:
enabled: true
path: /mcp
maxSessions: 10000
maxSessionsPerClient: 100
tools:
- name: weather
description: Get current weather for a city.
targetHost: http://weather-api:8080
path: /weather
method: GET
endpoint: weather@call
apiType: http
inputSchema:
type: object
properties:
city:
type: string
required:
- city
The endpoint field is the stable policy key used in rule.yml. If it is
omitted, the gateway derives one from the tool name and method, such as
weather@call.
maxSessions caps the total in-memory MCP frontend sessions for this gateway
process. maxSessionsPerClient caps sessions for one authenticated client or,
when no principal is available, one MCP clientInfo.name and
clientInfo.version pair.
For downstream MCP servers, set apiType: mcp. For downstream REST API
servers, use apiType: http or omit it when the default is acceptable.
access-control.yml
access-control.yml controls whether policy is active and how rules combine:
enabled: true
accessRuleLogic: any
defaultDeny: true
defaultInclude: false
skipPathPrefixes: []
logFullCelContext: false
Fields:
enabled: turns access-control evaluation on or off.accessRuleLogic:any(allow if any rule passes) orall(allow only if every rule passes) forreq-accrule IDs on an endpoint.defaultDeny: whentrue, deny calls with no matching endpoint rule.defaultInclude: whenfalse, a response row filter with no matching caller role, group, position, attribute, or user entry returns no rows. Settrueonly to preserve the legacy include-all row-filter behavior.skipPathPrefixes: endpoint prefixes that bypass access control entirely.logFullCelContext: controls CEL context values inlight_rule::celtrace events. The defaultfalsereports only statically referenced paths and structural metadata. Set it totrueonly for local or development debugging to include the bounded values of statically referenced properties. This property does not enable trace logging; use a filter such asRUST_LOG=light_rule::cel=trace,info.
The file name is access-control.yml. The loader also accepts
access-control.yaml.
rule.yml
rule.yml holds the CEL rule bodies and maps them to endpoints:
ruleBodies:
allow-scp-group.lightapi.net:
ruleId: allow-scp-group.lightapi.net
ruleName: Allow request when scp claim contains the required group
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
version: "1.0.0"
common: "Y"
actions: []
expression: |
'scp' in auditInfo.subject_claims.ClaimsMap
&& 'groups' in permission
&& permission.groups in auditInfo.subject_claims.ClaimsMap.scp
endpointRules:
weather@call:
req-acc:
- allow-scp-group.lightapi.net
permission:
groups: weather.r
Key fields in each rule body:
| Field | Required | Description |
|---|---|---|
ruleId | yes | Unique identifier, referenced from endpointRules. |
ruleName | yes | Human-readable description. |
ruleType | yes | req-acc for request authorization, res-fil for response filtering. |
conditionLanguage | yes | Must be cel. |
conditionSecurityProfile | yes | Must be strict (see Security Profile). |
expression | yes | CEL expression that must return true to allow the request. |
version | yes | Semantic version string. |
common | no | "Y" marks the rule as shared across hosts. |
actions | yes | Must be an empty list [] — action-based dispatch is not supported. |
Key fields in each endpoint rule entry:
| Field | Description |
|---|---|
req-acc | List of rule IDs evaluated before calling the downstream tool. |
res-fil | List of rule IDs evaluated after the downstream response. |
permission | Arbitrary key/value map injected into the CEL context as permission. Keeps rule bodies generic and reusable. |
The file name is rule.yml. The loader also accepts rule.yaml.
Rule Context
For every MCP tool call the gateway builds a CEL evaluation context containing the following top-level variables:
| Variable | Type | Description |
|---|---|---|
auditInfo | map | Decoded JWT claims and correlation metadata. |
permission | map | The per-endpoint permission object from endpointRules. |
headers | map | Normalised (lowercased) HTTP request headers. |
toolName | string | MCP tool name from the tools/call request. |
toolArguments | map | Tool call arguments from the tools/call request. |
endpoint | string | Endpoint identifier, e.g. weather@call. |
correlationId | string | Correlation ID when one is present. |
JWT claims are nested under auditInfo.subject_claims.ClaimsMap. The gateway
normalises common fields automatically:
| CEL path | JWT source | Type |
|---|---|---|
auditInfo.subject_claims.ClaimsMap.scp | scp | list<string> |
auditInfo.subject_claims.ClaimsMap.roles | roles | list<string> |
auditInfo.subject_claims.ClaimsMap.positions | positions | list<string> |
auditInfo.subject_claims.ClaimsMap.groups | groups | list<string> |
auditInfo.subject_claims.ClaimsMap.attributes | attributes | map<string,string> |
auditInfo.subject_claims.ClaimsMap.sub | sub | string |
auditInfo.subject_claims.ClaimsMap.client_id | client_id / azp | string |
auditInfo.subject_claims.ClaimsMap.uid | user ID injected by gateway | string |
auditInfo.subject_claims.ClaimsMap.role | role (singular) | string |
For a full reference with worked examples for every claim type, see Request Access Control Rules.
Security Profile
Rules must declare conditionSecurityProfile: strict. The strict profile:
- Restricts available functions and macros to a safe, well-known subset.
- Prevents access to undeclared variables, guarding against injection.
- Causes the expression to return an error (treated as
denied) if it references a missing variable rather than silently returningfalse.
Always guard list membership with an in check before accessing a key.
Claims absent from the token will be missing from ClaimsMap, and an
unguarded access will be denied:
# WRONG — will error if 'scp' is not in the token
permission.groups in auditInfo.subject_claims.ClaimsMap.scp
# CORRECT
'scp' in auditInfo.subject_claims.ClaimsMap
&& permission.groups in auditInfo.subject_claims.ClaimsMap.scp
Common CEL Patterns
Scope (scp) — OAuth 2.0 access token
expression: |
'scp' in auditInfo.subject_claims.ClaimsMap
&& 'groups' in permission
&& permission.groups in auditInfo.subject_claims.ClaimsMap.scp
Role
expression: |
'roles' in auditInfo.subject_claims.ClaimsMap
&& 'role' in permission
&& permission.role in auditInfo.subject_claims.ClaimsMap.roles
Position
expression: |
'positions' in auditInfo.subject_claims.ClaimsMap
&& 'position' in permission
&& permission.position in auditInfo.subject_claims.ClaimsMap.positions
Attribute
expression: |
'attributes' in auditInfo.subject_claims.ClaimsMap
&& 'attributeKey' in permission
&& 'attributeValue' in permission
&& permission.attributeKey in auditInfo.subject_claims.ClaimsMap.attributes
&& auditInfo.subject_claims.ClaimsMap.attributes[permission.attributeKey] == permission.attributeValue
AND — require both a scope group and a role
expression: |
'scp' in auditInfo.subject_claims.ClaimsMap
&& 'roles' in auditInfo.subject_claims.ClaimsMap
&& permission.groups in auditInfo.subject_claims.ClaimsMap.scp
&& permission.role in auditInfo.subject_claims.ClaimsMap.roles
For OR logic and a fully generic multi-claim rule, see Request Access Control Rules.
Endpoint Matching
When the gateway looks up the rule list for an incoming request it checks:
- Exact endpoint key — e.g.
weather@call. - Path templates — e.g.
accounts/{id}@get. - Parent path — e.g.
accounts@getmatchesaccounts/123@get.
For MCP tools, always set endpoint explicitly in mcp-router.yml so the
policy key remains stable even if the downstream path changes.
Reload Behavior
light-gateway supports live reload for MCP and access-control config:
- Reloading
mcp-router.ymlrebuilds the MCP router runtime. - Reloading
access-control.ymlorrule.ymlrebuilds the MCP and WebSocket policy runtimes.
This matches the product model where light-portal manages configuration and
config-server delivers the resolved files.
Operational Notes
- If
access-control.ymlis missing, MCP tools are allowed unless another handler blocks the request. - If
access-control.ymlis enabled anddefaultDeny: true, a tool call with no matchingreq-accendpoint rule is denied. - If
access-control.ymlis enabled anddefaultInclude: false, ares-filrow filter with no matching caller claim returns no rows rather than all rows. - If the
securityhandler does not run beforemcp, JWT claims are absent and CEL expressions that referenceauditInfowill deny. - Rule execution is local to the gateway. No database call is made per request.
x-maskandx-mask-patternin MCP toolinputSchemaare applied before the downstream call.x-tokenizeis reserved for the tokenization service integration.
Verification
Useful checks:
cargo tree -p light-gateway -i light-rule
cargo test -p light-pingora access_control
cargo test -p light-gateway gateway_loads_mcp_router_when_mcp_handler_is_active
The first command verifies the binary linkage. The test commands verify the MCP access-control path, default deny behavior, CEL-based allow behavior, and gateway MCP runtime loading.
See Also
- Request Access Control Rules — full reference for
CEL
req-accrules: JWT claim paths,permissionobject structure, worked examples for scopes, roles, positions, attributes, subject, client ID, combined conditions, and a generic dynamic rule pattern.
LLM Gateway Design
Status
Implemented through REL-1 and the request-scoped PII profile. Production enablement remains fail-closed pending committed PERF-3/PERF-4 measurements, live Python/TypeScript SDK evidence against both provider formats, and live canary/rollback evidence.
The checked-in implementation gates are evidence validators, not substitutes
for those external runs. llm-router.enabled remains false, release
canaryAllowed remains false, and PII promotion remains independently gated
by functional, security, durability, and performance lanes.
Implementation And Qualification Status
| Contract | Current implementation | Remaining promotion evidence |
|---|---|---|
| LF-1 through LF-6B | Deterministic baselines, canonical provider contract, OpenAI/Anthropic codecs, compiled single-attempt runtime, accounting/circuits/replay, buffered HTTP, and early SSE are implemented. | PERF-1 measurements remain an external architecture-checkpoint input. |
| PDB-1, LP-1, GC-1/GQ-1, PV-1 | Host-scoped schema, event persistence, commands/queries, atomic publication, governed-alias UI, defensive secret redaction, and component-level control-plane tests are implemented. | Operational Portal deployment and publication approval. |
| DIST-1, LF-7, LA-1 | Monotonic projection, two-replica convergence contracts, secret rotation, retained runtime state, and agent alias isolation are implemented. Production deployment resources require a complete current conformance result. | Captured provider evidence for the exact physical deployments. |
| LF-8, LF-9, PERF-2 | Durable WAL/sink, ownership lock, replay/reclamation, accounting-aware streaming, deadlines, and protocol checks are implemented. | Declared external performance captures. |
| PERF-3, OBS-1, SEC-1, REL-1 | Qualification contracts, bounded telemetry, SSRF/body-access controls, rollout stages, and monotonic rollback are implemented. Release evidence is bound to the current commit and critical-source digests. | Five-run PERF-3, live SDK/provider smoke, canary, and rollback-drill evidence. |
| PII-1, PERF-4 | Authenticated request-scoped placeholders, exact fragmented-stream recovery, typed promotion identity, vault boundary, and four independent promotion lanes are implemented. | Functional, security, durability, and PERF-4 lane evidence; session/host scope stays unavailable until the durable-vault lane passes. |
Production projection defaults requiredConformanceProvenance to
captured_sanitized. Synthetic corpus results remain useful regression
evidence but cannot make a production deployment eligible. A PASS Portal
deployment stores the complete canonical conformanceResult; compact state or
capability flags alone are quarantined or rejected.
The live SDK closure harness pins the official OpenAI Python and TypeScript
packages and exercises /v1/models, buffered chat, streaming with usage, and
tool calls against both an OpenAI deployment and an Anthropic-backed governed
alias. Its sanitized evidence is bound to the release commit, projection
digest, and both conformance digests.
Decision Summary
Add an LLM inference handler to light-gateway with these initial decisions:
- Expose an OpenAI-compatible client API. Implement
GET /v1/modelsandPOST /v1/chat/completionsfirst, including Server-Sent Events (SSE) streaming. AddPOST /v1/responsesafter the provider abstraction can preserve its richer content and event model. - Treat the request
modelas a public, governed model alias. Clients do not select provider credentials, provider base URLs, or physical deployments. - Keep the wire protocol separate from the provider abstraction. The gateway translates OpenAI-compatible requests into a provider-neutral internal representation and translates normalized provider results back to the selected public protocol.
- Reuse
crates/model-providerimplementations, but do not expose the currentProvidertrait directly as the HTTP contract. It needs typed errors, streaming events, structured content blocks, cancellation, richer request options, and per-model capabilities before it is a production gateway boundary. - Reuse the existing Light handler chain for correlation, authentication, authorization, request rate limits, metrics, and common traffic policy. Add LLM-specific routing, token and cost budgets, provider health, and usage accounting in a dedicated runtime.
- Keep LLM inference and MCP tool execution as separate protocol boundaries. The LLM gateway can accept tool definitions and return tool calls, but the client agent remains responsible for executing those calls through the MCP router and returning tool results to the model.
- Keep configuration and administration in the Light control plane. Do not add a second gateway-specific administration UI or public mutation API in the first implementation.
- Store the host-scoped model catalog, deployments, public aliases, routing policy, capability snapshots, and pricing metadata in the Light Portal control plane. Agent definitions reference a governed alias or model policy; they do not own provider credentials or select a physical provider model.
- Separate control-plane, inference-record, and reversible-PII storage. Portal PostgreSQL remains authoritative for configuration and canonical agent-domain events. A dedicated local or regional audit store owns gateway inference records, while a separately credentialed PII vault is used only when token mappings must survive the request.
- Keep request-scoped PII mappings in memory by default. Use a local bounded
WAL/spool for audit delivery, not as the authoritative audit corpus or a
replica-local long-lived PII vault. Distinguish
bounded-asyncadmission fromlocal-durablepre-dispatch commit; never claim that queue capacity is crash durability. - Represent the client format, logical operation, and selected upstream format separately. Preserve unknown fields in a bounded compatibility envelope for same-format forwarding, and upgrade to fully typed canonical content only when policy mutation or cross-provider conversion requires it.
- Pre-bind provider dispatch, resolved alias policy, eligible priority groups, pricing references, and content-access requirements into an immutable runtime snapshot. Static enum and preconstructed dynamic dispatch are both acceptable; the benchmark decides. A request must not repeatedly lock configuration stores, merge policy layers, construct a provider client, or look up provider implementations by string.
- Publish one small atomic root containing structurally shared routing,
provider, policy, and pricing sub-snapshots. A pricing-only or single-alias
update reuses unchanged
Arcgraphs, while one root load still gives each request a generation-consistent view. - Treat every upstream credential as an authorized quota and billing principal. Credentials in one deployment set are lifecycle versions within the same quota group; separately approved accounts/capacity are separate deployments. The gateway must not rotate keys to evade a provider’s RPM/TPM, account, contract, or abuse limits.
- Make performance a release contract, not an implementation claim. The gateway must meet an absolute latency and capacity SLO and must also equal or outperform a pinned Bifrost build under the same open-loop workload, hardware limits, protocol, payloads, provider mock, and enabled features.
Context
light-gateway already has most of the surrounding gateway capabilities:
- Pingora HTTP and HTTPS listeners and proxy transport.
- Ordered handler chains configured by
handler.yml. - JWT, API key, basic, unified-security, and agent-delegation authentication.
- Access control, request rate limits, correlation, metrics, headers, and CORS.
- MCP request handling through the
mcpapplication handler. - Browser-to-agent WebSocket routing through the
websockettraffic handler. - Config registration, config-server bootstrap, atomic
ConfigManagerswaps, and reloadable modules.
crates/model-provider already contains provider clients for OpenAI, Azure
OpenAI, Anthropic, Bedrock, Gemini, GLM, Ollama, OpenRouter, Telnyx, and generic
OpenAI-compatible endpoints. It also contains account- or CLI-oriented clients
such as Codex, Copilot, Claude Code, Gemini CLI, and Kilo CLI, plus two wrapper
providers:
RouterProviderresolves ahint:<name>to a configured provider and model.ReliableProviderperforms retries and walks provider/model fallback chains.
The current common types are intentionally small and agent-oriented:
ChatMessagehas a string role and string content.ChatRequesthas messages and optional tools.ChatResponseis buffered text, tool calls, usage, and optional reasoning content.ProviderCapabilitiescontains only native tool calling, vision, and prompt caching flags.- Provider errors are returned as
anyhow::Error.
That is enough for the current light-agent and light-workflow call paths,
but it cannot faithfully implement a public LLM gateway. For example, it has no
common incremental stream, typed provider status and Retry-After, structured
multimodal content blocks, response-format contract, cancellation signal, or
per-operation capability declaration.
Three open source gateways provide useful feature signals:
- Bifrost emphasizes an OpenAI-compatible API, provider-native compatibility adapters, retry and fallback, weighted routing, virtual-key governance, hierarchical budgets, semantic caching, plugins, observability, and MCP integration.
- LiteLLM exposes OpenAI-format and native endpoints across many providers and adds proxy authentication, virtual keys, spend tracking, rate limits, routing, fallback, caching, guardrails, and logging.
- agentgateway is a Rust
multi-protocol gateway with purpose-built local and xDS configuration, an LLM
model router, OpenAI and provider-native formats, typed provider conversion,
virtual models, health-aware priority failover, guardrails, token/cost
telemetry, and an atomically replaceable pricing catalog. The implementation
review in this document is based on commit
857281d.
The Light design should adopt the durable product capabilities without copying any reference project’s control plane. Light already has a portal, config server, controller, security handlers, access control, and an MCP router.
Goals
- Give applications and agents one stable base URL and one common API across supported model providers.
- Let existing OpenAI SDK users migrate by changing the base URL and client credential rather than rewriting request and response handling.
- Support buffered and streaming chat, tool calling, structured output, and supported multimodal input without losing provider semantics silently.
- Route public model aliases to one or more physical provider deployments.
- Let each organization/host register only the models and deployments it is authorized to use, and manage their routing metadata through GenAI Admin.
- Provide retry, fallback, load balancing, circuit breaking, health-aware routing, and cancellation with well-defined streaming behavior.
- Enforce model access, data-boundary constraints, token limits, cost budgets, concurrency limits, and request rate limits per authenticated identity.
- Record normalized usage, cost, latency, time to first token, route decisions, retry/fallback activity, and policy outcomes.
- Deliver audit records without adding synchronous Portal-database work to the normal inference path, and support governed content capture for later audit, evaluation, and curated dataset export.
- Tokenize policy-selected PII before cloud-provider dispatch and recover only exact authorized placeholders before returning the model response.
- Protect provider credentials and prevent clients from choosing arbitrary upstream URLs or passing provider secrets through the gateway.
- Reload provider, alias, route, and policy snapshots atomically without interrupting in-flight requests.
- Make provider conformance measurable so an alias is offered only when every eligible target can satisfy its declared capabilities.
- Sustain Bifrost-class request rates without entering a queueing collapse: keep the hot path typed and allocation-conscious, reuse upstream clients, shed excess load promptly, and verify comparative throughput and tail latency before release.
Non-Goals
- Do not run an autonomous agent loop in the LLM gateway. It does not execute model-returned tool calls or decide when an agent task is complete.
- Do not replace the MCP router or merge MCP JSON-RPC with the LLM HTTP API.
- Do not proxy arbitrary client-supplied provider base URLs, API keys, or cloud credentials.
- Do not expose every provider-specific option through the common API. Provider-native compatibility endpoints can be added deliberately when a real client need justifies their maintenance cost.
- Do not support fine-tuning, training, file storage, assistants, or durable conversation state in the first implementation.
- Do not enable account- or CLI-oriented providers in a shared gateway until their credential isolation, licensing, concurrency, and multi-tenant behavior have passed a separate security review.
- Do not log prompts, completions, images, tool arguments, or reasoning content by default.
- Do not use the Portal database, a gateway replica’s embedded database, or the inference content store as the production reversible-PII security boundary.
- Do not treat operational audit content as an automatically approved training dataset.
- Do not promise identical model output after fallback. Fallback preserves the API and required capabilities, not model behavior.
- Do not advertise a fixed multiple such as
40xor50x. Those ratios depend on the benchmark definition and can be dominated by overload queueing. Report gateway-added latency, sustainable throughput, success rate, and resource use from a reproducible benchmark instead.
Reference Feature Comparison
The comparison is a requirements input, not a compatibility promise.
| Capability | Bifrost signal | LiteLLM signal | agentgateway signal | Light direction |
|---|---|---|---|---|
| Common inference API | OpenAI-compatible API plus provider SDK adapters. | OpenAI input/output format plus native endpoints. | OpenAI Completions/Responses plus Anthropic Messages, embeddings, rerank, realtime, token count, detect, and opaque routes. | OpenAI-compatible API first; preserve source-format identity so selected native adapters can be added without flattening through Chat Completions. |
| Provider abstraction | Many hosted and local providers. | Broad provider and endpoint coverage. | Rust enum dispatch with typed request/response conversions and provider-format selection. | Reuse model-provider, gated by per-operation conformance tests; pre-bind provider execution in the runtime snapshot and let allocation benchmarks choose enum or dynamic dispatch. |
| Routing | Provider/model/key routing and weighted strategies. | Deployment router and load-balancing strategies. | Public/internal concrete models plus weighted, conditional, and health-aware priority-failover virtual models. | Public alias to eligible deployment targets with weighted, priority-, health-, policy-, and capability-aware selection. |
| Reliability | Retries, key rotation, and sequential fallbacks. | Retries, cooldowns, and cross-deployment fallback. | Generic HTTP retry integrates with endpoint outlier eviction so the next attempt can move to the next priority group. | One typed attempt coordinator, Retry-After, health/outlier state, and no fallback after visible stream output. |
| Tenant credentials | Virtual keys. | Virtual keys and proxy keys. | General gateway authentication and authorization policies apply to LLM routes/models. | Reuse Light authentication; map the authenticated client, user, agent, and host to an LLM policy. |
| Cost governance | Hierarchical budgets and rate limits. | Spend tracking and budgets by several scopes. | Token-aware rate limits and an ArcSwap pricing catalog with detailed usage classes and source overlays. | Atomic token/cost reservation and usage reconciliation by configured Light policy scopes; publish pricing as an independent versioned projection. |
| Caching | Exact/provider and semantic caching. | Configurable response caches. | Provider prompt-caching policy and cache-token accounting. | Exact cache later; semantic cache is opt-in and tenant/policy isolated. |
| Guardrails | Plugin-based request and response controls. | Per-project guardrails and callbacks. | Local regex/PII masking, external safety services, and bounded-window SSE/realtime response blocking. | Ordered local/remote hooks, reversible PII profiles, and explicit buffered versus bounded-window streaming semantics. |
| Observability | Metrics, tracing, and request logging. | Logging callbacks, usage, cost, and latency. | CEL-selectable LLM attributes, normalized usage/cost, and streaming completion accounting. | Existing correlation/metrics plus bounded-cardinality events and a dedicated durable audit pipeline. |
| Performance architecture | Compiled Go, fasthttp, typed provider codecs, object pools, and per-provider workers. | FastAPI/ASGI with a generic Python router, SDK dispatch, and callback pipeline. | Compiled Rust, typed/minimally parsed codecs, reusable clients, bounded bodies, endpoint sets, and atomic pricing snapshots; some configuration reads and policy merging remain request-time work. | Rust/Pingora, compatibility fast path, pre-bound provider dispatch, one structurally shared request snapshot, bounded admission, and benchmark-enforced parity or better. |
| MCP | MCP gateway and tool filtering. | MCP support and model-tool integration. | MCP, A2A, HTTP, and LLM backends share the gateway policy/runtime. | Keep the existing MCP router authoritative for tool discovery and execution while sharing identity, policy, and telemetry primitives. |
| Administration | Built-in configuration and monitoring UI. | Admin UI and APIs. | Human-friendly watched local config or granular purpose-built xDS resources mapped to a shared IR. | Use Light Portal, config server, and controller; publish granular resources but compile a complete request-ready runtime snapshot. |
agentgateway Architecture Review
The reviewed agentgateway path is not merely “Rust instead of Python.” Its architecture makes several explicit choices that reduce compatibility work and keep most LLM processing inside compiled code:
local YAML/JSON or purpose-built xDS resources
-> shared gateway IR and targeted policies
-> HTTP route -> LLM model router
-> concrete model or weighted/conditional/failover virtual model
-> provider endpoint set and merged backend policy
-> client-format parser -> provider-format renderer
-> reusable upstream transport
-> provider stream parser -> client-format stream renderer
-> optional bounded-window guard -> client
The following decisions are worth adopting or deliberately refining:
| agentgateway choice | Why it is useful | Light decision |
|---|---|---|
Separate endpoint RouteType, client InputFormat, and provider ChatFormat. | A Messages request may go to a Completions upstream while the gateway still knows which response contract it owes the client. | Define ClientFormat, Operation, and ProviderFormat separately from day one. Never infer the response contract from the selected provider route. |
Parse operated fields and preserve unknown fields in a flattened rest; upgrade to fully typed forms only for conversion. | Same-format compatibility survives provider/API additions without requiring an immediate full schema update. | Use a bounded compatibility envelope for same-format routes. Validate operated fields strictly; allow unknown fields only under an alias/provider allowlist and never forward an unknown extension across formats blindly. |
| Public/internal concrete models and virtual models with weighted, conditional, or priority-failover routing. | Client aliases stay stable while internal targets and rollout policy change. Internal targets need not be directly invokable. | Keep public aliases separate from internal deployments. Add priority groups and metadata-conditional routing after the ordered MVP, with expressions compiled at publication and evaluated only over an allowlisted, sanitized context. |
| Provider selection uses endpoint sets with health, latency, pending-work scoring, priority buckets, and outlier eviction. | A retry can reselect after a bad target is ejected instead of repeatedly hitting the same endpoint. | Treat retry and failover as one attempt coordinator. Feed typed outcomes into per-deployment health and reselect from the same immutable eligible plan, preserving capability and residency constraints. |
| Built-in providers use exhaustive enum dispatch and typed conversion code. | Provider choice is resolved in compiled Rust without constructing a dynamic SDK/client per request. | Preserve the pre-bound compiled path but benchmark sealed-enum and preconstructed trait-object implementations. Never do string-to-provider registry lookup or client construction on each attempt. |
| Streaming is translated to the client-visible SSE format before response guards inspect text. Held semantic frames are released in bounded windows with overlap. | Binary or provider-native streams are not mistaken for OpenAI SSE, and patterns spanning chunks can be detected before held frames are released. | Apply provider decoding first, then exact PII-token recovery and post-policy in a documented order over semantic events. Buffer complete frames plus a bounded overlap; a whole-response rule forces buffered mode. |
| Prompt/completion attributes are materialized only when a CEL expression asks for them, and the raw LLM request exists only during the LLM-policy phase. | Large, sensitive content does not become a universal request-context cost or remain available to later logging accidentally. | Make content access lazy, phase-scoped, and policy-authorized. Remove raw content from the general handler/log context after the content-policy phase; later stages receive normalized metadata or explicit encrypted references. |
Pricing sources merge into a validated ArcSwap snapshot and retain the last valid catalog on reload failure. | Pricing reads are wait-free and a broken file does not turn known prices into zero. | Publish pricing as an independently versioned immutable projection, capture its version per attempt, support explicit source precedence, and keep unknown pricing fail-closed for hard budgets. External reference catalogs never authorize a host or activate a deployment. |
| Local/xDS resources closely mirror user resources, while policies remain separate and merge at runtime. | Small control-plane changes avoid large route-list fan-out and keep control-plane translation mechanical. | Preserve granular Portal/config events and delta publication, but compile affected alias/route policy combinations before activation. The LLM request path loads one request-ready snapshot and performs no shared-store policy merge. |
There are also boundaries Light should not copy:
- The reviewed gateway store uses shared
RwLock-protected bind/discovery stores, and the HTTP/LLM path reacquires bind reads and clones/merges some policy layers during request processing. Light should spend the additional reload-time work to publish pre-resolved LLM plans behindArcSwap. - agentgateway may inject
stream_options.include_usage=truewhen the client omitted it, which adds a client-visible final SSE event. Light may request upstream usage internally, but it must remember the original client contract and suppress an injected usage frame unless the client requested it. - Its bounded-window streaming guardrail explicitly cannot provide full-stream accuracy, cannot retract earlier windows, and does not support streaming masking. Light policy publication must reject an incompatible streaming/policy combination or force buffering; reversible exact-token recovery is a separate bounded streaming transform, not ordinary masking.
- Its local PII recognizers mask or reject content; they do not provide the separately scoped reversible-token vault required by this design. Its telemetry path also does not replace Light’s logical-request/physical-attempt durable audit ledger and governed dataset export.
- A provider response parsing helper in the reviewed source logs up to the first 1,024 response bytes on a parse failure. Light must never put raw provider error/response bodies into ordinary logs. Record only bounded error classification, byte length, content type, and a keyed digest unless an explicitly authorized encrypted-content policy captures the body.
These differences are opportunities to be faster as well as safer than the reference: agentgateway validates the value of typed Rust codecs, static provider dispatch, endpoint priority groups, and atomic catalog replacement; Light can combine those ideas with a stricter one-published-root request path and no request-time configuration merging.
Common Client API
API Choice
Use the OpenAI API shape as the public compatibility profile.
| Option | Strength | Limitation | Decision |
|---|---|---|---|
| OpenAI Chat Completions | Widest existing SDK and agent-framework compatibility; maps closely to current provider clients; supports SSE and tool calls. | Its message model is less expressive than newer event/item APIs. | MVP. |
| OpenAI Responses | Better fit for reasoning models, structured content items, and richer streaming events. | Requires a significantly richer internal contract and has less uniform third-party coverage. | Phase 2, built on the same canonical internal types. |
| Provider-native APIs | Maximum fidelity for one provider and drop-in support for provider SDKs. | Multiplies codecs, tests, and long-term compatibility obligations. | Add selected adapters later; not the canonical client API. |
| Light-specific inference API | Full control over versioning and semantics. | Requires new SDKs and creates avoidable client migration work. | Do not use as the primary external API. |
The OpenAI-compatible contract is a compatibility profile, not a claim that every provider supports every OpenAI option. Capability checks and explicit errors are part of the contract.
Internally, keep three dimensions distinct:
ClientFormatis the request and response contract owed to the caller, for example OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages.Operationis the semantic action, for example chat, responses, embeddings, rerank, token count, or realtime.ProviderFormatis the selected upstream wire contract, for example OpenAI Completions, Anthropic Messages, Bedrock Converse, or a provider-native embedding route.
A provider selection can change ProviderFormat; it never changes
ClientFormat. This prevents a fallback or cross-provider conversion from
accidentally returning the upstream provider’s shape to the client.
Endpoint Roadmap
| Endpoint | Priority | Notes |
|---|---|---|
GET /v1/models | MVP | Return only public aliases authorized for the caller. Do not enumerate raw provider deployments. |
POST /v1/chat/completions | MVP | Buffered and SSE streaming; text, tool calls, supported image input, and structured output where the alias declares support. |
POST /v1/responses | Phase 2 | Use native response items and events rather than flattening through Chat Completions. |
POST /v1/embeddings | Phase 2 | Add only after an embedding operation exists in the canonical provider trait and pricing model. |
POST /v1/moderations | Phase 2 or policy service | Decide whether this is a provider operation or a Light policy endpoint before implementation. |
| Images, audio, rerank, batches, and files | Later | Each needs its own capability, size, cost, storage, streaming, and retention contract. |
| Realtime WebSocket/WebRTC | Later | This is provider realtime inference, not the existing UI-to-agent WebSocket router. Implement it as a separate protocol handler. |
A later provider-native compatibility adapter may use one of three explicit processing modes:
normalized: strict typed validation, full policy support, and conversion to any conforming provider format.detect: same-format forwarding with shallow LLM metadata and usage extraction. It is eligible only for policies whose required controls can be enforced without full content normalization.opaque: bounded HTTP/WebSocket forwarding with no LLM interpretation. It cannot satisfy token, content-guardrail, reversible-PII, or normalized-audit requirements. If enabled for a separately governed compatibility route, it still enforces identity, destination allowlists, request/response bytes, request rate, concurrency, duration, egress, and a pessimistic fixed per-request cost reservation or externally reconciled account-spend ceiling. A policy requiring authoritative token or exact per-call cost accounting cannot select it.
The public OpenAI-compatible endpoints use normalized. detect and opaque
are explicitly configured compatibility tools, never automatic fallbacks when
normalization fails.
Opaque traffic is therefore not free or unlimited; its governance unit is a
bounded request rather than a token. The audit record marks token usage and
per-call realized cost as unknown, records the reserved fixed envelope and
byte counts, and reconciles provider-account spend asynchronously when billing
data is available. If no conservative envelope or authoritative account cap is
configured, publication rejects the route.
Authentication And Headers
No Light-specific header is required for a normal SDK call.
Authorization: Bearer <credential>carries a Light-issued API key, JWT, or agent-delegation credential accepted by the configured handler chain. It is never a provider API key.Content-Type: application/jsonis required for JSON request endpoints.X-Correlation-IdandX-Traceability-Iduse the existing Light correlation contract.X-Light-Session-Idis an optional routing hint for session stickiness. It must be bounded, treated as untrusted input, and scoped by authenticated principal so two tenants cannot collide.Idempotency-Keycan enable request deduplication where the selected operation and storage policy support it. It cannot guarantee that a provider did not bill a timed-out upstream attempt.x-request-idshould be returned for OpenAI SDK diagnostics and linked to the Light correlation and trace identifiers in server-side telemetry.- Provider name, physical model, base URL, key ID, raw error body, and internal policy details are not returned by default. Authorized diagnostic tooling can retrieve them from audit events.
The OpenAI user field is optional attribution metadata. It does not establish
identity and cannot override the authenticated principal.
Public Model Names
The request model is a logical alias such as chat-fast-v1,
reasoning-standard-v1, or private-code-v2.
An alias defines:
- Allowed operations and request features.
- Maximum input and output sizes.
- Data classification and residency requirements.
- Eligible provider deployments and physical model identifiers.
- Routing, retry, fallback, timeout, and budget policy.
- Pricing policy and capability snapshot version.
- Deprecation and replacement metadata.
Provider-prefixed names such as openai/gpt-x can be convenient for local
development, but public production policy should disable them. Otherwise the
client can bypass alias-level routing, residency, and lifecycle controls.
GET /v1/models returns only aliases visible to the caller. A separate
authenticated control-plane view can show target deployments and detailed
capabilities.
Chat Completions Compatibility Profile
The MVP should support these fields when the selected alias declares the required capability:
modelmessageswith text content and supportedimage_urlcontent partstemperatureandtop_pmax_tokensandmax_completion_tokens, normalized to one internal output limit with a conflict error if both disagreestopstreamandstream_options.include_usagetools,tool_choice, andparallel_tool_callsresponse_formatfor text, JSON object, and JSON Schema where supporteduserand bounded metadata for attribution
The gateway must not silently drop a non-default option. The default
unsupportedParameterPolicy is reject. A per-alias allowlist can permit
provider-specific pass-through fields only when every eligible route handles
them consistently. Unknown null or SDK-default fields can be ignored if the
compatibility profile explicitly documents them.
Reasoning summaries may be exposed only through a documented public field or Responses event. Hidden chain-of-thought or raw provider reasoning content must not be logged or returned merely because a provider client captured it.
Example:
curl https://gateway.example.com/v1/chat/completions \
-H 'Authorization: Bearer <light-credential>' \
-H 'Content-Type: application/json' \
-H 'X-Correlation-Id: example-request-1' \
-d '{
"model": "chat-fast-v1",
"messages": [
{"role": "user", "content": "Summarize the attached incident."}
],
"temperature": 0.2,
"stream": true
}'
Streaming Contract
Chat Completions streaming uses text/event-stream, OpenAI-compatible
chat.completion.chunk data frames, and a terminal data: [DONE] frame.
Streaming changes reliability semantics:
- Before the gateway emits the first semantic output event, it may retry or choose an eligible fallback according to policy.
- After any text, tool-call argument, or other semantic output is visible to the client, the gateway must not retry or switch providers. Doing so can duplicate text, corrupt incremental JSON arguments, or create a second tool call.
- A failure after streaming starts emits a sanitized error event when the
compatibility profile permits it, then closes the stream without
[DONE]. - A downstream disconnect cancels the provider request promptly and records the final known usage. Cancellation is best effort because a provider can continue billing work already accepted upstream.
- Time to first token, stream duration, client cancellation, and upstream cancellation outcome are recorded separately.
- The gateway preserves the client’s
stream_options.include_usagechoice. A provider adapter may request usage upstream for accounting, but an internally injected usage event is removed from the client stream when the public contract did not request it.
The current MCP stream writer demonstrates that Pingora can write incremental frames, but LLM SSE framing, usage events, disconnect cancellation, and post-stream accounting need their own implementation and tests.
Error Contract
Map typed provider and gateway errors to the OpenAI error envelope:
{
"error": {
"message": "The selected model is temporarily unavailable.",
"type": "server_error",
"param": null,
"code": "model_unavailable"
}
}
The gateway should normalize at least these categories:
| Category | Typical HTTP status | Retry behavior |
|---|---|---|
| Invalid request or unsupported parameter | 400 | Never retry. |
| Authentication failure | 401 | Never retry. |
| Model or policy access denied | 403 | Never retry or reveal hidden aliases. |
| Unknown authorized model alias | 404 | Never retry. |
| Request or token limit exceeded | 413 or 422 | Never retry without changing the request. |
| Request/token/cost rate limit | 429 | Honor Retry-After; a different target is eligible only when policy permits. |
| Provider timeout | 504 | Retry or fallback only before visible stream output. |
| Provider unavailable or circuit open | 502 or 503 | Retry/fallback only to a capability-equivalent target. |
| Internal gateway failure | 500 | Do not expose raw provider or configuration details. |
Proposed Architecture
Client application or agent
-> Pingora listener
-> handler chain
correlation -> CORS -> unified-security -> limit -> access-control -> llm
-> OpenAI-compatible HTTP codec
-> client format + operation + bounded compatibility envelope
-> authenticated LLM request context
-> alias and policy resolver
-> token/cost/concurrency reservation
-> cache lookup when eligible
-> capability, residency, health, and budget target filter
-> route selection -> retry/fallback coordinator
-> optional request-scoped PII tokenization
-> statically selected model-provider adapter -> provider-format API
-> provider decode -> client-format semantic events, usage, and typed errors
-> exact-token PII recovery, policy post-processing, and quota reconciliation
-> buffered JSON or SSE response
-> metrics and trace
-> bounded audit queue -> local spool when needed -> dedicated audit store
The control plane publishes immutable, versioned runtime snapshots to the gateway. Portal, config server, secret manager, audit database, and PII vault lookups are not part of the normal alias-resolution or routing path.
The llm application handler terminates the HTTP request inside the gateway;
it is not an ordinary upstream proxy. Consequently, body-dependent security
and transformation stages must execute inside the application-handler flow
before provider dispatch. Merely listing access-control, tokenize, or
another body handler earlier in handler.yml does not prove that Pingora’s
later proxy body filters will run.
Performance Architecture And Release Contract
Performance is part of the public reliability contract. A fast provider does not compensate for a gateway that consumes excessive CPU, accumulates an internal backlog, or delays stream chunks. Conversely, a microbenchmark that excludes JSON, middleware, or response processing does not represent what a client experiences.
Interpreting The Bifrost Reference
The published Bifrost comparison contains two distinct results:
- At the advertised 500-RPS load, Bifrost reported about
9.5xLiteLLM’s completed throughput, while the reported P50 and P99 latency ratios grew to approximately48xand54xafter LiteLLM saturated and requests queued. - In a separate test with a 60-ms mock provider, the end-to-end medians were
60.99 ms and 100 ms, a
1.64xdifference. The40xclaim comes from subtracting the assumed 60-ms mock time and comparing 0.99 ms with 40 ms. - Bifrost’s 5,000-RPS internal-overhead figures exclude at least the upstream call and some codec work. They are useful implementation signals, but they are not directly comparable to an end-to-end proxy latency percentile.
The Light target is therefore not “be 50 times faster than LiteLLM.” The target is to remain below the saturation knee, equal or exceed Bifrost’s sustainable throughput, and equal or improve its gateway-added P50, P95, and P99 latency in a controlled comparison. Both the absolute and comparative gates below must pass.
Release Performance Gates
The first implementation establishes a checked-in benchmark manifest with the exact Light commit, Bifrost image digest or commit, load generator version, kernel and CPU architecture, instance limits, configuration, payload corpus, and mock-provider build. A result without those inputs is diagnostic only and cannot satisfy a release gate.
| Gate | Required result |
|---|---|
| Comparative non-inferiority | On identical hardware and feature-equivalent profiles, Light sustainable throughput must be at least Bifrost’s, and Light gateway-added P50, P95, and P99 must be no higher. Compare five or more steady-state runs and require the 95% confidence interval to remain inside a 5% non-inferiority margin. The engineering target is at least 10% better throughput or P99, not merely a statistical tie. |
| Rust architecture reference | Run the same compatible subset against a pinned agentgateway commit. Report results even though Bifrost remains the MVP release comparator. Any regression against agentgateway in routing-only, same-format, or streaming profiles requires an explained architectural cause and an accepted optimization plan. |
| 500-RPS small-payload baseline | On 2 vCPU and 4 GiB RAM with keep-alive and a 60-ms mock provider, admit and complete the full 500-RPS offered load with 100% success, add no more than 1 ms at P50 and 5 ms at P99, and show no growing internal queue during the steady-state window. |
| 5,000-RPS high-throughput profile | On 4 vCPU and 16 GiB RAM with a local mock provider and small buffered responses, admit and complete the full 5,000-RPS offered load with 100% success, keep P99 admission wait below 1 ms, and meet the comparative Bifrost latency and throughput gate without unbounded memory growth. |
| Production handler profile | Repeat the comparison with correlation, cached authentication, authorization, request limits, metrics, routing, usage accounting, and metadata-only bounded-async audit enabled. No feature may be disabled only for Light if its equivalent remains enabled for Bifrost. |
| Durable-audit profile | Run local-durable metadata audit on declared persistent storage and report commit-batch size, fdatasync duration, commit-wait P50/P95/P99, throughput, incomplete recovery, and overload behavior. It must meet its configured commit timeout with no unaudited dispatch; do not average it into or use it to weaken the normal 500/5,000-RPS gates. |
| Streaming profile | At matched concurrent streams and chunk cadence, Light time-to-first-byte overhead and P99 per-chunk processing delay must be no worse than Bifrost. Slow consumers must remain bounded and cancellation must release permits and upstream work promptly. |
| Overload profile | Increase fixed offered load beyond capacity. Admitted-request latency must remain bounded; excess requests must receive a prompt 429 or 503 instead of waiting in an unbounded queue. The report must show the capacity knee, rejection rate, queue wait, memory, and recovery after load falls. |
| Resource profile | At matched throughput, Light peak RSS and CPU per completed request must be no worse than Bifrost. Any optional pool or cache must have a configured bound and a measured benefit. |
The numeric absolute targets are initial release floors. After the first stable baseline they may be tightened, but a configuration or feature addition cannot silently weaken them. If a stricter policy profile performs synchronous remote work by design, publish it as a separate profile with its own SLO rather than averaging it into the normal data-plane result.
Benchmark Method
- Use a fixed-rate, open-loop generator for capacity and overload tests. A fixed number of virtual users is a separate closed-loop test and must not be labelled as RPS.
- Measure the mock provider directly in the same run. Report complete end-to-end latency and gateway-added latency, but never use subtraction as the only release metric.
- Warm DNS, TLS, connection pools, provider codecs, and lazy metrics before the measurement window. Report cold-start behavior separately.
- Run small, 10-KiB, and tool/schema-heavy request profiles; buffered and SSE response profiles; HTTP/1.1 and HTTP/2 where supported; and TLS on and off.
- Use the same upstream protocol, keep-alive policy, connection count, mock latency distribution, response payload, timeout, retry count, and logging policy for both gateways.
- Record histograms rather than averages: P50, P95, P99, P99.9 and maximum for end-to-end latency, gateway-added latency, admission wait, route selection, request/response codecs, time to first token, and stream-chunk processing.
- Record offered, admitted, completed, rejected, failed, retried, and cancelled requests separately. A rejected request is not a successful completion, and a request completed after the measurement window cannot inflate throughput.
- Capture CPU, RSS, allocation rate, task count, open connections, queue depth, and upstream pool reuse throughout the run. Preserve raw results as CI artifacts so regressions can be investigated.
Hot-Path Rules
The normal request path follows these rules:
- Read, decompress, and bound the HTTP body once. Parse the operated routing and policy fields once into a typed compatibility envelope. For same-format forwarding, preserve allowlisted unknown fields without a second generic JSON parse/serialize cycle. Upgrade to full canonical typed content only when an enabled policy mutates content or the selected provider format differs. A failed typed parse never falls back to opaque forwarding.
- Capture one immutable
Arc<LlmPublishedSnapshot>root at request admission. The root contains versionedArcsubgraphs for routing, provider bindings, effective policies, and pricing. Alias maps, capability masks, policy decisions, eligible route lists, weights, pricing references, and compiled hook lists are prepared during reload. Request processing does not scan configuration files or reacquire a config lock at each stage. - Make the root read wait-free, for example with
ArcSwap. A pricing-only or single-alias publication creates a small new root that reuses every unchanged subgraph; it does not rebuild one monolithic allocation. The current runtimeConfigManageruses anRwLock<Arc<T>>; it is a functional reload baseline, but the LLM data plane must not multiply read-lock acquisitions across routing, policy, provider, and streaming stages. The existingconfig-loaderArcSwapimplementation is a reusable pattern. - Reuse one configured HTTP client and connection pool per provider
deployment. Never create a
reqwest::Client, TLS configuration, DNS resolver, or credential object per request. Maintain separate bounded clients only when streaming timeouts or transport policy actually differ. - Keep provider translation in Rust using typed request, response, error, and stream-event codecs. Do not route a request through a scripting runtime, generic SDK dispatcher, thread-pool bounce, or serialize/deserialize bridge.
- Compile optional hooks into the snapshot. A disabled hook creates no task, future, dynamic lookup, log object, or channel message on the hot path.
- Keep local routing and policy decisions in memory. Database, Redis, portal, config-server, secret-manager, and pricing refreshes run outside the normal request path. A dependency needed for fail-closed policy is warmed and projected into the snapshot before it becomes active.
- Update counters and histograms in process. Enqueue audit and usage records to
bounded asynchronous sinks.
bounded-asyncreserves envelope capacity but does not claim crash durability.local-durablewaits only on the bounded single-writer WAL commit watermark defined by the audit contract; when capacity or durability is unavailable, fail before dispatch instead of performing an unbounded synchronous database write. - Preserve byte buffers with
bytes::Bytesor equivalent ownership where the Pingora and provider boundaries allow it. Allocate owned strings only for values that must outlive the input buffer or be transformed. - Add pooling only after allocation profiles identify a benefit. Bifrost’s large prewarmed pools trade memory for speed; Light should prefer bounded buffers, connection reuse, and fewer allocations over a large speculative object pool.
- Resolve provider dispatch while building the snapshot. The hot path invokes
one pre-bound executor and does not hash a provider name, build a client, or
construct a chain of provider wrappers. Benchmark a sealed enum/static
executor against a preconstructed
Arc<dyn InferenceProvider>under the 5,000-RPS and allocation profiles. Use dynamic dispatch when its confidence interval remains within the release margin; use static dispatch only when it provides a material measured benefit worth the maintenance cost. - Compile policy precedence and content requirements at publication. Each
alias plan contains its effective policy, compiled conditional expressions,
priority groups, and whether prompt/completion materialization is needed.
Request processing never reacquires a control-plane
RwLockor clones and merges policy maps.
Publication limits build CPU, peak temporary memory, and retired generations.
Unchanged nodes are structurally shared, dynamic health/in-flight counters stay
outside immutable configuration snapshots, and only affected alias plans are
recompiled. In-flight requests retain old Arc generations; cleanup drops
large retired graphs incrementally on a non-request worker so a frequent update
cannot cause a latency spike. A retirement manager keeps the final non-request
reference until request references drain, ensuring an inference task is not the
thread that recursively frees a large graph. If retained generations exceed a
bound, publication is coalesced or backpressured rather than growing memory
without limit.
The production benchmark covers the complete Light handler chain, not only
crates/llm-gateway. If repeated shared-runtime lock acquisitions or body
copies outside the LLM crate prevent the target, migrate those reads to a
request-scoped immutable handler bundle or a wait-free snapshot. They cannot be
excluded from the reported gateway overhead.
Admission, Concurrency, And Backpressure
Use bounded admission before expensive parsing, token counting, guardrails, or provider dispatch:
- Maintain global, per-principal, per-alias, and per-target in-flight permits. Reuse the fail-fast semaphore pattern already used by MCP resource admission.
- The default provider queue length is zero: dispatch immediately when a permit
is available or return a sanitized
429/503. An explicitly enabled queue is bounded by both depth and wait deadline and exposes its wait time in metrics. - Reserve separate capacity for buffered requests and long-lived streams so a stream flood cannot starve short inference calls.
- Apply per-principal and per-source limits before a stream acquires global capacity. Bound request-header/body read time, stream setup time, absolute stream lifetime, downstream write-progress time, and idle time separately; a heartbeat or one-byte read must not renew every deadline indefinitely.
- Select a target only after a permit can be acquired, or retry selection from the remaining eligible targets. Do not select a saturated target and then build a deep hidden backlog behind it.
- Token counting, JSON Schema compilation, DLP, and other CPU-heavy policies use bounded dedicated executors and admission limits. They must not block Pingora request processing or Tokio worker threads.
- Release permits on every success, error, timeout, downstream disconnect, failed stream setup, and panic boundary. Tests must prove permit recovery.
For hard multi-replica budgets, acquire bounded token/cost leases from the authoritative store and reserve from local atomics. Refresh leases asynchronously before exhaustion. A policy that requires a central transaction for every request is a separately named strict-accounting profile and cannot be the default high-throughput path.
Streaming Data Path
- Parse each upstream SSE event once and translate directly into one canonical event and one client frame. Do not accumulate the full completion unless a configured policy explicitly requires buffering.
- Decode provider-native transport before applying client-visible response policy. A Bedrock event stream, Anthropic SSE event, or OpenAI chunk must first become a semantic client-format event; guardrails must not scrape arbitrary raw byte chunks.
- Use a small bounded channel or direct backpressured writer between provider decoding and Pingora. A slow client must pause bounded upstream reads and eventually cancel; it must not create an unbounded per-stream queue.
- Enforce a downstream write-progress deadline and a minimum sustained drain rate after a bounded grace period. When either is violated, close the client stream, cancel upstream work, finalize partial usage/audit evidence, and release all permits. The maximum stream lifetime is an absolute deadline; SSE comments, TCP trickle reads, and provider heartbeats do not extend it.
- Avoid per-chunk task creation, tracing spans, JSON maps, and log writes. Maintain request-scoped counters and emit one summarized completion event.
- Detect downstream closure promptly, cancel the upstream request, close the channel, release permits, and reconcile the best available usage evidence.
- Keep stream event buffers bounded independently from maximum response bytes and test fragmented UTF-8, large tool arguments, rapid tiny chunks, provider stalls, and slow downstream consumers.
- A bounded-window response guard holds complete semantic frames until its threshold or maximum held-byte limit, evaluates the pending text with a bounded overlap from the previous window, and then releases or blocks the held frames. Publication records the accepted false-negative/context tradeoff explicitly. Policies needing full-response context force buffered mode.
- A text guard may prefer a sentence or punctuation boundary when one occurs inside its configured byte/time window, improving local DLP context without waiting for arbitrary raw chunks. This is only a flush heuristic: hard maximum bytes, maximum hold time, and overlap still apply because generated text and tool-call JSON may contain no sentence boundary. Exact PII placeholder recovery continues to use its bounded token-prefix state machine, not sentence segmentation.
Component Boundaries
apps/light-gateway
The application should own wiring rather than provider logic:
- Register an
llmapplication handler. - Load and hold an
Arc<LlmRuntimeStore>that exposes one wait-freeArc<LlmPublishedSnapshot>root load per request. It may reuse the existingconfig-loaderArcSwapmanager or a runtime-wide equivalent. - Register an
llm-router.ymlreloader. - Pass the existing authenticated principal, agent delegation, correlation, and trace context to the LLM runtime.
- Delegate buffered and streaming response writing to the shared Pingora LLM integration.
The handler can be placed in a chain with existing security and traffic handlers. A typical inference chain is:
handlers:
- correlation
- metrics
- cors
- unified-security
- limit
- access-control
- llm
MVP Application-Body Execution Contract
The current handler registry constructs descriptors whose executable contract
is only PingoraHandler::id(). GatewayProxy::request_filter resolves the
ordered IDs and dispatches behavior with a match, while body-dependent traffic
and access-control work normally completes later in Pingora’s
request_body_filter. MCP is an application-handler precedent: it reads and
answers its request directly from request_filter. Copying that pattern
without an explicit LLM body-policy stage would let the application response
bypass the later generic body filters.
For the MVP, do not make a repository-wide executable-handler-trait refactor a
prerequisite. Register llm in the existing handler registry and add a narrow
branch in GatewayProxy, but immediately delegate to a shared
LlmHttpIntegration owned by light-pingora. That integration executes this
contract exactly once:
- Pre-body handlers before
llmrun in configured order. Correlation, authentication, CORS, request-rate limits, and header policy populate the request context or terminate the request. llmverifies the method, route, media type, content encoding, declared length, and body-read deadline, then collects at most the configured body limit into oneBytes-backed capture. No downstream handler rereads the socket or independently buffers the JSON body.- If
access-controlappeared earlier in the resolved chain, the integration invokes the existing endpoint authorization once with the authenticated principal, trusted headers, endpoint, correlation ID, and parsed request data. It does not rely onrequest_body_filterto perform that check later. - The OpenAI codec validates the request and resolves the public alias. The protocol-neutral LLM runtime then applies host registration, alias/model policy, capability, data-boundary, and admission checks. Both authorization layers must allow the request before an audit marker or provider attempt is created.
- Buffered response policy runs before the response header/body is written. SSE aliases may use only streaming-safe LLM policy; a generic whole-response filter forces buffered mode or makes the alias invalid at publication.
- The shared writer emits buffered JSON or SSE, propagates disconnect cancellation, finalizes audit/usage, and releases every permit.
The LLM request context captures the active access-control runtime and the
published LLM root once. A reload cannot change either decision halfway through
the request. Generic tokenize/detokenize handlers are rejected in an LLM
chain until they are explicitly adapted to this application-body contract;
LLM-aware PII policy belongs in crates/llm-gateway and its normalized content
pipeline. This prevents a configured handler from appearing active while
silently doing nothing.
After the vertical slice is benchmarked, the same integration interface can be generalized for MCP and other application handlers. That refactor is accepted only if it preserves handler ordering and improves maintainability or measured copies/locks; it is not required to obtain the first LLM benchmark.
frameworks/light-pingora
The shared framework should own Pingora-specific integration:
- Request body bounds and media-type checks.
- One-pass bounded body collection with reusable byte buffers; downstream handlers receive the same captured body rather than independently reading or copying it.
- HTTP route matching under
/v1. - Extraction of authenticated and correlation context.
- Buffered response and SSE frame writing.
- Downstream disconnect detection and cancellation propagation.
- Config registration and secret masking metadata.
Provider selection, retries, cost calculations, and cache semantics should not
be embedded in apps/light-gateway/src/main.rs.
New crates/llm-gateway
A protocol-neutral crate should own the inference gateway runtime:
- Public alias and deployment snapshots.
- Precomputed, wait-free request snapshots and bounded admission permits.
- Request policy and capability validation.
- Route eligibility and selection.
- Retry, fallback, circuit breaker, and concurrency coordination.
- Exact/semantic cache interfaces.
- Usage normalization, pricing, reservation, and reconciliation hooks.
- Provider-neutral audit and metrics events.
- Translation between normalized gateway requests and
model-providercalls.
Keeping this logic outside Pingora makes it testable without a network server and reusable by a future sidecar or embedded inference client.
crates/model-provider
Provider clients should own provider-specific authentication, request encoding, response decoding, event parsing, and error classification. They should not own tenant policy or public alias routing.
The current server-oriented providers already construct and retain a
reqwest::Client, which supplies connection pooling across calls. The gateway
runtime must preserve that lifecycle. Runtime construction creates provider
clients once per validated deployment snapshot; request execution borrows the
client and must not rebuild transport state.
Add a gateway-capable operation contract while retaining an adapter for the
current agent trait during migration. Each published ProviderBinding
contains a preconstructed client, per-model capabilities, supported provider
formats, typed request/response/stream codecs, and one pre-bound executor.
The contract does not mandate enum or trait-object dispatch before measurement:
- A sealed built-in enum provides exhaustive compile-time dispatch and avoids a boxed async future, but centralizes provider variants and can increase maintenance coupling.
- A preconstructed
Arc<dyn InferenceProvider>keeps provider crates and optional extensions independent, but may add a virtual call and boxed future. - Both implementations must expose the same provider conformance suite and be benchmarked with real allocation profiles. Only a material, repeatable release-profile difference justifies making static dispatch mandatory.
Whichever representation wins, it is bound during publication. The request path never resolves a provider by string, constructs a trait object, creates a transport client, or stacks retry/routing wrapper providers dynamically.
The request codec owns the bounded compatibility envelope and canonical typed
content. The provider renderer receives only validated fields and the
allowlisted extensions for its exact ProviderFormat; it cannot forward a
generic client JSON object to an unrelated provider.
The canonical contract needs:
| Area | Required types or behavior |
|---|---|
| Content | Text, image URL/data, audio/file references when supported, tool calls, tool results, refusal, and public reasoning summary blocks. |
| Requests | Operation kind, messages/items, tools, tool choice, response format, sampling, output bound, stop conditions, metadata, and streaming. |
| Events | Response start, content delta, tool-call delta, usage update, finish reason, error, and response complete. |
| Usage | Input, output, cached input, reasoning, image/audio, and provider-specific billable units where available. |
| Errors | Stable category, HTTP status, retryable flag, provider request ID, sanitized message, Retry-After, and whether the provider may have accepted work. |
| Capabilities | Per model and operation, including streaming, tools, parallel tools, vision, structured output, reasoning, embeddings, audio, and prompt caching. |
| Control | Deadline and cancellation propagation. |
Capabilities must be per model/deployment when provider offerings differ. A provider-wide boolean is not enough for route safety.
Provider Eligibility
Initial gateway work should prioritize non-interactive, server-oriented providers. Account-login and local CLI providers remain disabled by default in shared deployments.
Before a deployment can serve an alias, it must pass a conformance profile for that alias’s required features:
- Buffered chat request and normalized response.
- SSE stream parsing and termination.
- Usage extraction for buffered and streaming calls.
- Tool call name, ID, and incremental JSON argument preservation.
- Multimodal content conversion where advertised.
- Structured-output enforcement where advertised.
- Timeout, rate-limit, authentication, invalid-request, and server-error classification.
- Cancellation and body-size limits.
- Secret redaction from errors and debug logs.
Strict typed codecs do not require brittle closed-world schemas. Request and
response types distinguish fields the gateway operates on from bounded raw
extensions, preserve unknown same-format fields, and use explicit
Unknown(raw) handling for forward-compatible enum values where safe.
Cross-format conversion remains strict because an unknown construct cannot be
silently translated.
Provider drift is managed operationally as well as through releases:
- Pin provider API versions where the provider permits it and record the negotiated/versioned contract in each deployment snapshot.
- Run scheduled and pre-publication canary fixtures against provider sandboxes or mocks for required operations, errors, and streaming events.
- Quarantine a deployment automatically when a required response shape or capability probe fails; aliases continue only through already conforming targets.
- Preserve sanitized unknown response evidence by digest for diagnosis, never by logging raw content.
- Track adapter compatibility and provider deprecation dates in GenAI Admin so an upgrade can be tested and rolled out before the upstream cutoff.
An alias must fail closed when no eligible deployment can satisfy every required capability. It must not quietly drop tools, images, JSON Schema, or a data-residency restriction to make a fallback succeed.
Control Plane And Model Catalog
The GenAI Admin LLM Model area is the authoritative administration surface.
It should present model registration as related objects rather than one large
record that mixes public policy, provider transport, credentials, and dynamic
health.
Catalog Entities
The initial Portal projection should model these concepts. Exact table and aggregate names can follow existing Portal conventions, but the boundaries are contractual.
| Concept | Scope and responsibility |
|---|---|
| Model catalog | Platform or host-visible description of a physical provider model: provider type, provider model ID, family/version, lifecycle, context/output limits, modalities, supported operations, and declared capabilities. It contains no credential. |
| Model registration | Host authorization to use a catalog model. It records ownership, allowed environments and regions, data classifications, lifecycle state, and any host-specific capability restriction. |
| Provider deployment | Host-scoped callable endpoint with provider type, physical model ID, base URL, region, transport limits, provider-account/quota-group identity, versioned server-owned credential references, and conformance status. Credentials belong here, never on the agent definition. |
| Public model alias | Client-visible logical name such as chat-fast-v1, with allowed operations, required capabilities, token limits, data boundary, logging/PII policy, deprecation, and replacement metadata. |
| Alias route | Ordered or weighted alias-to-deployment relationship with priority, fallback-only status, residency constraints, and rollout/canary policy. |
| Pricing version | Effective-dated rates for input, output, cached input, reasoning, image, audio, service tier, and contract override. Unknown pricing remains explicit. |
| Model policy | Which principals, clients, agents, and product profiles may use an alias, plus budgets, content-logging mode, PII profile, caching, and provider-native extension policy. |
Every mutable tenant record includes host_id. Platform-wide reference rows
may be shared read-only, but a shared catalog entry does not grant a host the
right to use it. The host registration, deployment, alias route, and policy
jointly determine eligibility.
The existing agent_definition_t directly stores model_provider,
model_name, and api_key_ref. Preserve those fields only for a bounded
migration period. New definitions should reference a public alias or model
policy ID. The existing agent_model_rate_t can seed pricing migration, but
runtime accounting must bind a versioned pricing record rather than a mutable
provider/model string pair.
GenAI Admin Workflow
The LLM Model menu should expose focused views over the same aggregates:
- Catalog shows technically supported provider models, lifecycle, capabilities, context/output limits, modalities, and conformance status.
- Host registrations and deployments shows which catalog models the selected host may use, regional endpoints, secret references, transport bounds, provider account/quota groups, credential lifecycle state, approved capacity, and enablement state. Secret values are never displayed.
- Aliases and routes edits public names, required capabilities, eligible deployments, weights, fallback order, rollout percentage, and data boundary.
- Pricing and policies manages effective-dated rates, budgets, allowed principals/agents, content mode, caching, and PII profile.
Provide actions to validate a deployment, run its conformance profile, preview the eligible routes for an alias and sample identity, publish a complete candidate, inspect its digest/version, and roll back to the last valid version. Runtime health and recent latency/error observations may be displayed read-only for operators, but editing or viewing them does not mutate catalog truth.
Agent-definition forms select only authorized public aliases or model policies
for the current host_id. They do not offer free-form provider names, physical
model IDs, base URLs, or API-key fields after migration.
Static And Dynamic Routing Metadata
Portal owns relatively stable, reviewable routing inputs:
- Capabilities and conformance results.
- Context, output, request-byte, and modality bounds.
- Region, residency, data-classification, and provider allowlists.
- Lifecycle, deprecation, replacement, and rollout state.
- Effective-dated price and offline quality/evaluation scores.
- Alias weights, priorities, fallback rules, budgets, and PII/logging profiles.
The gateway owns rapidly changing runtime observations:
- Active/passive health and circuit state.
- Current in-flight work and admission saturation.
- EWMA latency, time to first token, error rate, and rate-limit signals.
- Local lease capacity and recent provider throttling.
Do not update Portal rows for every inference. The gateway combines one immutable catalog/policy snapshot with local runtime observations, and exports bounded telemetry asynchronously. Portal may receive aggregated operational views, but those views are not the routing authority for an in-flight request.
Publication And Consistency
Catalog changes follow the existing Portal command/event and projection model:
- Validate aggregate references, host ownership, secret-reference shape, capability compatibility, and lifecycle transitions in the command path.
- Append the control-plane event and project the Portal read model.
- Build a complete gateway candidate containing provider deployments, aliases, routes, capabilities, pricing, and policy digests.
- Reject an invalid candidate without disturbing the last valid snapshot.
- Atomically publish the candidate and retain its version/digests in audit and agent policy snapshots.
Control-plane resource cardinality should mirror the Portal aggregates: one
changed alias, route, deployment, model policy, or pricing version produces a
small delta rather than republishing every route for a host. Parent resources
should not embed unbounded child lists merely for transport convenience.
However, the gateway’s reload worker—not the request path—resolves those
granular resources into affected request-ready plans. It validates all
references, computes effective policy precedence, compiles expressions and
wildcards, and builds provider priority groups. It rebuilds only affected
subgraphs, structurally shares unchanged Arc data, and swaps one small
LlmPublishedSnapshot root only after the candidate is valid. The root carries
a publication manifest and compatible routing, provider, policy, and pricing
versions, so one atomic load cannot observe a half-published combination.
Pricing may refresh more frequently than model authorization. Publish it as a
separate immutable PricingSnapshot subgraph. A pricing-only update creates a
new root pointing to the existing routing/provider/policy subgraphs and the new
pricing Arc; it does not rebuild the routing graph. Multiple approved sources
can overlay in declared precedence order; invalid or unreadable updates retain
the last valid snapshot. A public catalog such as models.dev can seed proposed
rates, but an operator-approved effective-dated version remains authoritative
and a price entry never creates a model registration or route.
Rapid health, latency, in-flight, and circuit observations are bounded atomics owned by stable deployment-runtime objects, not reasons to republish configuration. Publication metrics include build duration, peak temporary bytes, reused/rebuilt nodes, active generations, and bytes retained by in-flight generations. The publisher coalesces superseded updates and stops admitting new generations when a configured retained-memory bound would be exceeded.
GET /v1/models reads the authorized public-alias view from that snapshot. It
does not query Portal or enumerate physical deployments on demand.
Routing And Reliability
Selection Pipeline
For each request, filter targets in this order:
- Resolve the public alias from one immutable config snapshot.
- Apply caller model/operation allowlists.
- Enforce host, tenant, agent, data-classification, and region constraints.
- Require all request capabilities.
- Remove disabled, unhealthy, open-circuit, or concurrency-saturated targets.
- Remove targets that cannot fit the remaining token or cost budget.
- Apply routing priority, weight, session stickiness, or configured strategy.
The resulting route decision and snapshot version stay attached to the request for its entire lifetime. A config reload does not change an in-flight fallback chain.
Routing Strategies
Implement strategies incrementally:
- Ordered primary/fallback chain for the MVP.
- Priority groups in which lower-numbered healthy groups are preferred and equivalent targets within a group use weight or health/latency score.
- Weighted random across equivalent deployments.
- Least in-flight requests with a bounded weight bias.
- Latency-aware selection using a rolling time-to-first-token and completion latency window.
- Cost-aware selection subject to a minimum capability and quality tier.
- Sticky routing by authenticated principal plus bounded session ID.
- Canary and A/B allocation with an auditable stable hash.
- Region and data-boundary routing as hard eligibility rules, not soft weights.
- Conditional virtual aliases evaluated in declaration order over an allowlisted context such as authenticated claims, requested operation, region, bounded headers, and policy-derived classification. A final explicit fallback is required. Raw prompt access is disabled by default because it is both sensitive and expensive.
Quality-based or semantic routers can be explored later. They must be deterministic enough to audit, include the router’s own latency and cost, and never weaken explicit policy constraints.
Retry, Fallback, And Circuit Rules
- Retry connection failures, timeouts,
408,429, and selected5xxresponses according to typed error policy. - Do not retry authentication, authorization, invalid-request, unsupported parameter, context-length, or safety-policy failures.
- Honor provider
Retry-Afterand apply exponential backoff with jitter. - Bound attempts by both count and the original request deadline.
- Retain one replayable, bounded provider-neutral request or rendered attempt body for pre-output retries. Retry eligibility is explicit for every body size and operation; do not silently disable reliability at an arbitrary small replay-buffer constant.
- Use a different credential or deployment only when policy allows it.
- Open a circuit after a configurable failure threshold; probe with bounded half-open traffic.
- Preserve required capabilities and data-boundary rules across fallback.
- Never begin a fallback after semantic stream output is visible to the client.
- Record each physical attempt separately but charge and report the complete logical request accurately.
- Finalize each attempt’s health signal before selecting the next target. A retryable unhealthy result can eject or penalize that deployment so reselection advances to another target or priority group rather than looping on the same endpoint.
The current ReliableProvider error-string heuristic and nested retry loops are
useful prototype behavior, but gateway reliability must use typed errors and a
single request-scoped attempt budget.
Security And Governance
Identity And Model Access
Reuse existing handler-chain authentication. The LLM runtime receives a trusted identity context containing the authenticated client, user, host, issuer, roles, agent delegation, and relevant policy snapshot.
LLM policy can then enforce:
- Allowed model aliases and operations.
- Maximum input, output, total, and reasoning tokens.
- Maximum request bytes, images/files, tool count, and JSON Schema size/depth.
- Requests per minute, tokens per minute, concurrent calls, and concurrent streams.
- Per-request, per-window, and lifecycle cost budgets.
- Allowed providers, regions, and data-classification boundaries.
- Whether prompts or responses may be cached or content-logged.
- Whether tools, multimodal input, structured output, or provider-native extensions are allowed.
The existing limit handler remains useful for request-count limits. Token,
cost, and model concurrency limits require usage-aware LLM accounting.
Credential Isolation
- Provider credentials are resolved only from server-owned configuration or a secret reference.
- Secret values are masked in module registration, config inspection, errors, metrics, and logs.
- A provider base URL is validated at config load. The client cannot override it, preventing an inference request from becoming an SSRF primitive.
- Provider credentials should be scoped by deployment and environment rather than shared globally.
- Key/deployment selection is audited by opaque ID; raw secret material never enters the request context or audit event.
- Local CLI or account-session credentials require isolated single-user runner profiles and are not enabled in the shared gateway by default.
Credential rotation and capacity routing are different features. A deployment
declares one opaque providerAccountId, quotaGroupId, region, and approved
capacity, and may reference overlapping current/next credential versions for
zero-downtime secret rotation. Every credential in that set inherits the same
quota group, so adding a key cannot increase capacity. A separately approved
provider account/quota is represented as another deployment and alias target.
A 429 can move to that deployment only when policy permits it and the provider
contract treats it as independent authorized capacity; ordinary key rotation
never becomes quota striping.
The gateway must not cycle keys to bypass an upstream RPM/TPM, account tier, fair-use control, or abuse limit. Such behavior creates financial and provider account risk and is rejected during configuration review. The 5,000-RPS performance gate uses a controlled local mock to measure gateway capacity; it is not an instruction to send 5,000 RPS through one or many production provider keys.
Guardrails And Data Protection
Support ordered pre-provider and post-provider policy hooks:
- Prompt and attachment size/type validation.
- PII/tokenization or DLP policy.
- Moderation and prohibited-content policy.
- Prompt-injection and secret-exfiltration signals where configured.
- Tool schema and tool-name allowlists.
- Structured-output schema validation.
- Output redaction and data-boundary checks.
Compile the effective hook order into the alias plan. The default normalized request order is: validate client fields; apply server-owned defaults and prompt enrichment; run local request guardrails and PII classification; tokenize protected spans; reserve the final token/cost bound; select an eligible target; apply only that target’s typed provider transformation and authentication; then dispatch. Provider-specific remote guardrails declare whether they receive original, tokenized, or metadata-only content, and policy validation rejects a data-boundary violation before activation.
On response, decode the provider format into semantic events before applying
content policy. Exact placeholder recovery runs only for the originating
authorized scope. Local safety policy can run before or after recovery as its
profile declares; a remote service receives recovered cleartext only when its
data boundary explicitly permits it. Finally render the original
ClientFormat and release buffered or approved streaming frames.
Streaming has a sharp policy boundary: a post-filter cannot retract bytes that have already reached the client. A policy requiring whole-response inspection must either force buffered mode, use an upstream/provider guardrail that runs before emission, or reject streaming for that alias. Chunk-local filtering is allowed only for policies explicitly designed and tested for bounded windows.
Privacy-Aware Logging
Default audit events contain metadata, not content:
- Request/correlation/trace ID.
- Authenticated policy scopes.
- Public alias, internal route ID, and config snapshot version.
- Timing, status, retry/fallback count, and cancellation reason.
- Normalized usage and computed cost.
- Cache and guardrail outcomes.
Content logging is separately authorized, sampled, redacted or tokenized, encrypted, and retention-bounded. Never put prompts, completions, tool arguments, user IDs, API keys, or model output into metric labels.
Content is not a general-purpose handler attribute. Materialize prompt, completion, tool, or raw-body views lazily only when the compiled policy proves that an authorized content hook needs them. Make the raw request available only inside that content-policy phase and remove it before general transformations, access logs, and telemetry expressions run. Later stages receive normalized metadata, policy outcomes, digests, or encrypted object references.
Parsing and provider-error paths follow the same rule. Ordinary logs never include raw request bodies, provider response/error bodies, malformed SSE data, or a “first N bytes” preview. Emit the normalized error category, provider and route IDs, status, content type, byte length, parser position, and a keyed digest. An authorized encrypted-content capture is an audit operation with its own purpose and retention, not a debug log statement.
Storage Ownership
Use storage according to the data’s workload and security domain:
| Data | Authoritative home | Notes |
|---|---|---|
| Model catalog, aliases, routes, policies, and pricing | Light Portal PostgreSQL | Control-plane data projected to immutable gateway snapshots. |
| Canonical agent conversation/action events | Portal agent event ledger | Bounded normalized event or content reference used to rebuild agent state; not a copy of every physical provider attempt. |
| Logical inference request and physical attempt metadata | Dedicated local/regional audit PostgreSQL | Append-oriented, time-partitioned, separately credentialed, and written asynchronously. |
| Authorized prompt/response bodies and multimodal objects | Encrypted content store | Store ciphertext or an immutable object-store reference plus digest; do not put large content in Portal OLTP rows. |
| Delivery backlog after sink interruption | Gateway-local bounded WAL/spool | Store-and-forward only. It is not the only copy after acknowledgement and is not queried as the audit corpus. |
| Immediate reversible PII mapping | Request memory | Default for synchronous inference; destroy after the response and audit finalization. |
| Durable reversible PII mapping | Separate regional PII vault | Used only for asynchronous, multi-turn, restart/failover, or explicit retention requirements. Never colocate with the audit content corpus. |
Here, “local” means within the organization’s approved trust zone or region. It does not mean that each gateway replica owns the only durable copy. An embedded database such as SQLite is suitable for a single-writer spool, but replica loss, rescheduling, failover, and cross-replica queries make it a poor authoritative audit or PII store.
Audit Record Model
Represent one client call separately from its physical provider attempts:
- Proposed
llm_request_tis the time-partitioned logical-request table. - Proposed
llm_attempt_tis the ordered physical-attempt table keyed to the logical request and partition period. - Proposed
llm_content_object_tstores only encrypted-object metadata, immutable reference, digest, media type, size, encryption-key reference, retention class, and deletion state. - Proposed
llm_dataset_export_trecords a curated audit/evaluation/training export manifest, purpose, approvals, transformations, source partition range, content digests, and retention/deletion state.
At minimum, those records contain:
- A logical request record contains request/correlation IDs, authenticated host/client/agent identities or opaque references, public alias, operation, policy/catalog/config versions, admission and completion timestamps, final status, usage/cost totals, retention class, content mode, PII profile, and content references/digests.
- An attempt record contains logical request ID, attempt number, internal route and deployment IDs, physical model, retry/fallback reason, provider request ID, connect/provider/first-token/total timing, status, cancellation state, provider usage evidence, and pricing version.
- A transformation manifest records whether request content was raw, tokenized, redacted, or omitted at provider dispatch and audit capture. It stores policy and detector versions plus digests, not reversible cleartext.
This distinction preserves evidence when a single logical request retries, falls back, times out after provider acceptance, or returns a partial stream. Streaming produces one summarized attempt record rather than one database row per chunk.
Content Modes And Purpose
Each resolved model policy selects one explicit content mode:
metadata-only: default; store no prompt, completion, or tool content.tokenized-content: store the provider-visible tokenized exchange for approved audit/evaluation use.encrypted-raw: exceptional; store envelope-encrypted pre-tokenization or post-recovery content under stricter authorization and shorter retention.disabled: emit only the minimum operational counters allowed by policy and no durable request-level audit record where regulations require that mode.
Audit, evaluation, and training are different purposes. Operational audit data does not automatically become training data. A separately authorized export job creates a versioned, immutable dataset manifest, applies consent and retention rules, records source digests and transformations, and excludes data that is not approved for the requested purpose.
Delivery And Failure Semantics
Audit policy separates admission pressure from crash durability. A model
policy selects one of these explicit profiles; required by itself must never
be interpreted as an unspecified durability promise:
| Profile | Before provider dispatch | Crash guarantee | Intended use |
|---|---|---|---|
best-effort | Try to enqueue; pressure may drop the record and increments a loss counter. | None. | Local development only; invalid for an alias requiring audit. |
bounded-async | Reserve a complete bounded envelope and queue/spool budget or fail admission. The request does not wait for a disk commit. | A declared tail window can be lost if the process or node fails before the writer commits it. | Default metadata-only, high-throughput production profile and the feature-equivalent Bifrost comparison. |
local-durable | Append the admitted/attempt-start event and wait for the WAL durable watermark before every provider attempt. | A crash can leave an explicitly incomplete attempt, but cannot erase evidence that dispatch was authorized. | Regulated workloads that require pre-dispatch evidence. It has a separately reported latency SLO. |
remote-durable | Wait for an idempotent authoritative-sink transaction. | Survives loss of the gateway node according to the sink’s durability contract. | Later strict-accounting profile; not part of the MVP fast path. |
For bounded-async and local-durable, admission reserves the worst-case
metadata budget for the logical request and configured maximum attempts before
expensive parsing or dispatch. If the queue and allowed spool path cannot honor
that reservation, fail before provider work. Optional content capture may be
dropped independently while retaining required metadata.
The MVP spool is a single-writer, append-only segmented WAL, not an embedded query database. Its stable on-disk format is:
- A segment header containing a fixed Light LLM-audit magic value, WAL format version, segment UUID, gateway-instance ID, and creation timestamp.
- Immutable records encoded as
[length][checksum][sequence][payload].lengthis bounded,checksumcovers the sequence and payload, and a corrupt or partial tail is truncated during recovery. - A UTF-8 JSON payload with its own
schemaVersion, UUIDv7eventId, logical request ID, optional attempt number, event kind, timestamp, published-snapshot digest, and metadata body. MVP WAL payloads never contain prompts, completions, tool arguments, provider error bodies, credentials, or PII. - Event kinds
request_admitted,attempt_started,attempt_finished, andrequest_finished. Records are append-only; completion never overwrites a start record.
Serialization evolution is additive within a payload schema version. A reader must skip a bounded unknown event kind, but it must reject an unknown segment format version rather than guessing record boundaries. Segment and record size limits are validated before publication.
A dedicated writer batches by maximum records, bytes, and commit delay. In
local-durable mode it calls fdatasync after the batch and advances a
monotonic durable-sequence watermark. A request waiting to dispatch succeeds
only when its start sequence is at or below that watermark; timeout, I/O error,
read-only filesystem, or a full volume fails the request without an upstream
attempt. This is bounded group commit, not per-request file opening or an
unbounded synchronous database call.
For buffered local-durable requests, request_finished is also committed
before a successful final response when terminalCommitBeforeResponse is
enabled. For SSE, attempt_started is durable before response headers or the
first semantic event; the terminal event is appended on normal completion. A
crash after streaming begins therefore leaves a durable incomplete attempt
rather than falsely recording success.
The sink consumes WAL events in sequence and writes them idempotently using the
unique eventId plus logical request/attempt keys. A segment is deleted only
after every record has an authoritative acknowledgement and the acknowledgement
checkpoint is itself durable. Startup scans and verifies segments, truncates
only a partial final record, replays unacknowledged events, and marks a durable
start without a terminal event as incomplete. Duplicate delivery is expected
and must not create duplicate request or attempt rows.
local-durable is valid only with a dedicated persistent volume whose
deployment contract survives process and pod restart. Configuration must
declare the persistence class; an ephemeral emptyDir, container filesystem,
or undocumented host path is rejected for this profile. The directory is
gateway-write-only, uses restrictive file permissions and encrypted storage,
and has explicit capacity, retention, and alert thresholds. Node loss beyond
the volume’s durability boundary requires remote-durable rather than a
stronger claim about a local WAL.
A background sink batches metadata into the dedicated audit database and places later authorized large encrypted content in the content store. It never performs an unbounded synchronous Portal or audit-database write on the normal high-throughput path.
Partition metadata by time and make retention removal a partition operation. Encrypt content with per-host or per-retention-class data keys, keep key references out of content rows, and use separate roles for gateway writes, auditor reads, dataset export, and deletion.
Reversible PII Tokenization
Reversible PII tokenization is feasible, but the LLM path needs content-aware token processing rather than only JSON-field replacement.
The existing light-pingora PII handler is a useful cryptographic and schema
baseline: it scopes lookup by host_id, encrypts cleartext values, stores a
keyed value hash, and can tokenize configured request fields and detokenize
configured response fields. It is not the final LLM implementation because:
- It replaces the complete string at a configured JSON path; it does not find multiple sensitive substrings inside a normal message or tool argument.
- Response detokenization expects a field value to be exactly one stored token; it does not recover authenticated placeholders embedded in model prose.
- Response transformation buffers the complete body, which is incompatible with transparent SSE streaming.
- The current insert path does not assign an expiry, host-stable value reuse is linkable across requests, and the default cache can retain cleartext.
- Its database URL can fall back to the shared application database, which is a deployment convenience rather than the desired production security boundary.
Policy And Transformation Flow
The resolved model policy includes a piiProfileId, token scope, detector and
rule versions, allowed data classifications, failure mode, and whether
streaming remains eligible. Apply it in this order:
- Normalize the client request and identify eligible message text, text content parts, structured fields, and tool arguments. Do not tokenize provider routing, schema names, tool names, or control metadata accidentally.
- Detect configured PII types using compiled local rules or a bounded detector profile. A remote DLP dependency is a separately named strict profile with its own admission and latency SLO.
- Replace each sensitive span with a high-entropy authenticated placeholder,
using a short fixed ASCII grammar such as
[LPII1_<base32-id>_<truncated-mac>], and keep the mapping in the request context by default. The exact grammar and tag length follow a security review; the example is illustrative. - Serialize and send only the tokenized canonical request to a cloud provider.
- Scan normalized response text and tool arguments for exact placeholders, validate the MAC and host/request/session scope, and recover authorized values before returning the response to the originating agent.
- Record transformation digests and policy versions in audit metadata. Store content only according to the resolved content mode.
- Destroy request-scoped cleartext mappings after response delivery, audit finalization, and any required retry window.
Never use fuzzy token recovery. A missing, expired, altered, or unauthorized placeholder remains masked or fails the response according to policy. It must never trigger a broader lookup or reveal a value from another host, principal, request, or session.
Exact recovery is a security invariant, but a mangled token does not need to
make the normal user experience brittle. Tokenization profiles therefore also
define unresolvedTokenPolicy:
leave-maskedis the default. Preserve the unresolved placeholder or replace an identifiable malformed token with a generic irreversible marker, record a near-miss outcome, and continue. Never guess the original value.reject-bufferedis for workflows that require complete recovery. It forces buffered output and rejects before any semantic bytes are emitted.- A streaming profile cannot select a policy that would fail the response after earlier content has reached the client.
The gateway adds a short server-owned instruction to copy placeholders verbatim when the alias permits prompt enrichment, and keeps placeholders in structured content/tool values where possible. Provider conformance measures exact preservation, alteration, omission, and hallucinated-token rates over a versioned corpus. A model/deployment that does not meet the profile threshold is ineligible for reversible PII; the alias must use irreversible redaction, buffered strict handling, or another deployment. Near-match detection may produce telemetry, but it never performs a vault lookup or detokenization.
Token Scope And Vault Boundary
Support these scopes explicitly:
request: default. The mapping exists only in request memory and avoids database I/O on the normal inference path.session: opt-in for multi-turn tokenized history. Mappings expire with a bounded session retention and are available to authorized gateway replicas.host: exceptional stable-token mode for a documented integration need. It increases linkability and requires explicit security approval.
A durable mapping is required for asynchronous/batch inference, multi-turn
history containing placeholders, restart/failover recovery, or a response that
may resume on another replica. Access it through a narrow PiiVault interface:
insert with expiry, resolve one exact scoped token, revoke, and expire. The
gateway role cannot scan or export the vault.
The production implementation can be a dedicated regional PostgreSQL vault, a vault service, or a Redis-compatible distributed KV deployment only when it meets the same security and durability contract: independent credentials and network boundary, TLS, encryption of values with external key references, atomic insert/resolve semantics, enforced TTL, bounded memory behavior, restart/failover persistence, backup/recovery objectives, access audit, and tested deletion. An ordinary volatile Redis/Dragonfly cache or an eviction policy that can discard live mappings is not a durable PII vault. PostgreSQL is the conservative durable default; a qualified KV implementation is an optional session-scale profile selected by measured latency and recovery requirements.
The durable vault entry includes at least host ID, opaque token ID, token
format/version, scope kind, a non-reversible scope binding, PII type, encrypted
value, nonce, key ID, creation/expiry timestamps, and active/deletion state.
Request scope does not use the current host-stable value_hash uniqueness
rule. Session/host deduplication, when explicitly required, uses a separate
keyed hash and policy so linkability is visible and reviewable.
Using a separate schema and role in the Portal PostgreSQL cluster is an acceptable transition for development or an initial low-risk deployment, but it shares administrator, backup, and failure boundaries. Production PII mappings should not live in the Portal database, an audit-content database, or a replica-local embedded database. Colocating reversible mappings with logged content would defeat the intended breach separation.
Streaming Recovery
Provider streams can split a placeholder across arbitrary SSE and UTF-8 chunk
boundaries. A streaming-compatible PII profile keeps a bounded suffix no larger
than the maximum placeholder length, emits only bytes that cannot begin a
placeholder, and validates/replaces complete placeholders before release.
An altered candidate follows leave-masked; it is not recovered fuzzily and
does not terminate a normal stream after earlier semantic output. A strict
reject-buffered profile is never published as streaming-compatible.
If detection or response policy requires whole-message context, force buffered mode or reject streaming for that alias. Do not emit raw partial token syntax and attempt to retract it later. Tokenization and recovery benchmarks must be reported as named policy profiles; the metadata-only/no-PII profile remains the baseline performance contract.
Token And Cost Governance
Usage accounting must distinguish estimated, provider-reported, and locally counted values.
When a provider lacks a native token-count endpoint, a local tokenizer may
answer an internal estimate or a later provider-native compatibility endpoint.
The result is labelled local-estimate with tokenizer/model-table version; it
can enforce a conservative admission bound but cannot be recorded as exact
provider billing evidence. A native count response is also not inference usage
and must not be charged as consumed prompt tokens.
Recommended flow:
- Validate the alias’s maximum output and estimate or count input units.
- Atomically reserve the maximum allowed tokens and cost before provider dispatch.
- Reject the request when the authoritative budget cannot reserve capacity.
- Reconcile the reservation with trusted provider usage when the request finishes.
- Retain a conservative charge or explicitly mark accounting incomplete when a timeout/cancellation prevents authoritative usage.
- Store the pricing-table version and evidence source with the ledger entry.
Scopes should include host, customer/organization, team, client, user, agent, public alias, and provider deployment as needed. Multi-replica deployments need an authoritative shared quota store; process-local counters are not sufficient for hard budgets.
Provider-side capacity is tracked by declared providerAccountId and
quotaGroupId, not merely credential ID. Local admission consumes RPM/TPM and
concurrency leases for that group before dispatch and updates them from
provider rate-limit signals. Adding or rotating another secret in the same
group does not create more capacity.
Opaque routes use request/byte/duration/concurrency units plus the configured fixed cost envelope and account-spend ceiling. Because realized token usage is unknown, reconciliation cannot release a pessimistic reservation based on a guess; only authoritative later billing evidence may amend it. This makes opaque compatibility intentionally less efficient than normalized routing rather than a way around financial governance.
The pricing catalog must be versioned and support:
- Input, output, cached input, and cache-creation rates.
- Reasoning, image, audio, and other provider-specific billable units.
- Tiered context pricing and provider service tiers.
- Contract-specific overrides.
- A clear
unknownstate. Unknown pricing must not silently become zero when a hard cost budget is configured. - Atomic replacement, declared source precedence, and last-valid retention on refresh failure. The logical request and every physical attempt capture the exact pricing snapshot and effective tier used for reconciliation.
Caching
Caching is valuable but should follow routing, security, and usage correctness.
Exact Response Cache
Add after the MVP with a cache key that includes at least:
- Authenticated tenant/policy partition.
- Public alias and alias snapshot version.
- Normalized messages/items, tools, tool choice, response format, and relevant sampling parameters.
- Data-boundary and guardrail policy version.
Do not cache tool-call responses, sensitive requests, or nondeterministic requests by default. Provider prompt caching and gateway response caching are different features and need separate metrics.
Semantic Cache
Semantic caching is a later, explicit opt-in because similar prompts are not necessarily interchangeable. It requires:
- Tenant- and policy-isolated vector indexes.
- A versioned embedding model and similarity threshold.
- Alias, tool, structured-output, locale, and safety-policy compatibility.
- No cross-tenant hits.
- Auditability of the matched entry and score.
- A deletion and retention model for source text and embeddings.
Its lookup is an explicit sub-operation in the request deadline and budget, not hidden preprocessing:
- Perform normal identity, alias, size, data-boundary, and fail-fast global admission first; check the cheaper exact cache before semantic work.
- Acquire separate bounded semantic-cache and embedding permits. Reserve the embedding token/cost envelope and a configured deadline slice without consuming all time needed for the primary LLM route.
- Call only a policy-approved embedding deployment, then run a bounded vector lookup. Account for both operations and audit their model/index versions.
- On a hit, apply current authorization and post-response policy before
returning. On a miss, release cache permits and continue with the remaining
LLM deadline and budget. A lookup timeout follows the alias’s explicit
cacheFailureMode, normallytreat-as-miss; it never waits without bound. - Coalesce identical in-flight cache fills and bound fill concurrency so a popular miss cannot create an embedding or provider stampede.
Client end-to-end time to first byte starts at gateway admission and therefore
includes semantic lookup. Report semantic_embedding_duration,
semantic_lookup_duration, time_to_cache_hit, and the later provider
time-to-first-token separately. A cache hit has no provider TTFT. Release and
performance gates use named semantic-cache profiles rather than hiding this
latency inside routing.
Cached responses still pass current authorization and post-response policy. Usage and cost clearly distinguish cache hits from provider calls.
MCP And WebSocket Integration
The normal agent tool loop remains:
agent -> MCP router tools/list
agent -> LLM gateway chat request with selected tool schemas
LLM gateway -> model provider
model provider -> tool call
LLM gateway -> agent
agent -> MCP router tools/call
MCP router -> backend API or MCP server
agent -> LLM gateway with tool result
This separation preserves the MCP router as the execution, authorization, and audit boundary. The LLM gateway must not accept a model-generated target URL or execute a tool solely because the model emitted its name.
A later feature may let a policy-selected tool profile inject a small set of MCP schemas into an LLM request. Even then:
- The tool set is selected by server-owned policy and the authenticated principal.
- Tool execution still goes through the MCP router.
- The client agent remains responsible for the tool loop unless a separately designed managed-agent service owns it.
The existing websocket handler routes browser UI traffic to agents. A future
OpenAI-compatible Realtime API has different session, audio, provider, and
billing semantics and must use a distinct handler and configuration rather than
overloading the UI router.
Configuration Model
Use llm-router.yml for the data-plane projection. Provider credentials are
masked values or secret references populated through the existing runtime
configuration flow.
Illustrative configuration:
enabled: ${llm-router.enabled:false}
pathPrefix: ${llm-router.pathPrefix:/v1}
maxRequestBodyBytes: ${llm-router.maxRequestBodyBytes:4194304}
maxResponseBodyBytes: ${llm-router.maxResponseBodyBytes:16777216}
requestTimeoutMs: ${llm-router.requestTimeoutMs:120000}
streamIdleTimeoutMs: ${llm-router.streamIdleTimeoutMs:30000}
maxConcurrentRequests: ${llm-router.maxConcurrentRequests:1024}
maxConcurrentStreams: ${llm-router.maxConcurrentStreams:512}
unsupportedParameterPolicy: ${llm-router.unsupportedParameterPolicy:reject}
admission:
maxQueuedRequests: ${llm-router.admission.maxQueuedRequests:0}
maxQueueWaitMs: ${llm-router.admission.maxQueueWaitMs:0}
overloadStatus: ${llm-router.admission.overloadStatus:503}
streamBufferEvents: ${llm-router.admission.streamBufferEvents:32}
streamSetupTimeoutMs: ${llm-router.admission.streamSetupTimeoutMs:5000}
maxStreamLifetimeMs: ${llm-router.admission.maxStreamLifetimeMs:120000}
downstreamWriteProgressTimeoutMs: ${llm-router.admission.downstreamWriteProgressTimeoutMs:10000}
minStreamDrainBytesPerSecond: ${llm-router.admission.minStreamDrainBytesPerSecond:256}
publication:
maxRetainedGenerations: ${llm-router.publication.maxRetainedGenerations:8}
maxRetainedBytes: ${llm-router.publication.maxRetainedBytes:536870912}
coalesceWindowMs: ${llm-router.publication.coalesceWindowMs:100}
opaqueDefaults:
maxRequestBytes: ${llm-router.opaqueDefaults.maxRequestBytes:1048576}
maxResponseBytes: ${llm-router.opaqueDefaults.maxResponseBytes:8388608}
maxDurationMs: ${llm-router.opaqueDefaults.maxDurationMs:30000}
fixedCostReservationUsd: ${llm-router.opaqueDefaults.fixedCostReservationUsd:}
requireAccountSpendCeiling: ${llm-router.opaqueDefaults.requireAccountSpendCeiling:true}
telemetry:
auditQueueCapacity: ${llm-router.telemetry.auditQueueCapacity:8192}
usageQueueCapacity: ${llm-router.telemetry.usageQueueCapacity:8192}
perChunkEvents: ${llm-router.telemetry.perChunkEvents:false}
audit:
admissionPolicy: ${llm-router.audit.admissionPolicy:required}
durability: ${llm-router.audit.durability:bounded-async}
contentMode: ${llm-router.audit.contentMode:metadata-only}
includeProviderRequestId: ${llm-router.audit.includeProviderRequestId:true}
contentSampleRate: ${llm-router.audit.contentSampleRate:1.0}
terminalCommitBeforeResponse: ${llm-router.audit.terminalCommitBeforeResponse:true}
spool:
enabled: ${llm-router.audit.spool.enabled:true}
path: ${llm-router.audit.spool.path:/var/lib/light-gateway/llm-audit}
persistenceClass: ${llm-router.audit.spool.persistenceClass:ephemeral}
maxBytes: ${llm-router.audit.spool.maxBytes:1073741824}
segmentBytes: ${llm-router.audit.spool.segmentBytes:67108864}
maxRecordBytes: ${llm-router.audit.spool.maxRecordBytes:65536}
maxBatchRecords: ${llm-router.audit.spool.maxBatchRecords:256}
maxBatchBytes: ${llm-router.audit.spool.maxBatchBytes:1048576}
maxCommitDelayMs: ${llm-router.audit.spool.maxCommitDelayMs:1}
commitTimeoutMs: ${llm-router.audit.spool.commitTimeoutMs:10000}
sink:
type: ${llm-router.audit.sink.type:postgres}
databaseUrl: ${llm-router.audit.sink.databaseUrl:}
batchSize: ${llm-router.audit.sink.batchSize:256}
contentStore:
type: ${llm-router.audit.contentStore.type:none}
bucket: ${llm-router.audit.contentStore.bucket:}
pii:
defaultProfile: ${llm-router.pii.defaultProfile:none}
vault:
type: ${llm-router.pii.vault.type:none}
url: ${llm-router.pii.vault.url:}
credentialRef: ${llm-router.pii.vault.credentialRef:}
durabilityProfile: ${llm-router.pii.vault.durabilityProfile:durable}
maxConnections: ${llm-router.pii.vault.maxConnections:8}
defaultTtlSeconds: ${llm-router.pii.vault.defaultTtlSeconds:86400}
profiles:
- id: cloud-request-scoped
scope: request
detector: local-rules-v1
tokenFormat: authenticated-placeholder-v1
unresolvedTokenPolicy: leave-masked
streamingMode: bounded-token-window
providers:
openai-primary:
type: openai
baseUrl: ${llm.providers.openaiPrimary.baseUrl:https://api.openai.com/v1}
providerAccountId: openai-account-primary
quotaGroupId: openai-tier-primary
credentials:
- id: openai-key-current
secretRef: ${llm.providers.openaiPrimary.secretRef:}
lifecycle: current
connectTimeoutMs: 3000
requestTimeoutMs: 90000
anthropic-primary:
type: anthropic
baseUrl: ${llm.providers.anthropicPrimary.baseUrl:https://api.anthropic.com}
providerAccountId: anthropic-account-primary
quotaGroupId: anthropic-tier-primary
credentials:
- id: anthropic-key-current
secretRef: ${llm.providers.anthropicPrimary.secretRef:}
lifecycle: current
connectTimeoutMs: 3000
requestTimeoutMs: 90000
models:
- name: chat-fast-v1
clientFormats: [openai-chat-completions]
operations: [chat]
processingMode: normalized
capabilities: [streaming, tools, vision]
maxInputTokens: 128000
maxOutputTokens: 8192
dataClassifications: [public, internal]
routes:
- id: openai-fast-primary
provider: openai-primary
model: provider-physical-model-a
providerFormat: openai-chat-completions
priority: 0
weight: 80
regions: [ca, us]
- id: anthropic-fast-fallback
provider: anthropic-primary
model: provider-physical-model-b
providerFormat: anthropic-messages
priority: 1
weight: 100
regions: [ca, us]
routing:
strategy: priority-weighted
maxAttempts: 3
baseBackoffMs: 100
maxBackoffMs: 2000
retryStatuses: [408, 429, 500, 502, 503, 504]
circuitBreaker:
failureThreshold: 5
resetTimeoutMs: 30000
halfOpenRequests: 1
sessionHeader: X-Light-Session-Id
This is a proposed shape, not a statement that these keys are already
implemented. Portal source records keep secret-bearing deployments, public
aliases/routes, model policies, audit-sink configuration, and PII profiles as
separate security and lifecycle objects. llm-router.yml is their validated
data-plane projection, not a second independently edited source of truth.
Configuration validation must reject:
- Duplicate alias, provider, route, or deployment IDs.
- Empty credential sets for an enabled provider unless its authentication mode explicitly allows them.
- Deployments without provider-account/quota-group identity, invalid credential lifecycle overlap, or a capacity policy that would cycle keys inside one quota group to bypass upstream limits.
- Invalid or unsafe provider base URLs.
- Fallback targets missing an alias’s required capability or region.
- A route whose
ProviderFormatcannot represent the alias’sClientFormatand operation without dropping a required field or event. detectoropaqueprocessing on an alias that requires content guardrails, reversible PII, normalized audit, cross-format fallback, authoritative token accounting, or exact realized per-call cost accounting.- An opaque route without bounded bytes, duration, request/concurrency limits, and either a pessimistic fixed cost reservation or an authoritative provider-account spend ceiling.
- Routes to interactive providers in a shared profile.
- Impossible token or timeout bounds.
- Unknown strategy, operation, capability, or unsupported-parameter policy.
- Unbounded or contradictory admission settings, including a non-zero queue without a positive queue-wait deadline.
- Stream admission without positive setup, write-progress, idle, and absolute lifetime deadlines, or publication without retained-generation and retained-memory bounds.
- Unbounded audit, usage, cache, stream-event, or provider concurrency buffers.
- An alias that requires audit selecting
best-effort, or an unknown admission/durability profile. local-durablewithout a bounded WAL, positive commit timeout, valid segment/record/batch limits, and a declared persistent-volume class.remote-durablewithout an authoritative idempotent sink and a bounded transaction timeout.- An alias or model policy referencing a missing/inactive deployment, pricing version, content mode, or PII profile.
- A durable PII profile without a separately credentialed vault, expiry, key reference, and exact host/session authorization policy.
- Required audit admission without a complete worst-case envelope reservation,
or
encrypted-rawcontent mode without an encrypted content store and retention class. disabledaudit selected by a policy that requires a durable request record, or sampling applied to required metadata instead of optional content.- Secret values that would be exported without a mask.
Reload builds and validates a complete candidate runtime before one atomic swap. The previous runtime remains active when candidate validation fails.
Feature Priorities
MVP: Compatible And Safe Inference
llmhandler and reloadablellm-router.ymlruntime.- OpenAI-compatible
/v1/modelsand/v1/chat/completions. - Buffered and SSE streaming with disconnect cancellation.
- Text, tool calling, usage, and provider-supported image input.
- Public aliases and ordered primary/fallback routes.
- Portal-backed host model registrations, deployments, aliases, model policies, pricing versions, and an atomic gateway projection.
- Server-side credentials with masking and base-URL validation.
- Existing Light authentication, access control, correlation, metrics, and request rate limits.
- Request, token, response, timeout, concurrency, and attempt bounds.
- Typed errors and explicit unsupported-parameter behavior.
- Separate client/operation/provider-format types, with a same-format compatibility fast path and no parse-failure downgrade to opaque forwarding.
- Provider conformance tests for the first supported deployment set.
- Provider API-version pinning where available, forward-compatible same-format envelopes, and deployment quarantine on required contract drift.
- Metadata-only logical-request and physical-attempt audit events, normalized
usage,
bounded-asyncdelivery for the default performance profile, and a separately measuredlocal-durableWAL profile feeding a dedicated audit sink. - Checked-in direct/mock/Bifrost benchmark harness, immutable benchmark manifests, and passing absolute and comparative performance gates.
- One wait-free structurally shared published-snapshot root, reusable provider clients, fail-fast bounded admission, bounded streaming/audit/usage channels, and slow-consumer setup/write-progress/absolute deadlines.
Production Hardening
- Weighted and least-in-flight routing.
- Health/latency-scored priority groups and retry-driven outlier reselection.
- Per-target circuit breakers and active/passive health signals.
- Shared per-principal token, cost, and concurrency budgets.
- Provider-account/quota-group capacity, governed credential lifecycle, and explicit prohibition of key cycling for limit evasion.
- Versioned pricing and usage reconciliation.
- Model/region/data-boundary policy.
- Structured outputs and expanded multimodal conformance.
- Exact response caching.
- Pre/post guardrail hooks and strict streaming policy modes.
- Policy-selected tokenized/encrypted content capture with purpose-specific retention and curated evaluation/training dataset export.
- Request-scoped LLM content tokenization and exact response recovery, followed by a separately credentialed regional PII vault for session, asynchronous, and failover profiles.
- Portal configuration, route inspection, usage, and budget views.
- OpenTelemetry traces and operational dashboards.
- Multi-replica chaos, failover, and stream soak testing.
- Continuous capacity-regression testing for the production handler profile, including allocation, CPU, RSS, connection reuse, and overload recovery.
Endpoint And Intelligence Expansion
/v1/responseswith native normalized events.- Embeddings, moderation, images, audio, rerank, and batch APIs when supported by dedicated provider operations.
- Semantic caching.
- Cost-, latency-, and quality-aware routing.
- Canary/A-B routing and policy-controlled session stickiness.
- Selected Anthropic or Google native compatibility adapters.
- Realtime inference over a dedicated WebSocket/WebRTC handler.
- Optional governed MCP tool-schema injection without in-gateway execution.
Observability
Emit bounded-cardinality metrics for:
- Logical requests and physical attempts.
- Successes and normalized error categories.
- Active and queued requests/streams.
- Request duration, provider latency, time to first token, and stream duration.
- Input, output, cached, reasoning, and total tokens.
- Estimated and reconciled cost.
- Retries, fallback depth, circuit state changes, and route saturation.
- Cache hit/miss/bypass and guardrail outcomes.
- Downstream disconnects and upstream cancellation outcomes.
- Slow-consumer write-progress/drain-rate terminations and stream permit hold time.
- Snapshot publication build/retire duration, rebuilt/shared nodes, retained generations, and retained bytes.
- PII placeholder preservation, alteration, omission, hallucination, unresolved handling, and recovery outcomes without token IDs or values as labels.
- Semantic embedding/lookup duration, time to cache hit, coalesced fills, and cache-failure-mode outcomes.
Labels can include public alias, operation, route ID, provider type, status class, and environment when their value sets are bounded. Never label metrics with prompt text, user-provided model strings, user IDs, session IDs, raw provider error text, or provider request IDs.
Each trace should separate:
- Authentication and policy evaluation.
- Quota reservation.
- Exact-cache lookup and, when enabled, semantic embedding/vector lookup.
- Route selection.
- Each provider attempt.
- First-token wait and stream transfer.
- Guardrail processing.
- Usage and cost reconciliation.
Tracing is sampled and request-scoped. The default profile does not create a span or export event for every stream chunk. High-detail diagnostic tracing is time-bounded, rate-limited, and excluded from release benchmark comparisons unless enabled identically for every candidate.
Testing Strategy
Protocol Tests
- Golden request/response fixtures from current OpenAI SDKs.
- Chat message roles, content parts, tool calls, structured output, usage, and error envelopes.
- SSE fragmentation at arbitrary byte boundaries, multi-byte UTF-8, tool-call
argument deltas, final usage,
[DONE], and midstream failure. - SDK smoke tests with at least Python and TypeScript clients using only a base URL and credential change.
- Same-format unknown-field preservation, operated-field validation, and proof that cross-format conversion rejects unknown extensions it cannot represent.
- Forward-compatible unknown enum/field fixtures prove safe same-format preservation while required or unsafe drift quarantines the deployment.
- Proof that an internally requested streaming usage record is stripped when
the client did not request
stream_options.include_usage.
Provider Contract Tests
- Provider mock servers for every supported success and error shape.
- Capability conformance by physical model/deployment.
- Authentication, rate-limit, invalid-request, context-limit, timeout,
5xx, malformed JSON, oversized body, and truncated stream behavior. - Provider request ID,
Retry-After, usage, and cancellation extraction. - Secret redaction from every error and log path.
- Malformed JSON/SSE and provider parse failures prove that raw body prefixes, prompts, completions, tool arguments, and provider error bodies never reach ordinary logs or traces.
Routing And Governance Tests
- Deterministic alias resolution from one immutable published root and its generation-compatible sub-snapshots.
- Host registration, deployment, alias-route, model-policy, pricing-version, and agent-definition referential validation.
- Proof that a shared catalog row does not grant an unregistered host access
and that
/v1/modelsnever exposes physical deployments. - Capability, policy, region, and data-boundary target filtering.
- Weighted distribution and stable canary/session allocation.
- Retry/fallback deadline and attempt limits.
- Priority-group failover proves that a retryable unhealthy response updates outlier state before reselection and does not choose the same ejected target.
- Replay tests cover requests above 64 KiB up to the configured maximum and prove that retry eligibility is explicit rather than silently disabled.
- Proof that no retry or fallback starts after the first semantic stream event.
- Atomic token/cost reservation under concurrency and correct reconciliation on success, error, timeout, and cancellation.
- Cross-tenant cache and session-stickiness isolation.
- Credential-rotation tests prove keys in one quota group share capacity and a
429cannot trigger quota-evasion cycling; independent approved accounts can fail over only according to explicit policy. - Opaque routes enforce byte/request/duration/concurrency and fixed-cost/account ceilings, mark token/realized cost unknown, and cannot satisfy a normalized alias accidentally.
Runtime Tests
- Candidate config rejection leaves the previous runtime active.
- Config reload does not mutate in-flight route snapshots.
- Pricing-only and one-alias publications structurally share unchanged subgraphs, preserve a generation-consistent root, coalesce rapid updates, and stay within retired-generation/byte bounds under long-lived requests.
- High-concurrency buffered and streaming load tests.
- Slow-client, disconnect, provider-stall, circuit-breaker, and provider-outage chaos tests.
- Slowloris tests trickle request bytes and downstream reads, send heartbeats, and hold streams across setup/idle/write-progress/absolute deadlines; every path cancels upstream and recovers per-principal/global permits.
- Multi-replica budget and cache tests against the selected shared stores.
- Metrics cardinality and content-leak checks.
Audit And PII Tests
- One logical audit record with all ordered physical attempts for retry, fallback, timeout-after-acceptance, partial stream, cancellation, and cache hit paths.
- Required-audit admission failure when the complete logical-request/attempt
envelope cannot be reserved;
best-effortis rejected for a required alias. - WAL fixtures cover versioned segment headers, record length/checksum/sequence, maximum record and segment sizes, partial-tail truncation, mid-segment corruption rejection, and unknown payload event kinds.
local-durableproves that no provider mock observes a request before the corresponding start sequence reaches the durable watermark. Commit timeout,fdatasyncfailure, read-only/full volume, and writer termination all fail before dispatch and release reservations.- Recovery replays unacknowledged events without duplicate rows, preserves a durable start as an incomplete attempt when no terminal event exists, and deletes a segment only after a durable authoritative acknowledgement.
bounded-asynctests and telemetry expose its documented crash-loss window; they never label an in-memory reservation as durable.- Time partition creation/retention, encrypted content references, key/role isolation, purpose-specific dataset export, and deletion evidence.
- PII detection and replacement inside message text, content-part arrays, and tool arguments, including multiple values and repeated values.
- Exact authenticated-token recovery, expiry, request/session/host scoping, unresolved-token policy, and proof that forged or cross-host tokens never reveal cleartext.
- A versioned per-model preservation corpus measures exact, altered, omitted,
and hallucinated placeholders;
leave-maskednever guesses or fails a partially emitted stream, whilereject-bufferedemits no partial content. - Placeholder fragmentation at every streaming byte boundary, including UTF-8 boundaries, slow consumers, cancellation, and maximum suffix-buffer bounds.
- Request-scoped profiles perform no vault I/O and destroy cleartext state; durable profiles survive the documented multi-replica failover cases.
- Every
PiiVaultimplementation passes the same TTL, restart/failover, eviction, encryption, authorization, backup/restore, audit, and deletion contract; a volatile cache fails the durable profile. - Content-mode tests prove that
metadata-only,tokenized-content,encrypted-raw, anddisabledproduce only the authorized records.
Performance And Capacity Tests
- A direct mock-provider baseline plus Light and pinned-Bifrost runs generated from one versioned manifest.
- Open-loop offered-load sweeps that locate the sustainable capacity knee and verify prompt load shedding above it.
- The 500-RPS, 5,000-RPS, production-handler, streaming, overload, large-body, and cold-start profiles defined by the release performance gates.
- Five or more steady-state repetitions with raw histograms, confidence intervals, CPU, RSS, allocation, connection, queue, and task telemetry.
- Assertions that one request captures one published root, does not build an HTTP client, performs no synchronous control-plane I/O, and creates no disabled-hook task.
- Assertions that the request path takes no control-plane/configuration
RwLock, performs no policy-map merge, and does not resolve a provider by string. Measure normalized conversion and same-format compatibility paths separately. - Static-enum and preconstructed dynamic provider dispatch run under identical 5,000-RPS/allocation profiles; dispatch is not standardized until the confidence interval demonstrates whether either is materially better.
- Snapshot-churn profiles vary pricing and alias publication rates while measuring build CPU, peak temporary/retained memory, generation retirement, allocator behavior, and request P99.
- Named benchmark profiles for metadata-only audit, tokenized content, encrypted-raw content, request-scoped PII, durable-vault PII, and buffered whole-response policy. Do not average strict-policy remote I/O into the default fast-path result.
- Run metadata-only audit separately as
bounded-asyncandlocal-durable. The former remains enabled in the production-handler comparison; the latter reports WAL batch/commit-wait histograms and never dispatches before its durable watermark. - Slow-provider and slow-client tests proving bounded queues, cancellation, permit release, and recovery without a latency backlog.
- Semantic-cache profiles include embedding and vector lookup in end-to-end latency/cost, exercise timeout-as-miss and stampede coalescing, and report cache-hit latency separately from provider TTFT.
- A CI non-inferiority check against the last accepted Light baseline on every change to the handler, canonical types, router, provider codecs, admission, telemetry, or streaming path. Run the external Bifrost comparison on release candidates and scheduled performance infrastructure.
Rollout Plan
- Check in the benchmark harness, mock provider, payload corpus, benchmark manifest, and direct-provider baseline. Pin Bifrost and agentgateway references and record the first capacity curves before implementing the Light path; Bifrost remains the release non-inferiority comparator.
- Freeze the MVP application-body contract and audit durability profiles.
Build focused prototypes for one-pass body capture/security ordering and the
single-writer WAL/group-commit watermark. Measure body copies, handler locks,
fdatasync, batch size, and commit wait before selecting implementation defaults. No provider request is part of this prototype. - Add canonical inference types, typed errors, cancellation, and streaming to
model-provider, preserving an adapter for existinglight-agentandlight-workflowcallers. - Build provider conformance tests and enable a small server-safe provider set. Start with at least two different provider formats so cross-provider normalization and fallback are exercised rather than assumed.
- Add
crates/llm-gatewaywith a checked-in localllm-router.ymlprojection, immutable root, alias resolution, request validation, ordered fallback, usage normalization, and protocol-neutral tests. This local projection is a vertical-slice fixture, not a second control-plane authority. - Add
LlmHttpIntegration, register the Pingorallmbranch, and implement/v1/modelsplus buffered Chat Completions. Prove that body-dependent endpoint authorization and LLM alias policy both execute before dispatch and that existing MCP/WebSocket paths are unchanged. - Run the first 500-RPS direct/Light/Bifrost comparison on the buffered vertical slice. Profile full-handler-chain locks, body copies, allocations, and provider dispatch. Benchmark sealed-enum and preconstructed dynamic dispatch here; refactor the shared handler bundle only when measurements show it is needed.
- Add the Portal model catalog, host registration, deployment, public alias, alias-route, provider-account/quota group, credential lifecycle, pricing, and model-policy aggregates and projections against the now-exercised data-plane contract.
- Publish Portal deltas into the validated, structurally shared gateway root, replace the vertical-slice fixture as production authority, and migrate agent definitions toward alias/policy references while retaining bounded compatibility for existing fields.
- Add metadata-only logical-request/physical-attempt events, envelope
reservation,
bounded-asyncdelivery, the versioned local WAL,local-durablegroup commit/recovery, and idempotent batched delivery to the dedicated audit store. - Repeat the 500-RPS production-handler comparison with
bounded-asyncaudit enabled and publish the separatelocal-durablecommit-wait/capacity profile. Neither profile may dispatch when its declared admission or durability contract is unavailable. - Add SSE streaming and prove bounded buffering, cancellation, no-post-output-fallback rules, slow-consumer deadlines/permit recovery, and the streaming performance gate.
- Pass the 5,000-RPS, production-handler, streaming, and overload profiles before the MVP is declared production-ready.
- Add shared token/cost budgets, richer routing, guardrails, and exact caching.
- Add request-scoped LLM PII tokenization/recovery and tokenized-content
capture. Measure placeholder preservation per model and default unresolved
tokens to
leave-masked. Add a durable regionalPiiVaultimplementation only for session, asynchronous, and failover profiles, and prove its common security/recovery contract and separate performance SLO. - Add governed encrypted-raw capture and curated evaluation/training dataset export after access, retention, deletion, and key-isolation reviews pass.
- Add
/v1/responseswithout translating it through the less expressive Chat Completions representation. - Expand operations and provider-native compatibility only from demonstrated client requirements and conformance coverage.
- Add semantic caching only with separate embedding/vector admission, deadline and cost slices, stampede control, and named performance profiles.
Config Snapshot Lifecycle And Agent Cutover
The current immutable config-server values.yml snapshot is authoritative for
llm-router, including providers, deployments, aliases, prices, policies, and
non-secret runtime material references. At startup the gateway resolves
llm-router.yml, compiles the complete graph, and publishes one immutable
runtime snapshot. During an explicit control-plane reload, the runtime fetches
the current values.yml again and invokes only the selected module reloaders.
Selecting llm-router compiles and atomically publishes a new LLM snapshot;
failure retains the last-known-good snapshot. Reloading another module cannot
change LLM routing.
There is no independent LLM filesystem projection, polling worker, checkpoint,
or publication acknowledgement channel. Control-plane publication writes typed
llm-router.* properties to the selected instance. The operator then creates
or promotes a normal config snapshot and explicitly restarts the gateway or
reloads llm-router.
runtimeMaterial.credentialEnvironment authorizes opaque credential://
references to application-owned environment-variable names. Direct env:
references are also resolved locally. Neither values.yml nor the rendered
module configuration contains secret values. A successful snapshot load proves
configuration application, not provider reachability; provider testing remains
a separate operation.
Migrated light-agent definitions resolve a direct alias or exactly one policy
default before a turn is persisted. The immutable turn stores provider
gateway plus that alias, and the gateway client sends it in the ordinary
OpenAI model field. INTERNAL_LEGACY aliases require an explicit agent
binding (aliasVisibility: INTERNAL_LEGACY plus boundAgentDefId in Portal),
are returned only to that agent, and remain absent from /v1/models.
Endpoint URLs and service credentials come only from
llm-gateway-client.yml; agent records cannot supply them.
Run scripts/run-llm-production-integration-gates.sh. Pass a disposable
PostgreSQL URL (or set PORTAL_LLM_TEST_DATABASE_URL) to include the additive
schema and internal-alias ownership gate.
Acceptance Criteria For The MVP
- An OpenAI SDK can call an authorized public alias by changing only its base URL and credential.
- The same request can route to at least two different provider types while returning the same public Chat Completions shape.
- Buffered and streaming tool calls preserve IDs, names, and JSON arguments.
- Fallback works before output and is proven not to run after output begins.
- Provider credentials, base URLs, raw errors, and hidden model names do not leak to clients, logs, metrics, or module inspection.
- Unauthorized callers cannot enumerate or invoke hidden aliases.
- Portal host registration and model policy determine the visible aliases; agents cannot submit a provider URL, credential, or unregistered physical model to bypass them.
- Request, token, output, timeout, concurrency, and attempt bounds fail closed.
- Credential rotation preserves provider-account/quota-group limits and cannot be used to manufacture capacity. Any enabled opaque compatibility route has enforceable request/byte/duration/concurrency and financial envelopes.
- Usage is recorded for success, failure, timeout, and cancellation with an explicit completeness/evidence state.
- A bad
llm-router.ymlreload leaves the last valid runtime active. - Pricing-only and partial routing updates reuse unchanged snapshot subgraphs, publish one generation-consistent root, and remain within configured retired generation/memory bounds.
- Every provider dispatch is represented by one logical request and its ordered physical attempts in the dedicated audit sink. Required audit fails before dispatch when its complete envelope cannot be reserved.
local-durablealiases never dispatch an attempt before its start event reaches the WAL durable watermark; recovery reports a durable start without a terminal event as incomplete and sink replay is idempotent.- The metadata-only MVP stores no prompt, completion, tool argument, or reversible PII in Portal, the audit metadata tables, metrics, or traces.
- Existing MCP and WebSocket routes continue to work through their current handler chains.
- The versioned release benchmark meets the absolute 500-RPS and 5,000-RPS gates and demonstrates throughput and P50/P95/P99 gateway-added latency no worse than the pinned Bifrost build under identical profiles.
- Overload tests show bounded memory and queue wait, prompt
429/503load shedding, and recovery without a residual latency backlog. - Slow or trickle-reading stream clients hit bounded write-progress and absolute-lifetime deadlines without leaking upstream work or permits.
- Allocation and trace assertions prove that the steady-state path reuses provider clients, captures one immutable published root, performs no synchronous control-plane I/O, and creates no work for disabled hooks.
- Request-path assertions prove there is no control-plane lock acquisition, policy-map merge, provider string lookup, or parse-failure fallback to an ungoverned detect/opaque path.
References
- Bifrost repository
- Bifrost versus LiteLLM benchmark
- Bifrost benchmark methodology
- Bifrost reproducible benchmark tooling
- Bifrost overview
- Bifrost drop-in replacement
- Bifrost retries and fallbacks
- Bifrost governance routing
- LiteLLM repository
- LiteLLM documentation
- LiteLLM benchmark methodology
- agentgateway repository
- agentgateway configuration architecture
- agentgateway LLM compatibility parsing
- agentgateway LLM route and format types
- agentgateway provider dispatch and streaming translation
- agentgateway model and virtual-model router
- agentgateway HTTP and LLM request path
- agentgateway streaming guardrail state machine
- agentgateway pricing catalog snapshots
- agentgateway runtime stores
- agentgateway request-time policy merging
- OpenAI API reference
- OpenAI Responses streaming events
- PostgreSQL table partitioning
- PostgreSQL row security policies
- SQLite write-ahead logging
LLM Gateway API Contract
Status
- Status: Core API implemented; live production qualification pending
- Date: 2026-08-14
- Scope: public inference APIs, provider adapters, and authentication boundaries
This document defines the stable HTTP contract that agents and applications use
to call llm-gateway. It also defines how the gateway selects a provider wire
protocol and obtains the provider credential after policy and routing have
selected a deployment.
The current Rust implementation supports model listing and retrieval, Chat Completions, buffered and streamed Responses, and buffered embeddings. Live SDK, Codex, provider, multi-replica publication, performance, canary, and rollback qualification remains required before production promotion. The endpoint tables below distinguish the implemented core from optional and deferred profiles.
The terms MUST, MUST NOT, SHOULD, and MAY are normative.
Decisions
- The preferred agent API is the OpenAI Responses-compatible
POST /v1/responsesendpoint. It is the contract used by Codex and other agents that need typed input/output items, tool calls, and event streaming. POST /v1/chat/completionsremains the broad application compatibility API. Existing OpenAI-compatible clients continue to work.- The required public contract is the OpenAI-compatible API family: model listing, Chat Completions, Responses, and embeddings. It is the stable provider-neutral surface for Light-controlled agents and applications.
- Provider-native client facades are optional compatibility profiles, not provider-routing mechanisms. An Anthropic Messages profile is added only when Claude Code or another Anthropic-format client is a certified product requirement. A Gemini profile remains deferred until a Gemini-native client or feature requires it.
- Every request names a governed public alias. A client never supplies a provider URL, physical model ID, route ID, or provider credential.
- Client protocol, canonical operation, provider protocol, and provider authentication are separate types. The selected provider never determines the response contract owed to the client.
- Client authentication and provider authentication are separate trust boundaries. An inbound Light credential MUST NOT be forwarded upstream. A provider-delegated user credential MAY be forwarded only by an explicitly typed, owner-scoped delegated route to that credential’s provider; it is never a Light credential and is never eligible for cross-provider fallback.
- Shared multi-user production routes use provider API or workload credentials. A personal deployment MAY define owner-scoped native session connectors where the provider supports that use. The connector is visible only to its owner and the owner’s agents and is not eligible for a common multi-user route pool.
- The minimum generally available application surface is model listing, Chat Completions, and embeddings. Responses is an additional first-class agent surface, not a replacement that delays those three application endpoints.
- Portability applies only to features represented by the selected client contract and every eligible provider route. Unsupported or lossy conversion MUST fail before dispatch; the gateway does not silently drop a behavior-changing field to manufacture compatibility.
- AWS Bedrock Converse is an upstream provider protocol, not a public client
API. Core OpenAI-compatible requests and an optional Anthropic Messages
facade normalize into the canonical representation before a Bedrock adapter
emits Converse requests. Raw
/conversepass-through is not exposed.
These decisions extend, but do not weaken, the accepted public compatibility ADR. OpenAI Chat Completions remains the first implemented compatibility surface; this document defines the additive target contract.
Goals
- Give agents and applications stable APIs that do not change when routing moves between OpenAI, Anthropic, xAI, Google, or a local provider.
- Support OpenAI-compatible SDKs and Codex through the required core profile.
- Support off-the-shelf clients such as Claude Code through optional, explicitly certified compatibility profiles when product requirements justify them.
- Preserve tools, structured content, reasoning metadata, usage, cancellation, and streaming semantics when both client and selected provider support them.
- Make unsupported conversion explicit and actionable instead of silently dropping fields.
- Keep provider keys, OAuth refresh material, workload credentials, and physical model names inside the gateway deployment boundary.
Non-goals
- The inference API is not a public control-plane mutation API. Alias, deployment, pricing, credential-reference, and routing changes remain event-sourced Light Portal operations.
- The gateway does not execute client-side tool calls. It returns tool calls to the agent, which may execute them through the MCP gateway and submit results in a later model request.
- The initial contract does not promise lossless conversion of every provider-specific feature.
- The gateway is not a complete clone of every provider API. A provider-native client surface is not implemented merely because the corresponding upstream provider is supported behind the OpenAI-compatible core.
- Consumer subscription tokens and CLI credential caches are outside the gateway boundary. Personal workflow automation invokes the provider’s CLI directly; gateway routes use API or workload credentials only.
Architectural Model
agent or application
|
| client protocol + Light credential
v
client adapter -> canonical operation -> policy and alias router
|
v
provider adapter + auth provider
|
| provider protocol + provider credential
v
provider model API
The implementation MUST model these dimensions independently:
| Dimension | Purpose | Initial values |
|---|---|---|
ClientProtocol | Request, response, stream, and error contract owed to the caller | Required: openai_responses, openai_chat, openai_embeddings; optional profiles: anthropic_messages, gemini_interactions, gemini_generate_content |
Operation | Provider-neutral intent used by policy and capability checks | generate, embed, rerank, count_tokens, list_models, get_result, cancel_result, delete_result |
ProviderProtocol | Wire contract used for the selected upstream | openai_responses, openai_chat, anthropic_messages, bedrock_converse, xai_responses, xai_chat, gemini_interactions, gemini_generate_content, vertex_generate_content |
ProviderProfileType | Credential, transport, and eligibility class of the route | openai, anthropic, aws_bedrock, xai, google_gemini, google_vertex |
ProviderAuthMode | How upstream authorization headers are produced | bearer_secret, x_api_key_secret, aws_bedrock_api_key, aws_sigv4, google_api_key_secret, oauth2_workload, google_adc |
The canonical representation MUST retain typed text, image and document input, tool definitions and calls, tool results, structured-output constraints, usage, finish status, safety results, and provider extensions that policy explicitly allows. A conversion MUST fail before dispatch when a required feature cannot be represented by the selected provider protocol.
Public Base URLs
The OpenAI-compatible base URL is the required public surface. Optional native compatibility profiles use namespaced base URLs so their request, response, stream, error, and model-list contracts cannot be confused with the core.
| Client | Configured base URL | Example effective endpoint |
|---|---|---|
| Codex and OpenAI-compatible agents/apps | https://gateway.example/v1 | POST /v1/responses |
| Claude Code and Anthropic SDKs, when the optional profile is enabled | https://gateway.example/anthropic | POST /anthropic/v1/messages |
| Google Gen AI SDK and Gemini REST clients, when the deferred profile is enabled | https://gateway.example/gemini | POST /gemini/v1beta/models/{alias}:generateContent |
An enabled namespaced path is a client compatibility surface. It does not select an Anthropic or Google upstream. For example, an Anthropic Messages request MAY route to a Google model if the selected alias declares a conformant Messages conversion. Supporting an Anthropic or Google provider behind the core API does not require enabling the corresponding client facade.
Alias policy
The public model value MUST be a governed virtual alias such as
coding-default, fast-chat, or embedding-default. Provider-prefixed names
such as openai/gpt-4o, anthropic/claude-sonnet, or
google/gemini-pro are deliberately not a second routing mechanism.
Provider-prefixed model names are convenient in a developer proxy, but in Light they would expose physical-provider choice, couple applications to a deployment, and let clients bypass alias policy and approved fallback groups. An administrator MAY create an alias whose display name contains a provider word for migration compatibility, but it is still an ordinary governed alias; the prefix has no routing semantics.
Endpoint Contract
The contract is divided into profiles so provider support does not imply an unbounded public API commitment:
| Profile | Requirement | Purpose |
|---|---|---|
core_openai | Required | Stable provider-neutral API for Light-controlled applications, OpenAI-compatible SDKs, and Codex. |
anthropic_messages | Optional | Drop-in Claude Code and Anthropic SDK compatibility after a client conformance gate passes. |
gemini_native | Deferred optional | Drop-in Google Gen AI SDK or Gemini CLI compatibility when a concrete client or native feature requires it. |
retained_results | Deferred optional | Retrieval, cancellation, and deletion after state ownership and retention are designed. |
rerank | Optional extended | Provider-neutral reranking for RAG applications. |
Required OpenAI-compatible core
| Method and path | Status | Canonical operation | Contract |
|---|---|---|---|
GET /v1/models | Required core, implemented | list_models | Return only authorized public aliases in OpenAI model-list format. |
GET /v1/models/{alias} | Required core, implemented | list_models | Return one authorized public alias or an indistinguishable not-found result. |
POST /v1/responses | Required core, implemented; preferred for agents | generate | OpenAI Responses-compatible buffered or SSE generation, including typed items and tool calls. |
GET /v1/responses/{response_id} | Deferred retained_results profile | get_result | Retrieve a stored or background response only when the alias and route support retained results. |
DELETE /v1/responses/{response_id} | Deferred retained_results profile | delete_result | Delete gateway-owned retained response state and request provider deletion where applicable. |
POST /v1/chat/completions | Required core, implemented | generate | OpenAI Chat Completions-compatible buffered or SSE generation. |
POST /v1/embeddings | Required core, implemented | embed | OpenAI-compatible embedding request and response. |
POST /v1/rerank | Optional extended profile | rerank | Cohere/Jina-style reranking after a canonical rerank operation and pricing contract exist. |
POST /v1/responses is the standard agent contract. It MUST support, subject
to alias capabilities:
- string and typed item input;
- system or developer instructions;
- client-side function tools and tool results;
- structured text output;
- reasoning controls and summaries where representable;
- stateless opaque reasoning continuity through the standard Responses reasoning item when the selected route requires it;
previous_response_idonly when retained state is enabled for the alias;- buffered JSON and OpenAI Responses SSE events;
- client cancellation propagated to the active upstream request.
The first release of Responses support MAY require store: false. If retained
responses are not enabled, store: true, previous_response_id, retrieval,
and deletion MUST return unsupported_feature; they MUST NOT be silently
ignored.
Opaque reasoning continuity is client-protocol-specific and does not imply
gateway-owned conversation state. For stateless /v1/responses, the gateway
returns a gateway-sealed encrypted_content value when the upstream turn
produces opaque continuation state; it does not manufacture a reasoning blob
for a response without such state. The caller returns the complete reasoning
item as input on the subsequent turn. The legacy
include: ["reasoning.encrypted_content"] request remains accepted but is not
required. The gateway MUST bind the opaque state to the tenant, public alias,
client protocol, selected deployment, and that deployment’s provider-client
material generation. It MUST reject tampering and cross-route replay and MUST
NOT log or render the provider continuation material as public reasoning text.
Reasoning envelopes use a gateway-wide, host/environment-scoped key set
distributed through the existing credential-reference and local secret
materialization boundary using credentialPurpose: REASONING_SEAL; it is not a
provider-endpoint credential. The immutable values.yml snapshot carries exactly one current key
ID and reference, at most one previous key ID and reference, key-set generation,
and bounded item/count/aggregate limits. It never carries key bytes. Every
serving replica MUST resolve the same key set before acknowledging readiness; a
random per-process boot key is forbidden. New envelopes use the current key.
The previous key decrypts while it remains in the active key-set generation.
Key retirement is generation-based, not time-based. A replacement generation is prepared and resolved by every required replica before fleet promotion. The serving set MUST NOT contain replicas on different active key-set generations. Removing the previous key in a later promoted generation retires it for the entire serving fleet; replica wall clocks do not participate in acceptance.
Rotating the selected provider endpoint credential changes that deployment’s
provider-client material generation and deliberately invalidates its older
reasoning envelopes with reasoning_state_stale; rotation of an unrelated
endpoint does not. The generation is derived from the endpoint authentication
policy and active credential identity/version. Routine refresh of short-lived
SigV4 session credentials does not change it. reasoning_state_stale is
non-retryable for the existing conversation: the caller MUST restart from a new
initial request. Dropping the stale item and resubmitting its tool result does
not create an escape hatch; it fails with reasoning_state_required.
A continuity-required turn containing assistant tool-call history and a tool
result without the associated reasoning item fails before dispatch with
reasoning_state_required. Tampered, replayed, unknown-key, retired-key,
oversized, or over-count state likewise fails before dispatch with a stable
typed error.
A continuity-required alias MAY have multiple deployments for new
conversations. Once sealed state is returned, the subsequent request is pinned
to its bound deployment before preference ordering. If that deployment is no
longer mapped or eligible, the gateway returns reasoning_route_unavailable
without revealing physical identity. A request carrying sealed state is never
eligible for fallback, even when the bound provider fails before output. The
gateway MUST NOT discard reasoning state to start a fresh chain implicitly.
/v1/chat/completions has no portable opaque reasoning-state item. A deployment
that requires reasoning continuity, including for a multi-turn tool exchange,
is ineligible for Chat Completions and MUST fail before provider dispatch. An
optional native client protocol, such as Anthropic Messages, MAY use its own
typed signed/redacted reasoning blocks after independent qualification.
The portable first-release function-tool profile accepts strict when omitted
or false. strict: true is rejected as unsupported_feature until strict
tool-schema preservation is represented as a routed capability across every
eligible provider protocol; it is never silently weakened. The deprecated
Responses user forwarding field is likewise outside the stateless portable
profile because its provider-side safety semantics cannot be preserved across
routes.
For embeddings, provider responses are decoded into canonical finite f32
vectors. The gateway requests provider float encoding, validates declared
dimensions and response bounds, and renders either client float JSON or
gateway-re-encoded little-endian base64. Consequently,
supported_encodings gates the canonical provider-side float seam; client
base64 output does not require provider base64 passthrough.
Durable embedding-space stability
OpenAI-compatible request and response bodies do not identify a vector space.
An embedding alias therefore declares an immutable embeddingSpace contract:
space ID and revision, dimension, normalization, distance metric, and document
input-transform version. Every eligible primary, fallback, and canary must
publish the exact same contract. Matching dimensions alone are insufficient.
Knowledge Base aliases set requireExpectedEmbeddingSpace: true. Their clients
send both x-light-expected-embedding-space-id and
x-light-expected-embedding-space-revision. A missing, partial, malformed, or
mismatched expectation fails before provider dispatch. Successful qualified
calls return x-light-embedding-space-id,
x-light-embedding-space-revision, and x-light-config-generation; ordinary
SDK calls that omit the expectation do not receive the configuration generation.
The gateway pins or injects the contract dimension on required-space aliases.
Budgeted embedding clients may also send
x-light-maximum-billed-cost-micros. The gateway combines that request ceiling
with the alias ceiling and rejects the request before provider dispatch when
its conservative multi-attempt reservation envelope cannot fit. Every
successful embedding response returns x-light-billed-cost-micros from the
gateway’s reconciled pricing ledger; callers must not infer billed cost from
token counts or provider-specific response fields.
Idempotency-Key is currently accepted only as forward-compatible request
metadata; the gateway does not deduplicate embedding dispatch or billing by
that header. Durable ingestion workers must commit vectors with their own
input-hash/job idempotency record. Server-side dispatch deduplication remains a
future extension and must not be assumed by clients.
Query and indexing traffic use separate, network-restricted gateway instances
configured with embeddingWorkloadLane: kb_query and kb_index. A request
header cannot select a lane. Both lanes may use different budgets and capacity,
but must advertise the same embedding-space contract. OpenAI-compatible local
servers such as llama.cpp or Ollama can participate as physical embedding
deployments after endpoint, model, dimension, transform behavior, and the
operator-approved space revision pass the same conformance and drift gates.
Optional Anthropic Messages profile
| Method and path | Status | Canonical operation | Contract |
|---|---|---|---|
POST /anthropic/v1/messages | Optional, planned only for certified clients | generate | Anthropic Messages-compatible buffered or SSE generation. |
POST /anthropic/v1/messages/count_tokens | Optional, client-driven | count_tokens | Count the canonical request using the resolved alias/model tokenizer. |
GET /anthropic/v1/models | Deferred compatibility convenience | list_models | Return authorized public aliases in Anthropic model-list format. |
GET /anthropic/v1/models/{alias} | Deferred compatibility convenience | list_models | Return one authorized alias in Anthropic model format. |
This profile is required only when Claude Code, the Claude Agent SDK, or an
existing Anthropic-format application is explicitly certified as a supported
client. Enabling it does not constrain the selected upstream to Anthropic.
When enabled, the gateway MUST support the headers and streaming events in its
pinned client conformance profile. anthropic-version MUST be validated
against an explicit supported-version list. anthropic-beta capabilities MUST
be allowlisted per alias and MUST NOT be copied upstream blindly.
Claude Code is configured with an Anthropic-format base URL, for example:
export ANTHROPIC_BASE_URL=https://gateway.example/anthropic
export ANTHROPIC_AUTH_TOKEN="$LIGHT_LLM_TOKEN"
ANTHROPIC_AUTH_TOKEN is a Light-issued gateway credential in this setup. It
is not an Anthropic API key. Claude Code sends it as an authorization header;
the gateway authenticates the developer, removes the inbound credential, and
later obtains the selected route’s upstream credential.
Deferred optional Gemini-native profile
| Method and path | Status | Canonical operation | Contract |
|---|---|---|---|
POST /gemini/v1beta/interactions | Deferred gemini_native and retained_results profiles | generate | Create a Gemini Interactions-compatible agent request; buffered, streamed, or background according to declared capabilities. |
GET /gemini/v1beta/interactions/{id} | Deferred retained_results profile | get_result | Retrieve or resume a retained interaction. |
POST /gemini/v1beta/interactions/{id}/cancel | Deferred retained_results profile | cancel_result | Cancel a background interaction. |
DELETE /gemini/v1beta/interactions/{id} | Deferred retained_results profile | delete_result | Delete retained interaction state. |
POST /gemini/v1beta/models/{alias}:generateContent | Deferred gemini_native profile | generate | Gemini GenerateContent-compatible buffered generation. |
POST /gemini/v1beta/models/{alias}:streamGenerateContent | Deferred gemini_native profile | generate | Gemini GenerateContent-compatible SSE generation. |
POST /gemini/v1beta/models/{alias}:embedContent | Deferred gemini_native profile | embed | Generate one embedding in Gemini format. |
POST /gemini/v1beta/models/{alias}:batchEmbedContents | Deferred gemini_native profile | embed | Generate multiple embeddings in Gemini format. |
POST /gemini/v1beta/models/{alias}:countTokens | Deferred gemini_native profile | count_tokens | Count tokens for a Gemini-format request. |
GET /gemini/v1beta/models | Deferred compatibility convenience | list_models | Return authorized public aliases in Gemini model-list format. |
Gemini models remain eligible upstreams for the required OpenAI-compatible
core even while this client profile is disabled. The profile is enabled only
when a Google Gen AI SDK, Gemini CLI, or native-only feature is a certified
requirement. If enabled, the {alias} path component is always a public alias
even though the native Gemini API calls that component a model. The gateway
MUST reject models/ resource names, provider project paths, and physical
model identifiers that do not resolve to an authorized alias.
Gemini Interactions is not used as the gateway’s internal canonical model. It
remains behind both the gemini_native and retained_results profiles because
its background and retained semantics require an explicit state design.
Deferred surfaces
The following APIs require separate capability and storage designs and are not part of the required OpenAI-compatible core:
POST /v1/images/generationsand other image/video generation APIs;POST /v1/audio/transcriptions,POST /v1/audio/speech, and realtime speech APIs;- provider-hosted files, vector stores, caches, and prompt resources;
- asynchronous batch inference;
- provider-hosted managed agents, sandboxes, skills, or environments;
- provider-specific search, code execution, and hosted MCP tools.
They MAY be added later as typed operations. They MUST NOT be exposed through opaque pass-through routes that bypass Light authorization, policy, accounting, or audit controls.
POST /v1/rerank is ahead of media APIs in the roadmap because it has a small,
bounded request/response contract and is directly useful to RAG applications.
It still requires provider-neutral documents, scores, token/cost accounting,
and an alias capability before it can be enabled.
Operational Endpoints
Operational endpoints are not inference endpoints and do not use a model alias. They SHOULD be exposed only on an internal listener or protected management network.
| Method and path | Status | Contract |
|---|---|---|
GET /health | Implemented by light-gateway | Process liveness only; it does not promise that an LLM route is eligible. |
GET /readyz | Planned | Readiness for accepting traffic, including a valid published snapshot; it MUST NOT fail merely because one optional provider is unhealthy. |
GET /metrics | Planned Prometheus compatibility | Bounded-cardinality request, stream, latency, usage, cost, route-health, and error metrics. No prompts, outputs, aliases with unbounded user input, or credential data. |
The existing Light metrics handler and durable LLM audit pipeline remain the authoritative integration points. A Prometheus endpoint is an additional scrape format, not a replacement for accounting or durable audit delivery.
There is intentionally no public POST /v1/gateway/keys. Gateway client keys,
aliases, deployments, budgets, and access policy are control-plane aggregates.
They MUST be created through authorized event-sourced Light Portal commands so
that projections, snapshot export, replay, and audit history stay consistent.
Request Rules
Model alias
- OpenAI Chat, Responses, and embedding requests use the
modelbody field. - An enabled Anthropic Messages profile uses the
modelbody field. - An enabled Gemini GenerateContent or embedding profile uses
{alias}in the path. Gemini Interactions uses themodelfield when the interaction is model-backed; managedagentresources are deferred. - The alias is resolved against the request’s host, environment, subject, operation, and current immutable routing snapshot.
- Responses MUST echo the requested public alias, not the physical provider model name, unless a protocol explicitly requires a distinct field. Physical names remain internal telemetry with restricted access.
Streaming
The gateway owes the caller the selected client protocol’s stream:
- Responses: named SSE events such as
response.output_text.deltaand a terminal response event; - Chat Completions:
data:chunks ending in[DONE]; - Anthropic Messages, when enabled: Anthropic message/content block SSE events;
- Gemini GenerateContent, when enabled: Gemini SSE response objects;
- Gemini Interactions, when enabled: Gemini interaction events with resumable event IDs when retained state is enabled.
Provider events are decoded and re-encoded; they are not copied as arbitrary bytes across different protocols. After semantic output begins, the gateway MUST NOT retry or fail over to another provider. Cancellation and disconnect MUST propagate upstream.
On the Anthropic Messages facade, reasoning block ordering is provider-dependent. For Bedrock-backed streams, reasoning blocks are emitted after text because Bedrock provides continuation state at message completion. Gateway round-trips remain supported, but the accumulated message is not guaranteed to be directly replayable to Anthropic’s native endpoint.
Each reasoning block is emitted as a self-contained content_block_start /
content_block_stop pair after any open text block has been closed, so at most
one content block is open at a time. The gateway does not buffer text deltas to
place reasoning blocks first; a strict-order buffering mode remains a possible
future per-deployment option. This non-buffering behavior is deliberately
pinned by regression tests.
Headers
Authorization: Bearer <Light credential>is the canonical inbound authentication form.- An enabled Anthropic facade MAY accept
x-api-keyfor SDK compatibility, but the value is a Light-issued credential, not a provider key. - An enabled Gemini facade MAY accept
x-goog-api-keyfor SDK compatibility, but the value is a Light-issued credential, not a Google provider key. traceparent,tracestate, and the Light correlation header MAY be accepted according to the common handler chain.- Provider-specific beta, organization, project, account, and routing headers MUST NOT be forwarded unless a typed, per-capability allowlist permits them.
- All inbound Light credential headers MUST be stripped before provider dispatch. Raw inbound headers are never copied generically.
The gateway returns x-request-id on every response and SHOULD also return the
client protocol’s conventional request ID header where it differs.
Error Contract
Internally, every failure maps to a stable GatewayError category. The client
adapter renders that category in the caller’s native error envelope.
| Internal code | Typical HTTP status | Meaning |
|---|---|---|
invalid_request | 400 | The request does not conform to the selected client protocol. |
unknown_alias | 404 | No authorized alias is visible to the caller. |
unsupported_feature | 400 | The alias or selected conversion cannot preserve a requested feature. |
authentication_failed | 401 | The Light client credential is absent or invalid. |
access_denied | 403 | The authenticated subject cannot invoke the alias/operation. |
budget_exceeded | 429 | A request, token, cost, or organizational budget rejected admission. |
no_eligible_route | 503 | No active, priced, credentialed, healthy route can serve the operation. |
provider_auth_failed | 502 | The selected upstream credential was rejected. Operators receive the route-safe diagnostic. |
provider_rate_limited | 429 or 503 | The selected upstream quota is exhausted; retry metadata is sanitized. |
provider_unavailable | 502 or 503 | The upstream failed before semantic output began. |
deadline_exceeded | 504 | The request exceeded its effective deadline. |
stream_interrupted | protocol terminal event | Upstream failed after semantic output began. |
Errors MUST include the request ID and an actionable, sanitized message. They
MUST NOT contain provider credentials, raw credential references, private
provider response bodies, or physical route details. A bare
GENERIC_EXCEPTION or “failed without an error response” is not a conformant
public error.
Provider Adapter Contract
A provider adapter is selected only after alias authorization and route eligibility have succeeded. It owns:
- canonical request validation for its protocol;
- conversion to the physical provider request;
- provider authentication headers;
- buffered and streaming response decoding;
- usage and finish-state normalization;
- typed provider error classification;
- cancellation and deadline propagation;
- a declared capability set used before dispatch.
The adapter MUST NOT read a client-supplied provider name, URL, or provider credential. The provider base URL must be validated control-plane configuration and must pass the existing SSRF and authority controls.
Supported provider profiles
| Provider profile | Provider protocol | Default upstream base | Production authentication | Notes |
|---|---|---|---|---|
openai | openai_responses, with openai_chat compatibility | https://api.openai.com/v1 | Authorization: Bearer from an OpenAI Platform API-key secret reference | Shared or owner-scoped server-to-server route using Platform API billing. |
anthropic | anthropic_messages | https://api.anthropic.com | x-api-key from an Anthropic Console secret reference, or short-lived bearer token from approved workload identity; fixed anthropic-version | Direct Claude API. Cloud-hosted Claude needs a separate Bedrock, Vertex, or other cloud adapter because IAM and wire contracts differ. |
aws_bedrock | bedrock_converse | https://bedrock-runtime.{region}.amazonaws.com | Authorization: Bearer from an AWS Bedrock API-key secret reference for development/evaluation, or SigV4 from an AWS workload identity for production | The endpoint owns the AWS runtime/signing region; the deployment owns the physical model or inference-profile ID. For example, a US cross-region inference profile may use us.anthropic.claude-sonnet-4-6; clients still send only a Light alias. |
xai | xai_responses, with xai_chat compatibility | https://api.x.ai/v1 | Authorization: Bearer from an xAI API-key secret reference | Grok supports Responses and Chat Completions. Prefer Responses for agent routes. |
google_gemini | gemini_interactions, gemini_generate_content | https://generativelanguage.googleapis.com | x-goog-api-key from a Gemini API-key secret reference | Developer API upstream profile. Supporting it behind the OpenAI-compatible core does not enable the optional Gemini client facade. |
google_vertex | vertex_generate_content | validated regional or global Vertex AI authority | Short-lived OAuth bearer token obtained through ADC or workload identity | Production Google Cloud profile. The gateway refreshes tokens; Portal stores configuration and references, not access tokens. |
Optional mTLS is a transport property layered on the provider profile. For example, xAI mTLS still requires its bearer API key. Certificate references must use the same secret-materialization boundary as other provider secrets.
AWS Bedrock Converse adapter
The Bedrock profile uses the Bedrock Runtime Converse and ConverseStream
operations. It is distinct from both the direct Anthropic Messages provider
protocol and the optional public Anthropic Messages facade.
- The provider endpoint owns the AWS runtime/signing region and validated Bedrock Runtime authority. The provider deployment owns the physical model or inference-profile ID. None may be supplied or overridden by the client.
- The adapter converts canonical system content, messages, tool definitions, tool choice, tool use, tool results, inference parameters, stop reasons, and usage to and from their typed Converse equivalents.
- Conversion is explicitly fallible. A required field or semantic that cannot
be represented by Converse MUST return
unsupported_featurebefore dispatch; it MUST NOT be silently discarded. ConverseStreamevents are decoded into canonical stream events and then encoded into the selected client protocol. Raw AWS event-stream frames are never exposed to clients.- Model IDs and inference-profile IDs are account-, region-, and availability-dependent deployment data. They are not a static global model catalog and must be live-qualified for the target AWS account and region.
- The AWS Mantle-compatible
/anthropic/v1/messagesendpoint is a separate upstream protocol. It MUST NOT be substituted for Converse merely because a model appears in the AWS catalog; it requires its own provider adapter and live qualification before use.
For the initial us-east-1 qualification, Claude Sonnet 4.6 through the US
inference profile is the baseline text and native tool-use target. Catalog-only
Claude 5 entries remain ineligible until the configured account can invoke
them successfully. An always-on-reasoning deployment is additionally eligible
only for client protocols that can round-trip its opaque continuation state;
OpenAI Chat is not such a protocol.
Provider Authentication
Shared production routes
Shared routes MUST use credentials intended for server-to-server API access:
- OpenAI Platform API key for OpenAI models;
- Anthropic Console API key or approved workload-identity bearer token for the direct Claude API;
- AWS Bedrock API key for development/evaluation routes, or SigV4 credentials from an IAM role or other approved workload identity for production Bedrock routes;
- xAI API key for Grok;
- Gemini API key for the Gemini Developer API;
- Google ADC, service-account impersonation, or workload identity for Vertex AI.
Static values are loaded only through a local secret reference such as
env:OPENAI_API_KEY; they are never published in the control-plane snapshot.
Refreshable auth modes produce request headers at dispatch time and refresh
before expiry without changing the published route generation.
Personal CLI automation boundary
Codex, Claude Code, Gemini CLI, and similar tools may authenticate with a personal subscription. Those sessions represent an individual product entitlement and are not provider API credentials. Light Gateway MUST NOT load, store, delegate, or proxy those sessions.
A personal workflow may invoke each supported CLI directly in its documented non-interactive or structured-output mode. The workflow owns process isolation, prompt and result conversion, tool execution, and retrying a task with another CLI. Such a retry is a workflow decision, not gateway route fallback, because it changes the agent runtime and subscription principal.
The same workflow may call Light Gateway when it wants API-backed routing. Those routes use configured API keys or workload credentials and may fail over between providers only under the normal capability, policy, accounting, and pre-output fallback rules.
Configuration Model
The following YAML is illustrative target configuration. The event-sourced Portal model remains authoritative; projection rows MUST be produced from events and secret values remain local to the gateway instance.
reasoningSeal:
state: active
keySetGeneration: 1
current:
keyId: reasoning-seal-2026-08
credentialRef: env:LLM_REASONING_SEAL_KEY
previous: null
limits:
maxEncodedItemBytes: 131072
maxDecodedProviderStateBytes: 98304
maxItemsPerRequest: 8
maxCumulativeEncodedBytes: 262144
maxCumulativeDecodedBytes: 196608
providerProfiles:
openai-primary:
providerType: openai
protocol: openai_responses
baseUrl: https://api.openai.com/v1
scope: shared
auth:
mode: bearer_secret
secretRef: env:OPENAI_API_KEY
anthropic-primary:
providerType: anthropic
protocol: anthropic_messages
baseUrl: https://api.anthropic.com
auth:
mode: x_api_key_secret
secretRef: env:ANTHROPIC_API_KEY
headers:
anthropic-version: "2023-06-01"
bedrock-us-evaluation:
providerType: aws_bedrock
protocol: bedrock_converse
baseUrl: https://bedrock-runtime.us-east-1.amazonaws.com
region: us-east-1
auth:
mode: aws_bedrock_api_key
secretRef: env:AWS_BEARER_TOKEN_BEDROCK
bedrock-us-production:
providerType: aws_bedrock
protocol: bedrock_converse
baseUrl: https://bedrock-runtime.us-east-1.amazonaws.com
region: us-east-1
auth:
mode: aws_sigv4
service: bedrock
xai-primary:
providerType: xai
protocol: xai_responses
baseUrl: https://api.x.ai/v1
auth:
mode: bearer_secret
secretRef: env:XAI_API_KEY
gemini-developer:
providerType: google_gemini
protocol: gemini_generate_content
baseUrl: https://generativelanguage.googleapis.com
auth:
mode: google_api_key_secret
secretRef: env:GEMINI_API_KEY
gemini-vertex:
providerType: google_vertex
protocol: vertex_generate_content
baseUrl: https://aiplatform.googleapis.com
project: example-project
location: global
auth:
mode: google_adc
scopes:
- https://www.googleapis.com/auth/cloud-platform
A deployment binds one provider profile to a physical model and declared
capabilities. A public alias binds policy and pricing to one or more eligible
deployments. API clients see only the alias. The persisted control-plane shape
routes through deployment aggregates rather than directly from an alias to a
provider profile. For Bedrock, physicalModelId may be a foundation-model ID
or an inference-profile ID; it belongs on the deployment, never in a client
request or global reference value that assumes universal account availability.
Agent and CLI Profiles
Codex CLI
Codex can use the gateway as a custom Responses provider. The gateway token is supplied through a dedicated environment variable or a command-backed token helper, not through the user’s OpenAI provider key.
model = "coding-default"
model_provider = "light_gateway"
[model_providers.light_gateway]
name = "Light LLM Gateway"
base_url = "https://gateway.example/v1"
wire_api = "responses"
env_key = "LIGHT_LLM_TOKEN"
Codex subscription authentication is not forwarded through this profile. A workflow that wants to use the personal Codex subscription invokes Codex CLI directly; a Codex CLI configured as a Light Gateway client uses the Light credential above and consumes an API-backed gateway route.
Claude Code
Claude Code requires the optional Anthropic Messages profile because it speaks the Anthropic gateway protocol. Light MUST advertise Claude Code compatibility only after the pinned client conformance gate passes. The gateway must then keep pace with documented required headers, stream events, beta headers, and message fields. Pointing Claude Code at a gateway credential replaces subscription billing for that session; the selected upstream account is billed. If Claude Code is not a committed product client, this profile remains disabled and creates no obligation to expose Anthropic-format endpoints.
Grok applications
Grok applications use the canonical OpenAI-compatible base URL and select a
public alias routed to an xAI deployment. No Grok-specific client path is
needed because xAI supports Responses and Chat Completions. The client receives
OpenAI-compatible output while the provider adapter authenticates to xAI with
the route’s XAI_API_KEY reference.
Gemini applications
Light-controlled applications use /v1/responses, /v1/chat/completions, or
/v1/embeddings with a Gemini-backed alias; no Gemini public client path is
needed for that routing. A Gemini-native client uses the /gemini base URL and
a Light-issued credential only after the optional profile is enabled and its
client conformance gate passes. Vertex AI remains an upstream deployment
profile, not a different required public client API.
Capability and Conversion Rules
Every deployment publishes a verified capability set. Route eligibility is the intersection of alias policy, requested client features, canonical operation, provider capabilities, credential readiness, price readiness, health, and environment.
At minimum, generation capabilities distinguish:
- buffered and streaming output;
- text, image, audio, document, and video input;
- client-side function tools and parallel tool calls;
- structured JSON output;
- reasoning controls, public summaries, and client-protocol-specific opaque continuation state;
- retained response/interaction state;
- prompt caching controls;
- safety configuration and safety-result visibility;
- exact usage and provider cost reporting.
Unknown client fields may be preserved only for bounded same-format forwarding
under an explicit compatibility allowlist. Cross-format conversion uses typed
canonical fields. Required or behavior-changing fields that cannot be mapped
cause unsupported_feature before provider dispatch.
Protocol conversion SHOULD use an explicit fallible codec or Rust TryFrom
implementation with structured conversion errors. An infallible From
implementation is appropriate only where every source value has a valid,
semantically equivalent target representation.
Rust Implementation Alignment
The generic recommendation to start with Axum is sound for a new standalone
service, but light-gateway is not a greenfield Axum application. It already
uses Pingora listeners, the ordered Light handler chain, shared correlation and
security handlers, and a compiled LLM runtime. The API work MUST extend that
path instead of introducing a second HTTP server or middleware stack.
- Reuse preconstructed provider clients and connection pools from the compiled runtime snapshot; do not construct an HTTP client per request.
- Represent buffered and streaming results with async streams and typed codec events. Provider SSE is decoded incrementally and encoded into the client protocol without buffering the entire completion.
- Treat Bedrock
ConverseStreamas a typed AWS event stream rather than SSE. Decode its content-block, tool-use, metadata, and terminal events into the same canonical stream consumed by the OpenAI and optional Anthropic client encoders. - Use typed
serderequest models. Unknown fields are not globally lenient: they may enter only the existing bounded compatibility envelope for approved same-format forwarding. A malformed known field is a terminal parse error. - Normalize provider usage into canonical input, output, cached, reasoning, and
total token fields before rendering OpenAI
prompt_tokens/completion_tokens, Anthropicinput_tokens/output_tokens, or Gemini usage metadata. - Preserve the existing handler-chain order so authentication, authorization, admission limits, policy, accounting, audit, and provider dispatch cannot be bypassed by a new compatibility path.
Delivery Plan
- Contract foundation: generalize
ClientProtocol,Operation,ProviderProtocol, capability validation, and provider auth without changing the existing Chat Completions behavior. Keep client protocol and upstream provider protocol independently selectable. - Required application core: add
GET /v1/models/{alias}andPOST /v1/embeddings, with operation-specific capability, pricing, accounting, audit, and provider conformance gates. - Responses and Codex: add
POST /v1/responses, Responses SSE, OpenAI and xAI Responses adapters, and a Codex CLI smoke test. - AWS Bedrock provider: add the
aws_bedrockprofile, API-key and SigV4 auth providers,Converse/ConverseStreamcodecs, inference-profile routing, and buffered, streaming, usage, error, and tool-use conformance gates. Route the existing OpenAI-compatible core to Bedrock before adding another public client facade. - Optional Claude profile: only when Claude Code is a committed client, add namespaced Messages, required token counting, Anthropic SSE, and pinned Claude Code conformance fixtures. Prove that the client facade can route to both direct Anthropic and Bedrock without changing its public contract. Keep the profile disabled otherwise.
- Optional Gemini profile: only when a Gemini-native client or native-only feature is committed, add the smallest GenerateContent, streaming, embedding, token-counting, and model-list surface required by its pinned conformance suite.
- Optional retained state: add Responses retrieval/deletion and Gemini Interactions only after retention ownership, route affinity, deletion, encryption, expiry, and audit rules are implemented.
- Optional rerank: add the provider-neutral rerank operation only after document limits, score semantics, pricing, accounting, and conformance are frozen.
- Workflow integration boundary: document and test that personal CLI sessions remain in workflow-owned adapters while Light Gateway provider profiles accept API keys or workload credentials only.
Acceptance Criteria
- Official Codex CLI can complete a tool-calling turn through
/v1/responsesusing a Light-issued bearer credential. - Provider configuration rejects personal subscription sessions, CLI credential caches, and delegated consumer credentials as provider auth.
- Official OpenAI SDKs can call OpenAI-, Anthropic-, xAI-, and Gemini-backed aliases without seeing a physical provider model.
- Official OpenAI-compatible clients can complete buffered, streaming, and
native tool-use turns against a Claude Sonnet 4.6 alias backed by Bedrock
Converse in
us-east-1, without seeing the AWS region, physical model, or inference-profile ID. - Bedrock API-key and SigV4 modes have separate authentication tests. Inbound Light credentials are never used as AWS credentials, and AWS credentials or signing material are absent from snapshots, logs, errors, audit payloads, and client responses.
- The required core passes model-list, Chat Completions, Responses, and embeddings conformance without enabling either native client facade.
- If
anthropic_messagesis enabled, official Claude Code completes the buffered, streaming, tool-use, and required token-counting flows in the pinned conformance profile through/anthropic/v1using a Light-issued credential. - If
gemini_nativeis enabled, the pinned Google Gen AI SDK or Gemini CLI fixtures call the advertised/geminisurface with explicit Light authentication headers. - Inbound gateway credentials are proven absent from all recorded upstream requests; provider credentials are proven absent from logs, errors, audit payloads, and client responses.
- Representative accepted and rejected payloads are parsed and validated for each client/provider pair; tests assert semantic output, errors, streaming order, tool-call identity, usage, and cancellation rather than text fixtures alone.
- A requested feature that cannot survive conversion fails before dispatch
with
unsupported_featureand an actionable message. - A route is ineligible when credential, pricing, capability, environment, or
health data is missing, with
no_eligible_routeexplaining the missing category without revealing secrets. - Existing Chat Completions and model-list qualification gates remain green.
- A disabled optional profile registers no public route and adds no request-path task, lookup, allocation, provider restriction, or fallback behavior.
Provider References
- OpenAI Responses API
- Codex custom model providers
- Codex authentication
- Claude API overview and authentication
- Claude Code gateway guidance
- Claude OpenAI SDK compatibility and limitations
- Amazon Bedrock Claude Sonnet 4.6 model card
- Amazon Bedrock Converse API
- Amazon Bedrock inference profiles
- Amazon Bedrock Anthropic Messages API
- Amazon Bedrock Runtime endpoints
- xAI inference API
- xAI API-key authorization
- Gemini API reference
- Gemini gateway integration trade-offs
- Gemini Interactions API
- Google Gen AI SDK custom base URL
- Vertex AI Gemini quickstart and ADC
Codex CLI Provider Option
Status: deferred design reminder, September 11, 2026. No provider implementation, route enablement, or API conformance is implied. The immediate priority is the Claude Personal Worker; revisit this option after that worker is qualified.
Purpose And Candidate Architecture
Explore whether an owner-only local connector can expose a deliberately limited LLM Gateway API surface backed by the official Codex harness. This would let a client use a standard request envelope for supported text-generation operations without implementing native harness integration itself.
Client -> llm-gateway -> owner-bound local connector
-> pinned Codex App Server -> native vendor service
Prefer the documented App Server interface of the Codex CLI over terminal scraping or repeatedly parsing human-readable output. The connector owns the native process and local authentication context; the gateway owns inbound Light authentication, alias policy, request validation, and the public response format. Native subscription credentials never enter gateway requests or credential stores.
This is separate from a coding worker. The coding worker deliberately edits and tests repositories under a runner lease. An inference connector must have no repository, shell, MCP, plugin, or other native execution authority. If that restriction cannot be enforced, use a workflow coding action instead.
Existing Building Blocks And Limits
The qualified codex-app-server-v1 worker provides useful lifecycle, cancellation,
version-pinning, and event-parsing experience. Reuse appropriately factored code,
but issue a separate connector contract and qualification record.
crates/model-provider/src/codex.rs is an HTTP provider, not a Codex CLI adapter.
Its name does not establish support for this option. Do not reuse direct native
subscription-token handling as a substitute for the official harness boundary.
The App Server speaks a harness protocol; it is not itself the public Responses
API. See Codex App Server and
authentication. Recheck both before
implementation; native login support does not by itself authorize a shared
subscription-backed inference service.
Proposed Initial Contract
Start with text-only requests, one in-flight operation per owner-scoped capacity slot, and explicit governed aliases. Do not silently substitute models or fall back to paid API routes. Each connector is visible only to its owner and the owner’s authorized agents, never a shared provider pool.
- Preserve supported message roles and instruction precedence; reject any conversion requiring lossy concatenation.
- Reject client tools, embeddings, images, reasoning controls, structured output, storage, or other unsupported options before invoking Codex.
- Choose one public API subset for the first probe. A Chat Completions response envelope cannot stand in for Responses or Anthropic Messages conformance.
- Generate a unique public response ID for every operation, distinct from the
native thread ID. Start stateless; reject
previous_response_iduntil owned state, branch, retention, retry, and replay semantics are qualified. - Map streaming only after item ordering, deltas, terminal outcomes, usage, and cancellation pass client-specific fixtures. Otherwise advertise buffered only.
- Record native usage as advisory. Do not invent billable cost, remaining quota, or subscription entitlement from token counts.
Function calling is a separate milestone. A native approval request authorizes
Codex to execute a tool; it is not equivalent to returning a function invocation
for the client to execute. Do not fabricate compatibility by translating an
approval into a generic request_permission tool call.
Decisions And Gates Before Work Starts
The gateway API contract mentions potential owner-scoped native connectors but explicitly excludes CLI credential caches and places personal CLI automation outside the gateway. Resolve that architectural boundary in an ADR before implementing this option; this note does not amend it.
Required gates are vendor eligibility for the intended use, enforced absence of native tool authority, pinned binary/protocol compatibility, owner isolation, message fidelity, public API subset conformance, bounded output, cancellation, process-tree cleanup, secret-free audit evidence, and quota/error behavior. Live qualification must use the intended authentication class and must not treat skipped subscription tests as passing evidence.
Open questions: which client actually needs this subset, whether its agent loop requires function calling, whether native instruction semantics preserve that client’s contract, and whether startup overhead justifies a resident connector. A resident process requires resource limits and idle eviction; it must not create hidden conversational history for otherwise stateless requests.
Related option: Claude Code CLI Provider.
Claude Code CLI Provider Option
Status: deferred design reminder, September 11, 2026. No provider implementation, route enablement, or API conformance is implied. Implement and qualify the Claude Personal Worker first. This option is a future feasibility study, not part of that worker’s delivery scope.
Purpose And Candidate Architecture
Explore an owner-only local connector that runs the official Claude CLI behind a limited LLM Gateway API profile. The intended benefit is standard client requests for supported text-generation tasks, while the local harness retains its native authentication boundary.
Client -> llm-gateway -> owner-bound local connector
-> pinned Claude CLI -> native vendor service
The gateway handles Light authentication, governed aliases, validation, and API response semantics. A supervised local connector owns the process and native login context. No subscription token is collected, copied, or forwarded by the gateway. A shared multi-user subscription pool is outside this proposal.
The coding worker intentionally owns repository tools. This inference connector must disable and externally constrain repository access, shell, MCP, hooks, plugins, and other side effects. Tool denial cannot rely on a prompt. If a task needs native repository execution or human approvals, use the coding-worker contract instead of presenting it as ordinary inference.
Existing Building Blocks And Transport
crates/model-provider/src/claude_code.rs already wraps Claude CLI, but its agent
path bypasses permissions, buffers output, and flattens conversation roles into
text. It does not establish Responses, Messages, or external-tool conformance.
Do not promote that wrapper unchanged as this connector.
Prototype bounded Tokio subprocess pipes with -p and structured output. A
resident process using --input-format stream-json is an optional optimization
after its input envelopes, per-turn outcomes, cancellation, and idle behavior
are qualified. Do not use PTY prompt matching or write y to approve tools.
The public CLI reference and
headless guide are qualification inputs;
recheck the exact supported flags against the selected release.
Native --model selection must come from a governed alias mapping and the user’s
entitlements. Unknown models fail explicitly. Subscription execution must not
silently switch to an API key or paid fallback. Configuration containment and
native login must work together; --bare is not a personal-login solution.
Proposed Initial Contract
Begin with an explicitly advertised text-only subset of one public API. Merely
wrapping the CLI’s final text in JSON does not implement /responses or
/messages.
- Preserve supported system/instruction and message-role semantics. Reject unsupported role combinations rather than serialize them into a user prompt.
- Reject client tools, images, embeddings, structured output, reasoning controls, storage, and any other unqualified fields before dispatch.
- Stateless Messages requests must not inherit hidden native conversation state. Do not combine full client history with automatic CLI resume and duplicate the conversation. Choose an explicit state model for each client profile.
- Responses continuation requires owner-bound public-response/native-session
mapping, branch semantics, retention, idempotency, and crash recovery. Until
qualified, reject
previous_response_idand related stateful features. - Keep request/response IDs separate from native
session_id. Parse the pinned result schema; do not assume example fields such assessionIdortext. - Qualify buffered completion first; add streaming only with correct API-specific event ordering, content blocks, completion reasons, usage, and cancellation.
- Bound queues, deadlines, frames, total output, retries, and process lifetime. Keep usage advisory and surface exhaustion without automatic paid fallback.
Returning a synthetic permission function call does not support generic client tools: permission authorizes harness execution, whereas a tool-result message reports client execution. A generic agent requiring function calling will need a separately qualified external-tool bridge or an ordinary API-backed provider.
Eligibility And Architectural Decision
Anthropic’s subscription support update says the proposed credit-pool changes were paused. Its credential-use guidance also restricts third-party subscription integration. Do not infer authorization for this gateway from CLI technical feasibility or the token remaining local. Record an eligibility decision for this specific use before supported deployment; recheck current guidance rather than freezing either page’s wording into policy.
The gateway API contract contains both a potential owner-scoped native-connector allowance and an explicit CLI boundary. A separate ADR must reconcile those statements and define any narrowly scoped exception. This reminder does not change the accepted gateway contract.
Qualification And Revisit Criteria
Revisit after the Claude worker is qualified and a named client needs the supported inference subset. Prove no native tool side effects, native-login and configuration containment, owner isolation, alias/model binding, message fidelity, correct errors, process cleanup, cancellation, output bounds, and sanitized telemetry. Pin the binary and fixture schemas; qualify the actual public client API rather than just a successful native prompt.
Measure latency and memory before adopting resident processes. Their reuse must not leak conversation or permissions between requests. Missing history or an uncertain execution must never cause silent duplicate generation.
Related option: Codex CLI Provider.
LLM gateway configuration ownership and publication
Status: Proposed design. This document describes intended behavior, not completed implementation.
Problem
LLM Model Control Plane publications generate several llm-router properties in
instance_property_t. The generic Instance Config editor also offers edits to
those rows, but the two write paths do not share an aggregate history.
The observed failure was an agentDelegation row at aggregate version 3 while
its ConfigInstance event stream was at version 1. The generic editor submitted
version 3 and the event store correctly rejected it. Retrying or reducing the
submitted version does not resolve the conflicting ownership.
There is also no dedicated control-plane editor for the endpoint policy inside
agentDelegation. The Publication tab offers a read-only generated preview.
Users should not need to edit a generated JSON object to decide whether direct
user inference is allowed.
Current implementation evidence
These paths are relative to sibling repositories under the workspace:
| Component | Current behavior |
|---|---|
light-portal/db-provider/.../persistence/LlmModelPersistenceImpl.java, applyInstancePublication | Applies a publication to multiple instance properties and increments their aggregate_version directly; records publication ownership. |
light-portal/db-provider/.../persistence/AgentGatewayProjection.java, compile | Derives bindings from current Agent publications and client registrations; reads issuer/audience and endpoint choices from the previous generated instance property; returns null when no user issuer is available. This generated-output input must be removed after migration. |
light-portal/db-provider/.../persistence/LlmModelPersistenceImpl.java, publication revision lookup | Currently deduplicates by host, environment, and property-set digest only, so changed source provenance can reuse an older revision. |
light-fabric/apps/light-gateway/src/main.rs, security handler selection | Looks up the trusted endpoint key in agent_authorization.endpoints; absent entries retain the legacy handler chain, while a selected profile colliding with HMAC returns 503. |
portal-view/src/pages/genai/llm-model/PublicationPanel.tsx | Generates a read-only preview and publishes it to an instance. |
light-portal/db-provider/.../persistence/GlobalSnapshotPersistenceImpl.java | Maps instance_property_t to ConfigInstance-created events and includes LLM publication ownership data in snapshot support. The import ordering and duplicate-materialization behavior need explicit qualification. |
light-fabric/crates/llm-gateway/src/authorization.rs, authenticate | On a selected dual-token route, a required workload token that is absent produces 401 workload_token_required. Optional workload authentication still validates the user and any supplied workload token. |
Decision
Use one authoritative control-plane write path for managed configuration. Keep one complete publication event as the deployment revision, and treat its instance properties as generated projections. Add a typed Agent Delegation editor and make ownership explicit in the generic configuration UI and backend.
A domain event may project into many rows. Generating a separate event for each row is not necessary for event sourcing. Conversely, multiple property events do not establish which system owns a value or prevent the next publication from overwriting a generic edit.
No automatic reverse synchronization from arbitrary generated JSON is proposed.
The current previous.endpoints fallback and previous issuer/audience inputs are
a legacy reverse-synchronization channel, not the intended ownership contract.
Remove them from normal compilation when authoritative authoring is enabled;
reading generated output is allowed only in the explicit validated bootstrap.
No new agent-to-model assignment store is introduced: existing internal aliases
and boundPrincipal remain authoritative for model binding.
Alternatives
| Approach | Benefits | Costs and limitations |
|---|---|---|
| One publication event with owned projections — selected | Coherent revision, transactional projection, straightforward provenance and replay | Requires managed-edit restrictions and ownership-aware export/import |
| Publication emits individual ConfigInstance events | Per-property aggregate histories; generic versioning can be valid | Requires coordinated expected versions, publication correlation, complete-batch activation, and partial-failure recovery; still needs ownership rules |
| Independent Config edits with reverse synchronization | Either interface can modify raw values | Ambiguous mapping to domain records, feedback loops, conflicts, and security-sensitive interpretation of generated bindings |
| Explicit instance overrides | Supports deliberate deployment-specific exceptions | Adds precedence, validation, provenance, and rollback rules; defer until required |
If property events are adopted later, commands must append them through the event store, with a publication manifest and explicit completion barrier. A projection consumer must not emit an uncontrolled cascade of new business events on replay.
Authoritative Agent Delegation policy
Scope and storage
Persist typed endpoint-policy authoring for the target host and gateway instance. The initial scope is the same instance selected for publication; avoid introducing environment-wide inheritance of endpoint requirements in this change. Reuse an existing suitable LLM policy aggregate if the implementation inventory identifies one. Otherwise add one narrowly scoped authoring aggregate, with its own event-store version.
The implementation must document the selected table, aggregate ID convention, event names, and command/query schemas before migration. An authoring aggregate is distinct from derived agent bindings and is not another assignment store.
The authoring record contains:
- Host and gateway instance identity.
- Supported endpoint requirements: whether an agent workload token is required.
- An explicit reference to the host-and-environment user issuer/audience profile.
- Schema version and normal authoring concurrency/version metadata.
Endpoint requirements remain per instance, but the user issuer/audience profile is shared by host and logical environment, matching the current binding compiler’s scope. Reuse an existing suitable security profile or introduce an authoritative profile at that scope. All instance references and Agent policies in that scope must agree; reject conflicts on authoring validation and recheck at publication. The compiler must receive the target instance’s endpoint policy and the referenced environment profile explicitly. It must not infer trust settings from generated output or require an active Agent registration to obtain the user profile.
UI
Add an Agent Delegation tab to LLM Model Control Plane with gateway instance selection and separate editable and derived sections.
| Endpoint | Editable setting |
|---|---|
/v1/chat/completions@post | Require agent workload token |
/v1/responses@post | Require agent workload token |
/anthropic/v1/messages@post | Require agent workload token |
Explain that optional means an authenticated direct-user request may proceed. It does not disable user verification, permit an invalid supplied workload token, or grant direct users access to agent-bound aliases.
Show bindings, client IDs, registration versions, Agent policy digests, and alias restrictions read-only, with navigation to their owning records. Do not provide raw binding edits. Save authoring changes separately from publishing and activating a runtime snapshot. Show unpublished changes and expected runtime consequences.
For the local mixed-use gateway, the intended chat-completions setting is optional: public alias tests use a user credential; Tech Support sends both credentials. This is a deployment choice, not a universal security default.
Validation
Require host-scoped authorization for edits. Check instance ownership, supported endpoints, boolean types, and issuer/audience compatibility. Use optimistic concurrency on the authoring aggregate. Retain strict JWT expiry, audience, issuer, TLS, workload registration, and model-binding checks.
For every authored endpoint, emit a non-null agentDelegation profile containing
that endpoint’s explicit boolean, even with zero active bindings or no previous
publication. Reject preview/publish if the user profile is unavailable or compilation
would emit null or omit an authored endpoint. In particular, required policy must
never disappear into legacy single-token parsing. Removing the last binding must
preserve the profile and endpoint requirements.
Parsing behavior remains selected by trusted route configuration. Never retry with the legacy parser after new-profile verification fails. An optional endpoint remains in the new profile; it is not the same as removing the endpoint entry.
Publication lifecycle
- Read authoritative model, alias, registration, Agent publication, and endpoint policy records into a consistent candidate.
- Derive bindings and the complete managed property set.
- Include authoring versions and compiler-relevant defaults in the source fingerprint. Include all generated values in the property-set digest.
- Present a read-only preview and differences from the target instance.
- On publish, regenerate and check the reviewed digest and source versions. Reject stale previews rather than overwriting concurrent changes.
- Append one publication event containing the validated complete revision and source provenance.
- Apply its generated rows, ownership, and publication status in one projection transaction. Preserve idempotency on replay and transactional failure recovery.
- Create and activate the runtime snapshot only after complete projection. A runtime reload is a separate observable step; publication success alone does not prove that a running gateway has loaded the revision.
Revision identity and reuse must include the source fingerprint as well as host,
environment, and property-set digest; the source fingerprint must include target
instance identity and the referenced profile version. Update the deterministic ID,
reuse query, and manifest contract together. The current manifest contains only
schemaVersion and propertySetDigest, while llm_gateway_publication_t enforces
UNIQUE (host_id, environment, manifest_digest) in portal-db/postgres/ddl.sql.
The new versioned manifest must also contain the source fingerprint, and its digest
must cover that field. This preserves the uniqueness constraint while allowing
identical output from different sources; changing only the revision ID and reuse
query would fail on insert. Preserve historical manifests and digests during
migration and restore. Byte-identical
properties from different authoring versions require distinct provenance-bearing
revisions. Deduplicate retries of the same source and output, not different source
histories; replay preserves the recorded revision identity.
If authoring changes between publication and activation, the reviewed publication remains immutable. Operators can activate that exact revision or publish a newer one. The UI must identify which revision is active and which is pending.
Instance Config behavior and versioning
Display managed properties with the owning product, publication ID/revision, active snapshot, and an Edit in LLM Model Control Plane action. Alternatively, open the same domain editor in place, submitting the same domain command.
Reject ordinary ConfigInstance create/update/delete commands for an actively publication-owned property. Enforce this on the server before appending an event; UI restrictions alone are insufficient. Publication claiming a previously generic property must validate the expected value/version and ownership transactionally.
Separate two meanings currently conflated by aggregate_version:
- Event-store aggregate version: concurrency position of the owning command stream.
- Projection/publication revision: provenance of the generated value.
Do not increment a ConfigInstance event version merely because a publication updates its projection. Use publication ownership metadata for generated revision tracking. Audit query contracts, snapshot comparison, generic editors, and publication fingerprinting before deciding whether the existing column can be retained with narrower semantics or requires a separate field.
Historical generic streams remain historical; managed ownership does not rewrite them. Deliver an explicit ownership-release command for transfer back to generic operation. It must authorize the host/instance, check expected publication ownership and versions, deactivate ownership rows, and establish a valid generic event-stream baseline transactionally. Preserve historical publication provenance. Reject stale release attempts and make replay idempotent. A subsequent publication must explicitly reclaim ownership with the same concurrency checks as an initial claim. Release must not happen implicitly when a publication or authoring record disappears.
Global snapshot export and import
Export must distinguish unmanaged properties from publication-managed output.
Managed restore
Restore authoritative control-plane records, exact published material and provenance, and ownership. Reconstruct managed properties once through the publication restore path. Exclude those same rows from independent generic ConfigInstance-created event generation in this mode.
Do not recompile historical publications with today’s compiler defaults during restore: that could silently change the deployed policy. Preserve the exported revision and digest, validate its supported schema, then allow a later explicit publication to adopt new defaults.
Restore dependencies before publication activation. Imported event stream versions and references must follow the import contract rather than copying a projection counter as an assumed event-store position. Repeated import must not create duplicate publications or multiply property writes.
Detached restore
If a deployment intentionally needs configuration without the control plane, provide an explicit detached mode. Materialize plain configuration, omit active publication ownership, and establish valid generic aggregate histories. Restoring into an already managed target must use the ownership-release contract in the restore transaction so no active ownership rows remain. Explain that subsequent control-plane publication is not automatically synchronized.
Downloaded immutable values.yml remains a valid offline runtime input. This
proposal does not introduce policy expiration or require Portal availability for
an already configured agent or gateway to continue operating.
Legacy snapshots
Define a versioned compatibility path for exports containing both generic property rows and publication ownership. Reconcile them by verified value/digest and provenance; reject conflicting content with an actionable report. Do not silently choose whichever event happens to arrive last.
Development setup and existing data
Apply the database patch before starting the command/query services. The domain editor, ownership-release command, and managed-write protection are part of the same implementation. Protection is always enforced, with no rollout flag.
For existing data, inventory managed properties, ownership records, generic event streams, and projection versions. Bootstrap authoritative authoring from each instance’s last accepted configuration, preserving endpoint requirements and recording provenance through supported domain/import commands. Reject conflicts; do not lower versions or fabricate historical edits. Normal compilation never uses previous generated endpoints or issuer/audience values as authoring inputs.
Publish the desired endpoint policy, activate a new snapshot, and verify public-user and agent-mediated inference. Tests cover source-aware revision identity, ownership release, and managed, detached, and legacy export/import handling.
Rollback retains immutable prior publications and restores an explicitly selected compatible revision. Generic editing requires explicit ownership release.
Verification matrix
| Scenario | Required result |
|---|---|
| Save endpoint choice and regenerate | Candidate reflects authoritative authoring; prior generated JSON cannot override it |
| Two instances in one environment | Independent endpoint choices; shared issuer/audience reference; conflicting profiles rejected |
| Fresh instance or last binding removed | Non-null profile retains every authored endpoint; missing trust profile rejects publication; required route rejects user-only requests |
| Concurrent edits/publications | Stale version or reviewed digest rejected; no lost update |
| Optional endpoint, valid user, public alias | Direct buffered and streaming inference succeed |
| Required endpoint, missing workload token | Request rejected before provider dispatch |
| Invalid supplied workload token on optional endpoint | Rejected; no user-only or legacy fallback |
| User-only request for agent-bound alias | Existing non-disclosing model authorization failure retained |
| Valid dual-token Tech Support request | Correct agent binding, alias authorization, billing identity, and audit provenance |
| Generic edit/delete of managed property | Clear domain-owned error and navigation guidance; no event appended |
| Publication projection failure midway | Transaction rolls back all managed changes; no partial snapshot activation |
| Required → optional → required with identical final output | Final revision records the new source fingerprint; older provenance is not reused |
| Explicit ownership release and retry | Ownership deactivated with valid generic baseline; stale release rejected; replay idempotent |
| Duplicate event/replay | Same values and ownership; no extra revisions or writes |
| Export/import managed publication | Same values/digests and ownership; one materialization path |
| Import then domain edit and republish | Valid versions; no ConfigInstance mismatch |
| Legacy conflicting snapshot | Explicit rejection/report rather than last-writer-wins behavior |
| Detached restore | No residual managed ownership; generic edits have valid event histories |
| Portal/config-server unavailable | Accepted runtime configuration continues without policy renewal |
Run make llm for public-user inference and make genai-chat-ui for the real
browser-to-Agent-to-LLM path in light-portal-test. A browser pass alone does not
prove direct-user access, export/import correctness, or durable audit delivery.
Use transactional PostgreSQL integration tests for projection and restore cases.
Delivery phases
- Contracts and inventory: settle authoring/profile ownership, aggregate and schema changes, managed-version semantics, and import compatibility fixtures.
- Authoring and UI: commands/queries, Agent Delegation editor, source fingerprint, managed-edit guards and ownership release, removal of previous-output compilation inputs, non-null profile validation, and links from Instance Config.
- Projection and restore: migration, atomicity/idempotency tests, version semantics, source-aware revision identity, ownership-aware global snapshot round trips and detached mode.
- Local qualification: publish reviewed endpoint settings, activate/reload, pass direct buffered/streaming and GenAI Chat tests, verify audit delivery, and exercise rollback.
Completion requires both configuration paths to remain consistent after replay and export/import, not merely a successful edit or one successful model request.
Configuration ownership implementation and qualification
The four delivery phases are implemented and locally qualified on 2026-09-08. Managed-write protection is always enforced. See the design.
Authoring contracts
| Aggregate | Table | Stream subject |
|---|---|---|
| LlmGatewaySecurityProfile | llm_gateway_security_profile_t | hostId / securityProfileId |
| LlmGatewayDelegationPolicy | llm_gateway_delegation_policy_t | hostId / instanceId / llm-gateway-delegation-policy |
| LlmGatewayOwnershipRelease | llm_gateway_ownership_release_t | hostId / ownershipReleaseId |
Stream components use the existing Portal pipe delimiter. Security profiles are unique by host and logical environment, with immutable environment identity. They contain userIssuer, userAudience, and schemaVersion 1. Instance policies reference securityProfileId and contain explicit endpoint booleans for all three supported inference routes. Updates require the observed aggregateVersion.
The genai service exposes createLlmGatewaySecurityProfile, updateLlmGatewaySecurityProfile, createLlmGatewayDelegationPolicy, updateLlmGatewayDelegationPolicy, and releaseLlmGatewayOwnership at version 0.1.0. The corresponding list queries are getLlmGatewaySecurityProfile and getLlmGatewayDelegationPolicy; getFreshLlmGatewayDelegationPolicy selects an instance, and getLlmGatewayOwnershipState returns current managed ownership.
Publication requires the preview’s expectedPropertySetDigest and expectedSourceDigest. The instance-properties-v2 manifest includes sourceDigest and propertySetDigest. Historical v1 manifests and digests remain unchanged. Replay policy v4 registers the new event types while retaining registries v1–v3.
Development setup
Apply portal-db/postgres/patch_20260908_llm_configuration_ownership.sql before starting the command/query services. Managed-write protection is unconditional; there is no environment variable or system property to disable it.
Run portal-db/postgres/tests/llm_configuration_ownership_inventory.sql against
the target Portal schema first. It reports accepted material, conflicts, authoring
versions, and generic event-stream heads. Bootstrap only reconciled records;
unmanaged or conflicting values require explicit operator review.
Bootstrap authoring explicitly from each instance’s last accepted profile, preserving endpoint requirements. Save the environment security profile before the instance policy. Normal compilation does not read generated JSON as authoring. An instance without an authoritative policy cannot publish.
Release submits hostId, instanceId, and the expected instancePublicationId. The server emits a release event and generic ConfigInstance baseline events, verifies ownership inside the append transaction, and deactivates ownership atomically. The current runtime snapshot is unaffected.
Snapshot contracts
The snapshot root accepts llmConfigurationMode with managed (default) or detached. Managed mode exports current ownership as metadata folded into publication restoration, excludes duplicate generic property events, and rejects conflicting legacy values. Historical applications use restoreOnly so they do not reactivate old ownership. Managed cross-host restore is rejected because rebinding would change host-bound immutable material.
Detached mode omits publication/authoring state and produces generic configuration. Import into a managed target prepares the explicit release and baseline events in the same append batch, with transactional stale-ownership validation.
Qualification
Run LlmGatewayOwnershipPostgresTest with LLM_OWNERSHIP_TEST_JDBC_URL pointing at an isolated PostgreSQL database. The disposable test database uses test-only postgres credentials; the test creates and removes its own schemas. Qualification requires zero skipped tests.
After publication, snapshot activation, and gateway reload, run make llm and make genai-chat-ui in light-portal-test. Verify durable audit rows and rollback separately. Unit and projection tests alone do not establish runtime activation.
Local qualification evidence
The local llm-gateway instance is 391447f5-fb0d-5c80-92a7-a8d98f6d07c7
in logical environment dev. Its security profile is
01a08267-912b-7816-86e1-7e0118fca834. The bootstrap compared the current generated
profile with its owning accepted publication and retained the publication ID and
profile digest in migration provenance. Endpoint policy version 4 permits direct
user inference on all three supported routes. Agent bindings remain derived.
Verified through authenticated Portal commands and queries:
- Required → optional → required generated identical final property bytes but distinct source fingerprints and immutable revision IDs. A stale preview was rejected before publication.
- Release emitted nine generic baselines, left no active ownership rows, and
reconciled every property version with its event stream. A generic edit then
succeeded; explicit publication reclaimed ownership. With managed-write protection,
another generic write returned
409 LLM_CONFIGURATION_MANAGED. - Publication, snapshot activation, and reload with required chat authentication
returned HTTP 401 to a valid user-only request. Exact rollback to the stored
optional revision returned HTTP 200. Invalid supplied workload tokens returned
HTTP 401 in optional mode. Durable audit records identify
workload_token_requiredandinvalid_workload_tokenrespectively. make llmpassed both buffered and streaming requests.make genai-chat-uipassed a newly accepted turn and a unique response marker through Tech Support. The browser test used a separate qualification session file because the older saved session no longer matched the running Agent’s durable session ownership.- The final snapshot contains all nine publication values, including the current Agent binding. Database audit rows contain successful direct-user and Agent requests, user attribution, and workload attribution for delegated calls.
The local services use the ownership-20260908 image tags. The local-only Compose
file portal-config-loc/all-in-lt/.runtime/llm-ownership.compose.yml selects those
images. Managed-write protection is built into the command service and needs no
Compose setting.
Additional fixes proven necessary by qualification:
- The delegation-policy event subject has an explicit suffix so it cannot collide
with the existing
hostId|instanceIdInstance event stream. - Binding compilation follows
agentPolicy.policySnapshot.snapshotIdin the current configuration snapshot of the same Agent instance. It uses the instance’s logical environment, allowing a different deployment tag, and excludes revoked policy evidence and inactive registrations. - New inference requests receive independent UUID request IDs rather than reusing
caller correlation headers. Older WAL records with non-UUID request IDs receive
stable host/day-scoped UUIDv8 storage IDs. Their original IDs remain in
transport_context.legacyRequestId; WAL records are not discarded or rewritten. This preserves the historical grouping by correlation ID; it cannot separate distinct old calls that reused the same ID. The retained local backlog drained successfully after deployment.
Regression coverage includes the actual exporter/converter/PostgreSQL restore, historical applications arriving after current ownership, replay, conflicting legacy snapshots, detached restore, authoring updates across a repeatable-read snapshot, and append failure rolling back event/outbox writes and nonce reservation with the publication. A generic payload cannot use restore metadata to bypass the managed-write guard. Fourteen focused ownership/restore tests ran with zero skips. Two PostgreSQL audit tests verify duplicate delivery and legacy-ID recovery. UI tests exercise authoring versions, explicit endpoint flags, manual instance selection, and release.
The full canonical DDL loads successfully in a disposable PostgreSQL database.
The repository-wide DDL documentation gate still rejects pre-existing historical
ADD COLUMN statements in ddl.sql; those statements also exist in HEAD.
Review follow-up
The source fingerprint now uses only the selected compiler resource payloads and source resource versions, together with the authoritative policy/profile versions and derived Agent delegation. It excludes row audit metadata, unrelated catalog rows, envelope sequence numbers, and synthetic next-publication versions. No additional source-table scans run while the publication lock is held. Historical publication manifests remain unchanged; refresh an outstanding preview after upgrading the compiler.
Detached imports advance each property’s stream version for every consumed event. Only snapshot restore can supply a revision UUID; normal publication derives its identity. Security profiles use an active-only unique environment index and a 16-character environment limit in SQL, request schemas, and backend validation. The rerunnable migration upgrades the initial unique constraint. Existing values longer than 16 characters must be reconciled before migration; they are not truncated. The canonical definitions precede the dump trailer and include column comments.
Managed-write protection is unconditional, including the host lock for generic configuration commands. Generic edits require explicit ownership release. The obsolete environment flag and its Compose entries have been removed.
Follow-up validation passed 54 focused persistence tests (including 11 PostgreSQL ownership tests) and four request-schema tests, with no skips. The canonical DDL loaded successfully in isolated PostgreSQL and the documentation build passed. These review corrections have not yet been redeployed to the local services.
ADR 0001: LLM Public Compatibility Profile
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 before LF-3
Decision
The first public surface is OpenAI Chat Completions: buffered JSON, SSE, and
GET /v1/models. Models are authorized aliases, never provider deployment
identifiers. Errors use the OpenAI envelope while retaining a typed, sanitized
internal category.
A typed parse failure is terminal. Same-format OpenAI forwarding may retain unknown fields only in a size-bounded compatibility envelope and only for an alias/deployment allowlist. Cross-format OpenAI-to-Anthropic routing uses canonical typed content and rejects unrepresentable fields. No detect-and-opaque fallback is allowed.
The checked-in corpus under benchmarks/llm-gateway/payloads is the Phase 0
compatibility baseline. Account/CLI providers are not eligible for the shared
gateway.
Consequences
LF-3 codecs can define one closed canonical model. Compatibility behavior is observable and cannot turn parse failures into arbitrary upstream forwarding.
ADR 0002: One-Pass Application Body Contract
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 before buffered HTTP integration
Decision
Register one llm application handler and delegate to a typed integration.
The integration runs pre-body handlers, validates route/method/media type,
content encoding, declared length, and deadline, then captures/decompresses one
bounded Bytes body exactly once.
If access control appeared earlier in the selected chain, body-aware
authorization receives that captured byte sequence before LLM JSON parsing,
alias policy, transforms, client selection, or provider work. Parsing and all
later content adapters borrow or clone the same immutable Bytes; they do not
read the downstream stream again. Every error and downstream disconnect
cancels/finalizes the request.
Generic tokenize/detokenize handlers are not assumed to have consumed the body. Content transforms require an explicit LLM adapter.
Evidence
benchmarks/llm-gateway/evidence/body-capture.json records bounded,
chunked, one-pass capture and proves authorization precedes parsing while both
observe the same digest. The production gateway already demonstrates the
relevant ordering in GatewayProxy::request_body_filter; LF-4 must bind this
ADR to the new application handler with an integration test.
ADR 0003: One ArcSwap Runtime Root Per Request
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 before LF-5
Decision
The LLM runtime publishes one immutable Arc<CompiledLlmRoot> through
ArcSwap. Request admission captures the root once and all routing, provider,
policy, pricing, accounting, and client choices come from that Arc. The request
path must not repeat current-config reads.
The reload worker builds and validates a complete candidate off-path, reuses unchanged Arc subgraphs, materializes clients/secrets, and performs one atomic store. A failed candidate leaves the previous root active. Dynamic counters and circuit state have stable identities and are not rebuilt merely because the configuration root changes. Retired roots live until the last in-flight Arc is dropped.
Evidence
benchmarks/llm-gateway/evidence/snapshot.json compares repeated
light_runtime::ConfigManager RwLock reads with a single capture through the
existing ArcSwap-backed config_loader::ConfigManager, and proves the captured
root remains generation-coherent across publication.
ADR 0004: LLM Configuration Uses the Standard Config Lifecycle
- Status: Superseded filesystem projection; accepted values-backed lifecycle
- Date: 2026-08-15
Decision
llm-router uses the same configuration authority and lifecycle as every
other reloadable gateway module. The config server’s current immutable
values.yml snapshot is the only source of LLM routing configuration.
At startup, LightRuntime downloads the current snapshot, resolves
llm-router.yml, compiles the complete providers/deployments/aliases graph,
and publishes one immutable runtime snapshot. During an explicit module reload,
the runtime downloads the current snapshot again and invokes only the selected
reloaders. LlmRouterReloader compiles from that fresh reload context and
atomically swaps the candidate only after validation succeeds. A failed reload
retains the last-known-good LLM runtime.
The former config-server /files manifest/resource projection,
LlmProjectionWorker polling loop, projection checkpoint, and gateway-to-Portal
publication acknowledgement are removed. LLM configuration cannot change merely
because files appeared in config-cache; it changes only at startup or when
llm-router is included in an explicit reload.
Configuration Boundary
The typed llm-router.* properties in values.yml include the complete
provider, deployment, alias, policy-derived, pricing, and non-secret runtime
material configuration. Map and list properties are whole typed nodes, not
quoted JSON strings.
Credential and reasoning-seal values remain outside config server. The snapshot
contains only env: or opaque credential references plus the authorized local
reference-to-environment mapping. Trust-bundle configuration similarly contains
only approved references, paths, and digests; it does not contain private key or
provider credential bytes.
The config snapshot is immutable. Publishing control-plane changes means
updating the target instance’s llm-router properties, creating/promoting a new
snapshot through the normal config workflow, and then explicitly restarting or
reloading llm-router.
Consequences
- Startup and reload observe one coherent config-server snapshot.
- Reloading an unrelated module cannot alter LLM routing.
- In-flight requests retain the immutable runtime snapshot they captured.
- There is no second polling, sequence, checkpoint, or acknowledgement protocol.
- Delivery success is reported by the standard module reload result; provider reachability remains a separate runtime qualification concern.
ADR 0005: Off-Path Secret Materialization
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 before production provider publication
Decision
Portal resources carry only credential:// reference IDs. The first
production resolver seam is the gateway’s already resolved runtime
configuration: config-loader decrypts CRYPT values, RuntimeConfig exposes
resolved values to authorized module construction, and ModuleRegistry masks
sensitive values in inspection output. A later provider integration must
implement the same narrow SecretResolver contract rather than changing the
request path.
Reload performs three stages: parse/validate the secret-free resource graph; authorize and resolve every enabled credential reference and construct reusable clients; publish only the fully materialized root. Resolution, decryption, token exchange, and client construction never occur during inference.
Missing, denied, expired, blank, or malformed references reject the candidate and preserve the last valid root. Runtime config reload is the rotation notification. Rotation rebuilds only affected provider subgraphs; in-flight requests may retain the old secret-bearing client Arc until their old root retires.
Ordinary logs, metrics, traces, audit events, projection/root digests, benchmark artifacts, crash reports, and module inspection contain neither secret values nor credential reference IDs. Repair-only operator diagnostics require explicit authorization and still prefer deployment/error IDs.
Evidence
projection-secret.json exercises success, missing, denied, rotation,
redaction, last-valid-root, and zero request-time lookup assertions.
ADR 0006: MVP Accounting, Circuit, and Replay Defaults
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 before LF-5B
Decision
The initial measurable configuration keys are:
llm-router.accounting.estimatorId,estimatorVersion,safetyMarginBps,maxInputUnits,maxOutputUnits,maxReservedCostMicros, andunknownPricingMode;llm-router.circuit.failureThreshold=5,openCooldownMs=30000, andhalfOpenProbePermits=1;llm-router.retry.maxAttempts=1for the first buffered slice andmaxReplayBytes=1048576.
Reservations are per replica and keyed by host/principal/alias; they are explicitly non-distributed. The conservative local estimator is identified by wire profile and is never reported as provider billing usage. Hard accounting fails closed on unknown pricing; observational profiles preserve unknown and incomplete evidence.
Timeout/cancellation reconciliation remains conservative. Passive circuits
count configured transport, timeout, throttling, and provider 5xx categories.
Retry-After cooldown belongs to provider-account/quota-group and deployment
state. Any future multi-attempt policy whose canonical replay body exceeds
maxReplayBytes is rejected at publication.
Numeric defaults remain manifest inputs and may change only with new benchmark and pricing evidence.
ADR 0007: Audit Durability Profiles and Group Commit
- Status: Accepted
- Date: 2026-07-18
- Gate: LF-2 evidence; implementation before LF-8
Decision
The named profiles are best-effort, bounded-async,
local-durable, and remote-durable. MVP implements bounded-async as the
production default and local-durable under a separate SLO. required is an
admission policy, not a durability level.
Bounded-async reserves the complete metadata envelope and bounded queue/spool
capacity before dispatch but does not wait for disk commit. Local-durable waits
until every attempt-start record reaches the WAL durable watermark. The WAL is
single-writer, length-delimited, checksummed, sequence-numbered, and uses group
fdatasync; recovery stops and reports any corrupt/truncated committed
record. Full/read-only storage fails admission for required profiles rather than
silently degrading.
Remote-durable is reserved for a later authoritative sink transaction. Best-effort is development-only and counts loss.
Evidence
benchmarks/llm-gateway/evidence/wal.json measures grouped synchronization,
durable watermark/recovery, truncated-tail detection, and fail-closed capacity
and read-only behavior. It is feasibility evidence, not the production WAL
implementation.
A2A Gateway
Status: Phases 0 through 7 are implemented for the governed JSON-RPC profiles. The checked-in gates qualify source, schema, SDK, deployment, reload, and operational contracts; live environment evidence is still required before a production release decision.
Related control-plane designs:
- AI Agent Registration In Task Center defines the logical Agent, native runtime link, base Agent publication, and explicit handoff into optional A2A publication.
- Control-Plane Policy Publication Through Config Server
defines the immutable snapshot,
(host, serviceId, envTag)workload identity,/configs, reload, acknowledgement, last-known-good, and rollback contract reused here.
This document defines how light-gateway, native A2A support in light-agent,
and a first-class light-a2a integration service should expose and govern
Agent2Agent (A2A) traffic while Light Portal remains the catalog and policy
authority. light-agent embeds the managed A2A boundary for Portal-native
agents. light-a2a supplies that boundary only for external business agents or
remote A2A federation.
The core principle is:
light-gatewayprotects the public edge, the selected native or integration runtime enforces A2A-specific protocol and policy from shared modules, and the selected agent implementation owns business reasoning and domain effects.
Summary
Introduce shared A2A protocol, policy, card, and task modules; embed them in
light-agent; provide a registered, horizontally scalable light-a2a external
integration service; and add a small a2a-router edge module in
light-gateway. Together they support three onboarding paths:
- Portal-managed generic agent: publish a
light-agentassembled from Portal-managed prompts, models, capabilities, skills, memory, knowledge, tools, and workflows.light-agentterminates A2A and enforces the shared A2A security and policy contract in-process; no sidecar is deployed. - External business agent with managed sidecar: run
light-a2abeside custom agent code. The developer implements a narrow business interface; the sidecar owns A2A, platform security, fine-grained access control, policy, task correlation, limits, audit, and telemetry. - Existing remote A2A agent: use shared-service
light-a2afederation to expose or call an already compliant remote agent through an approved Portal catalog binding.
The same light-a2a binary supports two external-integration deployment
profiles. Shared-service mode is the default for many remote agents and
horizontal scaling. Sidecar mode is reserved for private/local external
business implementations, isolated credentials, or backends that do not
implement A2A themselves. It is never inserted beside light-agent.
The first external-developer profile uses the private
light-a2a-backend/v1 HTTP/JSON contract over a fixed loopback origin, with
SSE when a backend declares streaming. Python, Java, and TypeScript SDKs are
production release requirements; Rust provides the reference implementation
and shared conformance harness. This local backend contract is not the public
A2A HTTP+JSON binding.
The first delivery supports A2A 1.0 JSON-RPC with an explicit A2A 0.3 compatibility profile and no activated protocol extensions. Phase 6 adds an independently enabled external-sidecar A2A 1.0 profile for extended disclosure, optional data-only extensions, and governed push notifications. Public A2A HTTP+JSON, public A2A gRPC, custom bindings, required extensions, and additional transport or SDK profiles remain disabled until independently qualified. The first production milestone includes both governed inbound publication and governed outbound invocation. Implementation remains sequenced so the inbound server, identity, policy, and task foundations land before outbound completion, but inbound-only operation is a development canary rather than the production release boundary.
The design intentionally does not copy the AgentGateway implementation. It
adopts the useful behavioral baseline—traffic classification, Agent Card URL
rewriting, and A2A-aware telemetry—but uses the existing light-gateway
handler chain, configuration publication, security, registry, and reload
models. It adds shared A2A modules because protocol-semantic authorization and
durable task handling do not belong in the gateway process, and adds the
light-a2a application for external-agent adaptation, federation, and
developer-facing runtime isolation.
Background
The AgentGateway repository implements A2A as an empty traffic-policy marker on top of its ordinary HTTP proxy. When enabled, it:
- recognizes the legacy and current well-known Agent Card paths;
- classifies JSON
POSTrequests and records their JSON-RPC method; - rewrites backend Agent Card URLs to the public gateway address;
- recognizes selected A2A 0.3 and 1.0 response fields for telemetry; and
- relies on general gateway policies for authentication, authorization, routing, rate limiting, TLS, and transformations.
That is a useful interoperability baseline, but it does not own agent registration, skill assignment, memory, task persistence, or agent execution.
Light-Fabric already has richer platform foundations:
- Light Portal registers an agent as an API version with API type
agt. agent_definition_tbinds that agent identity to an authorized model alias or model policy and to the remaining Agent profile. Direct provider/model/key fields are legacy compatibility inputs, not the native publication model.skill_t,agent_skill_t, andskill_tool_trepresent governed skills and their assigned tools.genai-query/getEffectiveAgentCatalogcompiles an agent-scoped catalog.- the immutable
light-agent/agentprojection carries definition, model, skills, catalog, memory, knowledge, execution, channel, and session policy. light-agentowns durable sessions, turns, actions, event history, memory banks, recall, and retention.light-gatewayalready provides handler-chain dispatch, JWT validation, access control, delegation, request and response filtering, rate limiting, TLS, service discovery, telemetry, and last-known-good configuration reload.
The implementation now provides an a2a-router handler in light-gateway,
native A2A routes in light-agent, and a registered light-a2a application
for external sidecar and federation profiles. light-agent continues to load
its effective catalog from an immutable Config Server snapshot; a live Portal
query is not runtime authority.
Use Cases
Existing A2A Server Behind Light Gateway
An organization already operates an A2A-compatible agent. It registers the
agent and backend binding in Portal, then exposes it through light-gateway
and shared-service light-a2a. The remote server stays authoritative for its
tasks; light-a2a validates the protocol and applies the approved binding,
delegation, fine-grained policy, limits, and telemetry.
External Business Agent With A Managed Sidecar
An external developer implements a narrow local business interface instead of
implementing A2A and every Light platform concern. A light-a2a sidecar
terminates A2A, loads its Portal-published policy, validates callers and
operations, manages protocol task correlation, and invokes the business
backend over the first-release fixed-loopback HTTP/JSON transport. Unix-domain
sockets and mutually authenticated network transports require separately
versioned and qualified backend profiles.
The sidecar passes a short-lived signed authorized invocation context. It does not forward the caller’s raw bearer token. The business implementation still enforces domain invariants, but it does not recreate platform authentication, tenant isolation, A2A task handling, audit, or observability.
Light Agent Published As A2A
A Portal-defined light-agent is published for external or internal A2A
clients. Portal compiles the public Agent Card. The gateway exposes its public
route and routes directly to the registered light-agent. Shared A2A server
modules inside light-agent serve the card and map A2A contexts and tasks onto
its durable Light sessions and turns without a sidecar or internal network hop.
Light Agent Calling An External Agent
A Light agent selects an external agent from its effective, policy-filtered
catalog. The call is sent through light-gateway using a stable server-owned
agent reference. light-a2a resolves the approved destination, attaches
server-owned credentials or delegation, and enforces outbound data policy
after gateway edge admission.
Public And Extended Discovery
An anonymous caller can receive a deliberately limited public Agent Card in the first production profile. The independently authorized Phase 6 external-sidecar profile lets an authenticated caller request an extended Agent Card containing additional policy-approved skills or interfaces. Neither form exposes internal skill instructions, tool bindings, credentials, memory, or topology.
Goals
- Support interoperable A2A Agent Card discovery and message/task operations.
- Deliver both governed inbound A2A exposure and governed outbound A2A calls in the first production milestone.
- Make the selected native
light-agentor external-integrationlight-a2aruntime the A2A protocol-semantic and fine-grained policy enforcement point while reusing the same shared Light runtime, A2A, and security crates. - Reuse
light-gatewaypublic-edge authentication, routing, filtering, rate limiting, TLS, service discovery, and telemetry without duplicating policy authority. - Let external developers implement business logic behind a narrow trusted backend contract without implementing A2A or Light platform plumbing.
- Use stable Portal agent identity instead of client-selected target URLs.
- Publish Agent Cards from immutable, versioned Portal projections.
- Preserve the distinction between public discovery metadata and the richer internal effective agent catalog.
- Map Portal-assigned skills to a safe A2A
AgentSkillprojection. - Keep durable session, task, model-loop, skill execution, knowledge, and memory responsibilities in the agent runtime.
- Support transparent external A2A backends without requiring them to adopt Light Portal’s internal runtime model.
- Offer one
light-a2abinary in shared-service and sidecar deployment modes for external integrations, never as a required companion tolight-agent. - Make version support, body limits, timeouts, streaming, and failure mapping explicit and testable.
- Support safe horizontal scaling and last-known-good configuration reload.
Non-Goals
- Do not embed an agent runtime or model loop in
light-gateway. - Do not make the external business backend parse public A2A requests, validate raw platform tokens, or query Portal policy.
- Do not make
light-a2aown prompts, model selection, memory recall, knowledge retrieval, tool selection, or business-domain decisions. - Do not store A2A tasks or conversation state only in gateway memory.
- Do not let an Agent Card, request body, model response, or memory grant tools, credentials, network destinations, or authorization.
- Do not publish raw
contentMarkdown, tool schemas, workflow configuration, memory content, API keys, or internal service locations in public cards. - Do not query Portal authoring/projection tables on the request path. A runtime may read its own operational task/correlation store when required by an authorized A2A operation.
- Do not accept arbitrary caller-provided upstream URLs, credential references, service IDs, or agent definition IDs.
- Do not claim public A2A HTTP+JSON, public A2A gRPC, push notification, or custom-binding support until each binding passes its own conformance and operational gates.
- Do not translate the existing
/chatWebSocket protocol inside the gateway into a second, gateway-owned task engine. - Do not deploy one sidecar per remote SaaS agent when a shared federation service provides the required network and credential boundary.
- Do not expose every registered Agent through A2A automatically. Registration and native runtime linking are prerequisites; A2A exposure is an explicit, independently authorized publication decision.
Decisions
Portal Owns Authoring And Publication
Portal owns agent identity, descriptive metadata, assigned skills, visibility, security declarations, supported interfaces, backend bindings, and publication lifecycle. It compiles those records into an immutable A2A runtime projection.
The gateway consumes only the approved projection. It does not reconstruct an Agent Card by joining Portal tables and does not decide which skills should be public.
Public Skill IDs Are Stable Publication Aliases
An A2A AgentSkill.id is a public protocol identifier. The
A2A contract
requires a unique string; it does not require or benefit from exposing an
implementation database key. Portal therefore publishes a stable,
tenant-scoped publicationAlias, such as billing.refund-review, and never
publishes the skill_t.skill_id UUID as the A2A skill ID.
Portal stores the UUID and alias as separate structured identities:
skillId internal Portal identity, joins, policy and audit
publicationAlias stable public A2A AgentSkill.id
Portal View suggests an alias from the skill name, lets an authorized owner confirm or change it before first publication, validates normalized uniqueness within the host, and displays it on the Skill form and effective-card preview. After the first successful publication, compatible skill revisions retain the alias and advance their version and digest. An incompatible semantic replacement requires a new alias and normally a new skill identity; an ordinary update must not silently transfer a public ID to different behavior.
Each immutable agent publication records the exact
publicationAlias -> skillId + skillVersion + skillDigest mapping used to
compile its card and runtime projection. That mapping supports deterministic
dispatch, rollback, audit, and historical task interpretation without a
request-path Portal lookup. The alias is discovery and correlation metadata,
not authority: a caller that knows it gains no skill, tool, workflow, or agent
permission.
Gateway Owns The Public Edge
light-gateway owns:
- public listener, host, path, and TLS policy;
- coarse caller authentication and endpoint admission;
- bounded edge request admission;
- routing only to the registered
light-agentorlight-a2aservice selected by the published implementation kind; - optional generic request/response filtering;
- edge rate limiting and denial-of-service controls; and
- public-edge audit, metrics, and trace correlation.
The edge may classify A2A traffic and extract bounded metadata for routing and telemetry. It is not the authoritative parser or fine-grained A2A policy decision point.
Agent APIs Use Deployment-Scoped Gateway Policy Identities
Each logical agent is modeled as an API and API version in Portal. The existing
identity rule remains agentDefId == apiVersionId. A deployable agt product
and product version describe the compatible Light agent runtime and its
configuration contract; they do not replace the API-version identity of an
individual agent.
Publishing an agent through a Gateway requires an active instance_api_t
association between that agent API version and the target light-gateway
instance. Its instanceApiId is the deployment-scoped binding identity used to
compile routing and coarse edge authorization. This relationship is distinct
from the binding that selects the native light-agent, external sidecar, or
remote A2A implementation.
For a native Agent, these are two distinct deployment relationships:
Agent API version -> native light-agent runtime
Agent API version -> public light-gateway instance
The Agent registration flow owns or verifies the first relationship. The A2A
publication flow owns or verifies the second. Their instanceApiId values are
not interchangeable. An internal instanceId or instanceApiId may appear in
Portal commands, manifests, associations, and audit evidence, but neither is a
Config Server workload identity or query parameter.
Multiple agent APIs intentionally expose the same A2A protocol endpoints. Raw
keys such as /@post or /message:send@post therefore cannot be used directly
as keys in the Gateway’s combined rule.endpointRules map. Portal compiles
opaque, exact-match policy endpoint keys in this namespace:
a2a:instance-api:<instanceApiId>:card
a2a:instance-api:<instanceApiId>:invoke
a2a:instance-api:<instanceApiId>:endpoint:<endpointId>
The initial profile requires card and invoke. The endpoint-specific form is
available when a later binding exposes independently governed HTTP surfaces.
These are authorization resource identities, not public URLs and not A2A
operation names. Portal View never asks an administrator to enter them.
Each agent also receives a unique, human-readable public path prefix on a
Gateway, for example /a2a/order-agent. Portal stores it with the Instance API
path-prefix association and rejects normalized public host-and-path collisions.
The route projection maps that public identity to instanceApiId,
apiVersionId/agentDefId, implementation kind, registered target service,
and the generated policy endpoint keys. Changing a public alias does not
silently transfer authority because authorization remains bound to the
Instance API association.
Gateway authorization is deliberately two-level. light-gateway uses the
generated card or invoke key for coarse access to one published agent. The
selected light-agent or light-a2a runtime then authoritatively parses A2A
and authorizes the specific abstract operation, skill, task/context ownership,
delegation, and data boundary. The edge-level invoke class must never be
treated as permission for every A2A operation inside the selected runtime.
The Selected Runtime Owns The Managed A2A Boundary
Shared A2A modules embedded in the selected runtime own:
- well-known Agent Card rendering and disclosure-class selection;
- A2A version, binding, extension, and operation negotiation;
- bounded protocol parsing and schema validation;
- caller, target-agent, skill, operation, tenant, data-boundary, task-ownership, delegation-depth, budget, and limit authorization;
- deterministic backend resolution;
- server-owned credential and trust-policy resolution;
- task/context/idempotency correlation required by adaptation;
- safe public interface construction;
- A2A-aware request/output validation and redaction;
- streaming and callback transport enforcement;
- response and error normalization at the protocol boundary; and
- protocol-level audit, metrics, traces, and policy-decision evidence.
For a LIGHT_AGENT binding, light-agent embeds these modules and loads the A2A
server policy in its immutable Agent audience projection. For
EXTERNAL_SIDECAR and REMOTE_A2A, light-a2a embeds the same modules and
loads an immutable light-a2a audience projection. Each real runtime registers
with the Controller. Neither runtime reconstructs live authority by joining
Portal authoring tables on the request path.
Agent Runtime Owns Work
The selected runtime owns:
- message interpretation and model execution;
- durable contexts, tasks, turns, actions, and cancellation;
- tool selection and execution through the approved placement;
- knowledge retrieval and evidence;
- memory-bank selection, recall, retention, and reflection; and
- final task results and artifacts.
For an external A2A backend, that backend is the task authority. For a
Portal-native agent, light-agent and its durable database model are the task
authority. For a non-A2A external business backend, light-a2a owns the A2A
task facade and correlation record, while the backend owns business execution
and effects.
A2A Task Artifacts Have Independent Retention And Visibility
An A2A artifact is the concrete output of a task, such as a document, image,
structured result, report, or generated file. It is not chat history and it is
not Hindsight memory. The
A2A specification likewise
separates task history messages from task output artifacts and permits an
expired or purged task to produce TaskNotFoundError. The selected runtime owns
the artifact lifecycle together with the durable task; light-gateway never
becomes an artifact repository.
The three data classes remain independently governed:
| Data class | Purpose | Runtime authority |
|---|---|---|
| A2A task artifact | Exact task deliverable, integrity evidence, and retrievable output. | A2A artifact policy and operational artifact store. |
| Chat/session history | Conversation reconstruction and continuation. | Durable session events and the session-history policy. |
| Hindsight memory | Derived facts, experiences, and mental models selected for later recall. | Memory-bank policy and bank-scoped authorization. |
TASK_OWNER is the default artifact visibility, not a separate or exclusive
ACL system. Artifact access is default-deny and uses the platform’s existing
fine-grained access-control decision with authenticated host, principal,
calling client or agent, delegated user when present, target publication,
skill, task/context ownership, requested artifact operation, data
classification, and applicable obligations. Policy may explicitly grant
additional principals or workloads access. Operators and Portal administrators
receive no implicit artifact-content authority. There is no special break-glass
authorization path: successful and denied decisions use the normal audit
pipeline, and audit recording never grants access.
At minimum, authorization distinguishes artifact metadata read, content read or
download, export, deletion, and promotion to memory. The same ownership and
policy checks apply when an artifact is reached through GetTask, ListTasks,
subscription, a task response, or a download URL. A task ID, context ID,
artifact ID, object reference, or URL is an identifier rather than a
credential. A resource that is absent, expired, or inaccessible produces the
binding-correct not-found result without disclosing which condition applied.
The initial production profile creates no public, tenant-wide, or anonymous
link visibility by default; any later sharing is an ordinary explicit
fine-grained policy grant, not a new artifact-specific security mechanism.
Artifact bytes live in tenant-scoped, content-addressed managed object storage. The operational database stores bounded metadata such as task and agent ownership, logical name, media type, size, content digest, storage reference, classification, policy and publication digests, provenance, verification state, retention deadline, legal hold, and deletion evidence. Object-store references and credentials are never exposed as public artifact identities. Small inline A2A parts remain subject to the same lifecycle and limits; an implementation must not escape retention by copying them into an ungoverned task JSON column.
The platform ships these conservative defaults, which Portal may replace with an approved host or agent artifact-retention profile:
| Record | Initial default |
|---|---|
| Incomplete streaming chunks | Compact into the final artifact and remove residual chunks within 24 hours. |
| Final artifact content | Retain for 30 days after the task becomes terminal. |
| External A2A task and artifact visibility | Retain for the same 30-day retrieval window. |
| Metadata, digest, provenance, and deletion tombstone without content | Retain for 365 days or the approved compliance period. |
| Legal hold | Suspend ordinary content deletion until the hold is released by authorized policy. |
The effective profile and deadlines are frozen when the task is admitted, so a
later configuration change cannot silently extend access to an existing
artifact. When the external retrieval window ends, GetTask treats the task as
expired even if non-content audit metadata remains internally. Expiration or
deletion removes managed bytes, previews, temporary objects, and caches,
verifies absence, and retains bounded deletion evidence. A privacy-erasure
workflow follows recorded lineage into authorized derived copies; an ordinary
artifact TTL does not implicitly delete separately governed chat or memory.
Chat history may store a bounded artifact ID, digest, and display reference but not duplicate the artifact bytes. Hindsight does not automatically retain an artifact. Promotion to memory is a separate fine-grained operation that extracts or summarizes bounded content, applies memory-bank visibility and redaction, records the source artifact ID and digest as provenance, and creates a new memory record with its own retention. The derived memory may outlive ordinary artifact expiration, while privacy erasure can still find it through the provenance lineage.
For an external result, a remote URI is not assumed durable. If Light-Fabric
promises local retrieval, light-agent or light-a2a validates, scans, digests,
and imports the content into the tenant-owned store. Otherwise it marks the
reference ephemeral and makes no availability promise beyond the upstream
server’s accepted lifetime. Raw remote or object-store URLs are never treated
as permanent platform artifact references.
Reuse the existing light-workflow artifact mechanics: tenant-scoped staging
and content-addressed promotion, digest verification, quarantine, legal holds,
retryable deletion, verified absence, and durable tombstones. Extract shared
artifact policy, validation, storage, and retention contracts rather than
inserting every A2A result directly into workflow_artifact_t, whose required
workflow execution identity does not fit native model-only or external A2A
tasks. The A2A operational schema may use a dedicated task-artifact ownership
table or a future generalized platform artifact table, but it must use the
same lifecycle state machine.
A2A Discovery Is Not Runtime Authority
An Agent Card is descriptive input for discovery. Its skill list, security schemes, interfaces, and capabilities do not grant access. Effective authority is the intersection of:
authenticated caller authority
intersect published agent visibility
intersect agent-definition policy
intersect gateway edge admission
intersect selected-runtime A2A operation and binding policy
intersect runtime policy snapshot
intersect live backend capability
Extensions Are Registered And Disabled By Default
An A2A extension can add structured metadata, constrain the core message
profile, introduce methods, or refine task-state behavior. It is therefore
executable protocol policy rather than an arbitrary metadata field. The first
production profiles advertise and activate no extensions: their Agent Cards
omit capabilities.extensions or publish an empty list, and their compiled
inbound and outbound allowlists are empty.
Every future extension that Light-Fabric advertises, activates, interprets, or
forwards must use an exact, versioned URI from a Portal-managed registry. All
activated extensions are allowlisted; required: true has a stricter gate and
is initially prohibited. A required declaration is a compatibility constraint
from an agent, not proof that the extension is safe. It becomes eligible only
after its implementation, dependencies, parameter schemas, directions,
operations, security review, and conformance evidence are approved for the
selected runtime profile.
For A2A 1.0 over HTTP, a client requests activation with the A2A-Extensions
service parameter. The selected light-agent or light-a2a runtime performs
authoritative negotiation:
- an allowlisted requested extension is activated only after its parameters and dependencies validate, and the response lists the extensions actually activated;
- an unknown optional extension is not activated or echoed, and its namespaced metadata is excluded from policy, model, backend, and artifact inputs;
- invalid content for a recognized and requested extension is rejected with the binding-correct validation or extension error;
- omission of an extension that the published card marks required returns
ExtensionSupportRequiredError; and - an unapproved required extension in a remote Agent Card prevents onboarding or publication rather than failing for the first time on a production call.
A2A-Extensions negotiation and ExtensionSupportRequiredError are 1.0-only
in this design. The 0.3 compatibility profile cannot advertise, require,
accept, or activate extensions. Portal publication and runtime projection
compilation reject any attempt to configure them for 0.3 rather than inventing
a non-standard wire parameter or error mapping.
See the A2A extension negotiation and required-extension rules. Light-Fabric authentication, delegation, tenant binding, data-boundary policy, and task ownership remain trusted platform context; the first release does not re-express them as a proprietary required A2A extension.
Memory Never Lives In The Gateway
The gateway may propagate authenticated subject, agent, context, task, and correlation identifiers. It must not recall memory, inject remembered text, choose a memory bank, or retain model output.
An A2A contextId is an external correlation identifier, not proof of session
ownership. A runtime must bind it to the authenticated host, principal, agent
definition, and policy snapshot before resuming a session or accessing memory.
Architecture
CONTROL PLANE
Portal API/Agent Catalog
|-- API version type agt
|-- agent definition and implementation kind
|-- assigned skills and tools
|-- managed extension registry and profiles
|-- public/extended disclosure policy
|-- environment-specific agent bindings
|-- backend, trust, and credential references
`-- publication/signing policy
| |
| v
| light-oauth signing authority
| |-- purpose-bound A2A key profiles
| |-- Agent Card JWS signing
| `-- public JWKS and rotation
| |
`----------- signed publication
|
v
Config Server Controller
immutable audience projections live agent/A2A instances
| ^
v |
DATA PLANE
A2A client --> light-gateway --> light-agent native A2A endpoint
public edge |-- shared A2A server modules and policy
`-- native durable sessions and turns
A2A client --> light-gateway --> light-a2a shared service
|-- A2A semantics and federation policy
`--> existing remote A2A server
A2A client --> light-gateway --> light-a2a sidecar --> external business code
Light agent --> light-gateway --> light-a2a --> approved external A2A server
Component Responsibilities
| Component | Responsibilities | Explicitly does not own |
|---|---|---|
| Portal | Agent authoring, implementation kind, catalog identity, skill assignment, extension registry and profiles, disclosure and access policy, backend binding, signing policy, publication lifecycle. | Request-path routing, extension negotiation, or task execution. |
light-oauth | Resolve host, environment, and purpose-bound signing profiles; sign final canonical Agent Cards; publish profile JWKS; execute key rotation and revocation; and record signing audit. | Agent metadata authoring, card rendering, A2A authorization, or reuse of OAuth token keys as Agent Card keys. |
| Config Server | Publish immutable, audience-specific Gateway, A2A, and Agent projections. | Live policy evaluation, agent reasoning, or per-request database joins. |
| Controller | Discover, register, health-check, and route to live light-agent and light-a2a instances. | External-agent catalog identity or complete agent policy. |
a2a-router | Resolve a published public route to its Instance API binding and generated policy endpoint, perform coarse admission, route by implementation kind to the registered native or integration runtime, apply generic filtering, and emit edge telemetry. | Authoritative A2A parsing, fine-grained A2A policy, durable tasks, skills, memory, or model calls. |
light-a2a | External-agent A2A server/client bindings, cards, semantic validation, fine-grained policy, secure backend resolution, adapter correlation, delegation, limits, redaction, audit, and telemetry. | Portal-native agent execution, model loops, memory recall, tool selection, or business-domain effects. |
light-agent | Native A2A server boundary, Agent Card serving, fine-grained A2A policy, model loop, durable sessions/turns/actions, effective tools, knowledge, memory, results, audit, and telemetry. | Gateway edge routing or external-agent adaptation/federation. |
| External business backend | Custom reasoning, domain validation, and domain effects through a narrow trusted interface. | Public A2A, raw token validation, Portal policy lookup, Controller integration, or platform telemetry. |
| External A2A server | Its own task semantics and results behind an approved federation binding. | Light Portal or Light policy authority. |
Agent Implementation And Binding Model
Portal should distinguish the stable agent definition from its deployable binding. The definition declares what the agent is; a binding declares how one environment reaches and governs an implementation.
Recommended implementation kinds are:
| Kind | Meaning |
|---|---|
LIGHT_AGENT | Portal-managed generic light-agent; prompt, model, skills, memory, knowledge, tools, execution policy, and native A2A server policy are projected to that runtime. No sidecar is used. |
EXTERNAL_SIDECAR | Custom business implementation reached through a local or private light-a2a sidecar backend contract. |
REMOTE_A2A | Existing A2A-compliant server reached through shared-service light-a2a federation. |
An agent binding contains environment, network zone, selected interfaces, backend location, trust and credential references, active publication, and operational limits. Multiple bindings may implement the same definition in different environments without creating a second agent identity.
The Controller registers actual Light runtime instances. It registers each
native light-agent, light-a2a shared-service replica, or light-a2a sidecar
process, not every remote external agent as a virtual service instance. A
Light-managed external implementation may register independently when it is
itself a real runtime.
Protocol Scope
Initial Profile
| Capability | Initial decision |
|---|---|
| A2A version | 1.0 primary, explicit 0.3 compatibility profile. |
| Binding | JSON-RPC over HTTP. |
| Agent Card | /.well-known/agent-card.json; optional legacy /.well-known/agent.json. |
| Message | Non-streaming and SSE streaming when the selected runtime or backend declares support. |
| Task lookup | Supported only when the native runtime or external integration provides durable task lookup. |
| Cancellation | Supported only through a durable backend operation; never gateway-local cancellation state. |
| Extended card | Deferred to Phase 6 as an independently authorized disclosure profile. |
| Push notifications | Deferred. |
| Public A2A HTTP+JSON and gRPC | Deferred binding profiles; this does not prohibit the private sidecar backend HTTP/JSON contract. |
| Extensions | 1.0 advertises and activates none initially; 0.3 extension configuration is rejected before publication. |
Every request must resolve one configured profile. After applying the version-specific missing-value rule below, an unsupported version, binding, operation, content type, or capability must produce the corresponding A2A-compatible error rather than silently falling through as an ordinary proxy request. For 1.0, extension handling follows the separate optional, required, and malformed rules above: an unsupported optional URI is not activated, while a missing required extension or invalid activated extension is an error. The 0.3 compatibility profile rejects extension configuration before publication or runtime activation.
For HTTP bindings, the selected native or integration runtime accepts the A2A
version from the normative A2A-Version header or request parameter. A 1.0
client must provide it; an absent value is interpreted as 0.3 and then checked
against the selected interface.
A 1.0-only interface therefore returns VersionNotSupportedError for an
absent value instead of inferring 1.0 from payload shape. The gateway edge does
not interpret the version. See the
A2A 1.0 versioning requirements.
Abstract Operation Layer
Define one internal operation model independent of transport:
enum A2aOperation {
GetAgentCard,
GetExtendedAgentCard,
SendMessage,
SendStreamingMessage,
GetTask,
ListTasks,
CancelTask,
SubscribeToTask,
SetPushNotificationConfig,
GetPushNotificationConfig,
ListPushNotificationConfigs,
DeletePushNotificationConfig,
}
Binding adapters parse the transport into this model. Authorization, target selection, task ownership, limits, and telemetry use the abstract operation instead of transport-specific method strings.
Managed A2A Runtimes
Embed the shared A2A server modules in apps/light-agent for LIGHT_AGENT
bindings. Create apps/light-a2a on light-axum and the existing
light-runtime lifecycle for EXTERNAL_SIDECAR and REMOTE_A2A bindings. Each
application uses startup.yml and portal-registry.yml, loads its own
audience-specific configuration through Config Server, registers its real
service instance with the Controller, and participates in managed startup,
reload, quiesce, and shutdown.
The gateway retains a small a2a-router edge handler whose configuration maps
approved public routes to their Instance API binding, generated coarse policy
endpoints, implementation kind, and registered light-agent or light-a2a
service identity. Protocol models, Agent Card rendering, fine-grained
decisions, backend bindings, and task correlation live outside the gateway.
Shared Crate Boundaries
light-agent and light-a2a depend on shared crates; neither application
depends on the other application’s internal modules.
crates/a2a-protocol/ versioned wire types, parsing and conformance
crates/a2a-client/ outbound bindings and Agent Card client
crates/a2a-core/ operations, errors, authorization envelope, policy,
task mapping, artifact lifecycle and retention types
crates/a2a-store/ durable integration-service task correlation
crates/agent-store/ durable native-agent A2A task correlation
contracts/a2a-backend/v1/ private backend protocol and SDK conformance
apps/light-a2a/ external integration server/client and adapters
apps/light-agent/ generic agent runtime with native A2A server
Phase 1 intentionally keeps the small mutually dependent runtime, policy,
envelope, and artifact value contracts in a2a-core. Split them into narrower
crates only when an independent consumer or release cadence requires it; do not
create pass-through crates that merely re-export the same types. The private
backend contract and reference adapter remain Phase 3 deliverables.
Both applications also reuse existing light-runtime, light-security,
light-client, agent-core, and agent-delegation. The common policy crate
contains only immutable cross-runtime authority. Prompt, model, memory,
knowledge, tool-selection, approval, and model-loop configuration remain in a
light-agent-specific policy layer.
Do not move light-agent SQL repositories, session/turn orchestration, memory
logic, or model execution into shared A2A crates merely because A2A exposes
tasks. Share stable contracts and validation, not application ownership.
Deployment Profiles
| Profile | Placement and scope |
|---|---|
native | Shared A2A modules embedded in each light-agent; handles only that runtime’s LIGHT_AGENT binding and durable agent state. There is no sidecar or adapter network hop. |
shared | Registered, horizontally scalable light-a2a service for remote external agents within one host/tenant policy boundary. Centralizes federation connections, cards, trust, policy and telemetry. |
sidecar | One light-a2a process beside a private/custom external backend that is not part of the Light-Fabric agent runtime. Its projection pins one or a small bounded set of agent bindings and denies arbitrary destinations. |
A native light-agent registers its A2A capability, supported versions, and
active configuration generation with its existing service identity. A sidecar
registers with tags such as mode=sidecar, network zone, supported A2A
versions, active configuration generation, and agent binding ID. A shared
replica registers mode=shared and its supported binding capabilities. Full
agent policy remains in Config Server, not Controller metadata. A
LIGHT_AGENT publication must never select mode=sidecar.
Each accepted runtime projection is single-host. All agent bindings in that
projection inherit and must match runtimePolicy.host. Shared mode means
many agents within that boundary, not one mixed-host policy document. A future
multi-host fleet must load separately signed and isolated host partitions; it
must not weaken the host check or combine entries under one envelope.
Proposed light-a2a Configuration Shape
runtimePolicy:
publicationId: ${runtimePolicy.publicationId:}
releaseVersion: ${runtimePolicy.releaseVersion:0}
policySnapshotId: ${runtimePolicy.policySnapshotId:}
policyVersion: ${runtimePolicy.policyVersion:0}
policyDigest: ${runtimePolicy.policyDigest:}
contentDigest: ${runtimePolicy.contentDigest:}
audience: ${runtimePolicy.audience:light-a2a}
host: ${runtimePolicy.host:}
serviceId: ${runtimePolicy.serviceId:}
envTag: ${runtimePolicy.envTag:}
sourceEventSequence: ${runtimePolicy.sourceEventSequence:0}
schemaVersion: ${runtimePolicy.schemaVersion:1}
createdAt: ${runtimePolicy.createdAt:}
validFrom: ${runtimePolicy.validFrom:}
revocationEpoch: ${runtimePolicy.revocationEpoch:0}
compatibilityGeneration: ${runtimePolicy.compatibilityGeneration:1}
a2aPolicy:
enabled: ${a2aPolicy.enabled:true}
mode: ${a2aPolicy.mode:shared}
maxRequestBodyBytes: ${a2aPolicy.maxRequestBodyBytes:1048576}
maxResponseInspectionBytes: ${a2aPolicy.maxResponseInspectionBytes:4194304}
maxAgentCardBytes: ${a2aPolicy.maxAgentCardBytes:262144}
maxJsonDepth: ${a2aPolicy.maxJsonDepth:128}
maxConcurrentRequests: ${a2aPolicy.maxConcurrentRequests:1024}
maxConcurrentRequestsPerPrincipal: ${a2aPolicy.maxConcurrentRequestsPerPrincipal:32}
requestTimeoutMs: ${a2aPolicy.requestTimeoutMs:120000}
streamIdleTimeoutMs: ${a2aPolicy.streamIdleTimeoutMs:30000}
cardCacheTtlSeconds: ${a2aPolicy.cardCacheTtlSeconds:60}
artifacts:
defaultVisibility: ${a2aPolicy.artifacts.defaultVisibility:TASK_OWNER}
transientRetentionHours: ${a2aPolicy.artifacts.transientRetentionHours:24}
contentRetentionDays: ${a2aPolicy.artifacts.contentRetentionDays:30}
taskVisibilityDays: ${a2aPolicy.artifacts.taskVisibilityDays:30}
metadataRetentionDays: ${a2aPolicy.artifacts.metadataRetentionDays:365}
memoryPromotion: ${a2aPolicy.artifacts.memoryPromotion:EXPLICIT_AUTHORIZATION}
externalReferenceMode: ${a2aPolicy.artifacts.externalReferenceMode:IMPORT_OR_EPHEMERAL}
requireMalwareScan: ${a2aPolicy.artifacts.requireMalwareScan:true}
cardSigning:
delivery: ${a2aPolicy.cardSigning.delivery:PRE_SIGNED}
profileId: ${a2aPolicy.cardSigning.profileId:}
jwksUrl: ${a2aPolicy.cardSigning.jwksUrl:}
signingServiceUrl: ${a2aPolicy.cardSigning.signingServiceUrl:}
profiles:
a2a-v1-jsonrpc:
binding: ${a2aPolicy.profiles.v1.binding:JSONRPC}
versions: ${a2aPolicy.profiles.v1.versions:["1.0"]}
acceptVersionRequestParameter: ${a2aPolicy.profiles.v1.acceptVersionRequestParameter:true}
missingVersion: ${a2aPolicy.profiles.v1.missingVersion:assume-0.3}
extensions:
unknownOptionalAction: ${a2aPolicy.profiles.v1.extensions.unknownOptionalAction:IGNORE}
advertised: ${a2aPolicy.profiles.v1.extensions.advertised:[]}
allowedInbound: ${a2aPolicy.profiles.v1.extensions.allowedInbound:[]}
allowedOutbound: ${a2aPolicy.profiles.v1.extensions.allowedOutbound:[]}
required: ${a2aPolicy.profiles.v1.extensions.required:[]}
maxCount: ${a2aPolicy.profiles.v1.extensions.maxCount:8}
maxHeaderBytes: ${a2aPolicy.profiles.v1.extensions.maxHeaderBytes:2048}
a2a-v03-jsonrpc:
binding: ${a2aPolicy.profiles.v03.binding:JSONRPC}
versions: ${a2aPolicy.profiles.v03.versions:["0.3"]}
acceptVersionRequestParameter: ${a2aPolicy.profiles.v03.acceptVersionRequestParameter:true}
missingVersion: ${a2aPolicy.profiles.v03.missingVersion:assume-0.3}
extensions:
unknownOptionalAction: ${a2aPolicy.profiles.v03.extensions.unknownOptionalAction:IGNORE}
advertised: ${a2aPolicy.profiles.v03.extensions.advertised:[]}
allowedInbound: ${a2aPolicy.profiles.v03.extensions.allowedInbound:[]}
allowedOutbound: ${a2aPolicy.profiles.v03.extensions.allowedOutbound:[]}
required: ${a2aPolicy.profiles.v03.extensions.required:[]}
maxCount: ${a2aPolicy.profiles.v03.extensions.maxCount:0}
maxHeaderBytes: ${a2aPolicy.profiles.v03.extensions.maxHeaderBytes:0}
agents: ${a2aPolicy.agents:[]}
The checked-in template exposes placeholders and conservative defaults. The
runtime agents list and each binding’s effective artifact access-control and
retention profile are compiled by the control plane rather than manually
maintained in every A2A deployment. The artifact fields above contain immutable
rules, never artifact rows, bytes, remote URLs, object-store credentials, or
legal-hold commands. Fine-grained grants remain in the normal access-control
projection. Activation rejects a profile whose task-visibility period exceeds
its managed-content period or whose metadata period cannot cover managed
content and deletion evidence. Retention values are bounded by deployment and
compliance limits rather than accepted as arbitrary integers. For external
bindings, Gateway configuration contains only the public route, catalog and
Instance API binding identities, generated coarse policy endpoints,
implementation kind, and registered light-a2a service destination.
Extension policy is profile-scoped, not instance-global, because each agent
binding selects exactly one profile and one runtime projection may serve both
generations. Every profile is single-generation: its versions list resolves to
one A2A generation, and a profile mixing 1.0 and 0.3 is rejected. A 0.3 profile
rejects every non-empty extension collection during publication and projection
compilation. A 1.0 profile’s extension sets apply only to agents that select
that profile and can never activate, advertise, or relax negotiation for an
agent bound to another profile.
The first-production compiler requires the four extension collections to be
empty in every profile. unknownOptionalAction: IGNORE means do not activate or
echo an unsupported optional URI; it does not permit its metadata to reach the
agent, model, backend, policy engine, or output. A future non-empty entry is a
compiled registry record containing the exact URI, direction, allowed
operations, parameter-schema digest, handler identity, dependency set, metadata
limits, and required-eligibility decision. Within its own profile, required
must be a subset of advertised and the applicable inbound or outbound
allowlist. The runtime rejects a projection whose extension handler or schema
digest is unavailable, whose profile is not single-generation, or whose 0.3
profile carries any extension configuration.
For a native LIGHT_AGENT binding, Portal compiles an optional a2aPolicy
section into the existing agent.yml Agent audience projection. The base Agent
policy and A2A overlay are compiled, validated, snapshotted, activated, and
acknowledged as one immutable generation for the target
(host, serviceId, envTag). The overlay is not a separately activated native
Agent configuration. It contains the
inbound binding, server profiles, disclosure policy, authorization policy, task
and artifact policy, limits, accepted signed card, and logical signing-profile
metadata. It
never contains private-key material or a KMS/HSM key reference. It does not
create a separate light-a2a projection or deploy a sidecar. The Gateway route targets the actual
registered service ID, currently com.networknt.agent.account-1.0.0; publication
must not invent a generic com.networknt.light-agent-1.0.0 identity. The same
Rust A2aPolicy type is embedded by light-agent and by the light-a2a
configuration model even though their enclosing audience projections differ.
PRE_SIGNED is the production default: Config Server delivers the immutable
signed card. signingServiceUrl is populated only for an explicitly approved
activation-time signing fallback; even then, the runtime sends a logical
profileId to light-oauth and never receives the selected key reference.
The compiler derives profileId, jwksUrl, and signingServiceUrl from the
approved signing profile and registered platform service; an agent binding or
runtime cannot supply an arbitrary signer or JWKS origin.
Proposed Gateway Route And Access Projection
The checked-in Gateway templates remain placeholder based:
# a2a-router.yml
routes: ${a2a-router.routes:[]}
# rule.yml
ruleBodies: ${rule.ruleBodies:{}}
endpointRules: ${rule.endpointRules:{}}
Portal compiles the effective Config Server values. The following is an illustrative expanded projection for one agent; the UUIDs and rules are generated or selected authoring references, not manually maintained YAML:
a2a-router.routes:
- publicPathPrefix: /a2a/account-agent
allowedHosts: [agents.example.com]
instanceApiId: 018f0000-0000-7000-8000-000000000001
apiVersionId: 01900000-0000-7000-8000-000000000001
agentDefId: 01900000-0000-7000-8000-000000000001
implementationKind: LIGHT_AGENT
targetServiceId: com.networknt.agent.account-1.0.0
targetEnvTag: prod
policyEndpoints:
card: a2a:instance-api:018f0000-0000-7000-8000-000000000001:card
invoke: a2a:instance-api:018f0000-0000-7000-8000-000000000001:invoke
rule.endpointRules:
a2a:instance-api:018f0000-0000-7000-8000-000000000001:card:
req-acc: [account-agent-card-access]
a2a:instance-api:018f0000-0000-7000-8000-000000000001:invoke:
req-acc: [account-agent-invoke-access]
permission:
role: account-agent-user
The route resolver selects the route and policy endpoint from trusted
configuration before access control. It passes the generated policy endpoint
as the exact authorization resource and separately records the actual request
host, path, instanceApiId, agentDefId, and public route in the rule context
and audit event. Rules can therefore use deployment or request attributes
without accepting identity fields supplied by the caller.
Publishing a second agent with the same A2A API endpoints produces a different
instanceApiId namespace and cannot overwrite the first agent’s entries. A
publication or reload fails closed if route and rule projections disagree,
refer to an inactive binding, contain a duplicate normalized public route, or
reuse a generated policy endpoint across Instance API owners.
External Integration Backend Kinds
light-agent is a direct A2A runtime target, not a backend kind behind
light-a2a. The integration service supports these explicit backend kinds:
| Kind | Purpose |
|---|---|
external-backend | Invoke the narrow trusted backend contract used by a sidecar deployment. |
remote-a2a | Call an approved remote A2A server through a pinned interface and trust policy. |
All destinations are validated configuration, not request data. Hostname,
scheme, port, DNS/IP ranges, TLS policy, redirect behavior, and network zone
must be checked to prevent SSRF and destination substitution. The default
public-destination profile rejects private, loopback, link-local, multicast,
unspecified, documentation, CGNAT (100.64.0.0/10), protocol-assignment
(192.0.0.0/24), and benchmarking (198.18.0.0/15) address space.
External Business Backend Contract
For EXTERNAL_SIDECAR, expose a small SDK contract such as:
#[async_trait]
trait AgentBackend {
fn capabilities(&self) -> BackendCapabilities;
async fn invoke(
&self,
context: AuthorizedInvocation,
request: BusinessRequest,
) -> Result<BusinessResponse, BusinessError>;
async fn invoke_stream(
&self,
context: AuthorizedInvocation,
request: BusinessRequest,
) -> Result<BusinessEventStream, BusinessError>;
async fn status(
&self,
context: AuthorizedInvocation,
) -> Result<BusinessOperationStatus, BusinessError>;
async fn cancel(
&self,
context: AuthorizedInvocation,
) -> Result<(), BusinessError>;
}
AuthorizedInvocation contains the approved principal and agent actor, host,
tenant, environment, selected agent and skill, allowed operation, policy and
data-boundary digests, task/context/idempotency IDs, deadline, budget, and trace
context. It is short-lived and signed. The sidecar never forwards the caller’s
raw bearer token to business code.
For cancellation, the signed context contains exactly one required task ID;
there is no separate unsigned task parameter. Any task or context ID repeated
inside BusinessRequest must equal the signed value. A response that starts
detached or long-running work also returns an opaque backendOperationId, which
light-a2a stores with its durable task correlation. A later status or
cancel context binds both that operation ID and the task ID; neither method
accepts an unsigned alternate target. This reconciliation operation is required
to meet the Phase 3 sidecar/backend restart guarantee rather than guessing the
outcome of work that survived a sidecar restart.
BackendCapabilities declares streaming, cancellation, status reconciliation,
accepted content modes, and other bounded features used to compile the Agent
Card. invoke_stream returns ordered status and artifact events and is callable
only when the published binding declares streaming support.
First-Release Sidecar Backend Transport
The first external-developer release defines one language-neutral private
application protocol, light-a2a-backend/v1. Its canonical source is a pinned
OpenAPI 3.1 document and referenced
JSON Schemas under
contracts/a2a-backend/v1/, with golden request, response, error, and event
fixtures. SDK models and documentation derive from that source; a language SDK
must not redefine the wire contract independently.
The required production transport is HTTP/1.1 with JSON over one fixed loopback
origin. Streaming uses text/event-stream on the same origin when the backend
declares support. The v1 surface is deliberately small:
GET /v1/capabilities
POST /v1/invoke
POST /v1/invoke-stream
POST /v1/status
POST /v1/cancel
GET /health/live
GET /health/ready
All business operations carry the short-lived signed
AuthorizedInvocation separately from developer-controlled business input.
The SDK validates its issuer, audience, expiry, deadline, replay identifier,
host, environment, agent, skill, operation, task/context/idempotency bindings,
backend operation ID when present, and policy and data-boundary digests before
calling business code. Loopback address or process placement is defense in
depth, not authentication. The configured origin, version, methods, and paths
are immutable; proxy environment variables, redirects, arbitrary destinations,
and wildcard backend listeners are rejected. The HTTP origin uses an explicit
loopback port in the valid TCP range 1..=65535; Portal rejects an invalid port
before publication rather than leaving it for runtime URL parsing.
Portal authoring selects only an approved backend transport profile. The
activated Config Server projection pins its contract version and digest,
transport profile, loopback origin, allowed methods and paths, signed-context
audience, timeouts, and resource limits for the target light-a2a instance.
Neither the external developer nor an A2A request may override those values.
HTTP over a peer- and filesystem-permission-protected Unix-domain socket may be qualified as an optional Linux hardening profile using the identical v1 semantics and fixtures. Its absence does not block the first external-developer release. Mutually authenticated private-network HTTP, gRPC, WebSocket, stdio, and in-process FFI or plugin transports are deferred. A backend requiring a different host is not treated as a local sidecar merely to bypass the remote A2A or future private-transport governance.
This private HTTP/JSON interface does not activate the public A2A HTTP+JSON
binding. light-a2a still terminates the selected public A2A JSON-RPC profile
and adapts it to the smaller backend contract. Making the business backend
implement A2A would defeat the sidecar’s purpose.
The backend accepts calls only from its sidecar and still validates business-domain invariants and durable business idempotency. It does not make platform authorization decisions, interpret Portal policy, or receive the caller’s raw token.
First-Release Backend SDKs
The first external-developer production release requires supported Python,
Java, and TypeScript/Node.js SDKs. A Rust implementation in crates/a2a-backend
is the reference adapter and conformance oracle, not a fourth externally
supported SDK release gate. Go, .NET, and other languages remain eligible after
the v1 contract is stable and demand justifies their compatibility burden.
| SDK | First-release requirement |
|---|---|
| Python | Production package, async unary/streaming adapter, examples, and full conformance evidence. |
| Java | Production library and standalone server adapter usable without reimplementing Light security, with full conformance evidence. |
| TypeScript/Node.js | Production package, unary/SSE adapter, examples, and full conformance evidence. |
| Rust | Checked-in reference adapter, golden-vector producer/consumer, and shared test harness. |
| Go, .NET, and others | Deferred; developers may use the published wire contract without a supported SDK claim. |
Each production SDK owns the local HTTP server adapter, signed-context and
replay validation, identifier equality checks, deadlines, limits, typed error
mapping, SSE framing, cancellation and status reconciliation, health endpoints,
artifact descriptors, and trace propagation. The developer implements only the
AgentBackend business callbacks. Generated models are useful but insufficient:
the thin handwritten runtime and all generated types must pass the same
cross-language conformance suite against the same light-a2a build.
Request Lifecycle
Handler Selection
The configured gateway handler chain resolves and admits the request before
forwarding it to the registered runtime selected by the published
implementation kind. LIGHT_AGENT routes directly to light-agent;
EXTERNAL_SIDECAR and REMOTE_A2A route to light-a2a. The edge handler
matches only configured public paths and their well-known card suffixes. It
must not treat every JSON POST as A2A traffic.
request
-> admission
-> handler-chain resolution
-> CORS where configured
-> JWT/session authentication
-> public A2A route and Instance API binding resolution
-> coarse authorization with the generated card or invoke policy endpoint
-> registered native or integration runtime routing
-> A2A version, content-type, body and operation validation
-> fine-grained A2A policy decision and obligations
-> backend selection and narrow delegation/credential injection
-> native light-agent execution or external integration invocation
-> response validation, filtering and protocol normalization
Agent Card Request
For a Portal-published card:
- Resolve the configured public host and path to an active Instance API
binding, its generated
cardpolicy endpoint, and disclosure class. - Authenticate before extended-card disclosure.
- Authorize coarse card access using that generated policy endpoint, then verify the publication generation is active and not revoked or replaced.
- Select the immutable card whose final public URL, optional fields, digest,
and
light-oauthsignature were accepted with the active publication. - Authorize disclosure before conditional-request evaluation, then attach an ETag derived from publication digest, disclosure class, applicable authorization-policy digest, and revocation epoch. Authenticated cards use private cache policy and never reuse an ETag across disclosure classes.
- Return the bounded card without a request-path signing call, starting business execution, or joining Portal authoring tables.
For a proxied upstream card:
- Fetch the card through the configured backend.
- Enforce status, content type, body size, and JSON-depth limits.
- Validate the declared versions, interfaces, and schemes against route policy.
- Rewrite only complete, approved interface URLs.
- Handle upstream signatures explicitly.
Rewriting a signed upstream card invalidates its signature. light-a2a must
either reject rewriting, publish a separately signed public facade card, or
remove the invalid signature and obtain an external-facade publication
signature from light-oauth. It must never forward an upstream signature over
mutated content. The controlled publication or cache-refresh path performs the
signing operation and caches the result; an ordinary Agent Card request does
not trigger signing. Signing keys never enter light-a2a, light-agent, the
gateway a2a-router, or Config Server runtime projections.
Message Or Task Request
- The gateway resolves the public route to an active
instanceApiId, verifies the publishedagentDefId, authorizes its generatedinvokepolicy endpoint, and selects the published implementation kind and a healthy registeredlight-agentorlight-a2ainstance. It does not accept any of those identities or a destination from the body. - The selected runtime independently validates the delegated binding identity,
then validates
A2A-Version, extensions, envelope, operation, IDs, and body limits through the shared A2A server modules. - Bind authenticated host and principal to the selected agent definition and environment binding.
- Authorize caller, calling agent, selected skill, abstract operation, tenant, data boundary, delegation depth, and budget.
- Validate context/task ownership for
GetTask,ListTasks,CancelTask,SubscribeToTask, and push-configuration operations. - For
LIGHT_AGENT, admit the operation directly into its durable native session/turn model. ForEXTERNAL_SIDECAR, create adapter correlation and a narrowly scoped, short-livedAuthorizedInvocation. ForREMOTE_A2A, select the pinned remote interface and server-owned credential or delegation. - Invoke the native agent, approved remote A2A server, or external business backend according to that binding.
- Validate the response or stream event within configured bounds, apply decision obligations such as redaction, and have the selected runtime scan, verify, classify, and materialize any artifact for which Light-Fabric promises managed retention.
- Apply gateway response filtering before returning data to the caller.
- Record edge, policy, protocol, backend, and runtime outcomes without logging sensitive content.
Streaming
Streaming is end-to-end. The gateway enforces generic edge limits; the selected runtime enforces A2A event framing, setup deadline, idle timeout, event limits, cancellation propagation, and disconnect behavior. Neither turns an ephemeral gateway stream into the authoritative business task record.
For a Portal-native agent, light-agent streams its durable turn events
directly through the embedded A2A server. A client that reconnects uses the A2A
task/context contract to recover state; it does not depend on the same gateway
process retaining a stream session. Incremental artifact chunks are bounded and
assembled under the selected runtime’s artifact policy; residual chunks are not
retained as an accidental second copy of the final artifact.
Portal Publication Model
Stable Identity
Use the existing Portal identity rule:
agentDefId == API version ID for the agent API asset
The A2A public path is a publication attribute, not a second agent identity. Task, audit, policy, skill, and memory records continue to refer to the stable agent definition ID.
The four related identities have separate purposes:
| Identity | Purpose |
|---|---|
apiVersionId / agentDefId | Logical, versioned agent catalog identity. |
instanceApiId | Binding of that agent API version to one Gateway instance; namespace for compiled edge policy. |
| Public host and path prefix | Human-readable routing and Agent Card interface identity. |
agt product and product version | Runtime software and configuration compatibility, not business-agent identity. |
Public Agent Card Metadata Authority
For public Agent Card metadata, authoritative means the Portal authoring record and deterministic precedence rule from which the publication compiler must obtain a value when several records could describe the same agent. It does not mean the Agent Card grants runtime authority. Cards remain descriptive; authentication, authorization, task ownership, and effective skills continue to come from policy.
There are three metadata authority stages:
- Portal authoring records are the source of truth for edits and review.
- The immutable, digest-bound A2A publication freezes the effective values.
light-agentorlight-a2aserves that accepted publication without joining Portal tables or allowing runtime configuration to override individual fields.
Use this source and precedence contract:
| Agent Card field | Authoritative Portal source | Precedence and validation |
|---|---|---|
| Name | api_t.api_name | Required. The A2A binding cannot rename the agent. |
| Description | api_version_t.api_version_desc, then api_t.api_desc | Use the nonblank version description first and the nonblank API description as fallback. A selected publication profile may require a result. |
| Provider | Referenced public provider profile, then the host’s default public provider profile | Provider name and URL come from an approved structured profile. Absence is allowed only when the selected A2A profile permits it. |
| Documentation URL | Version-scoped agent public metadata | Must be an approved absolute public URL. It is not inferred from api_version_t.spec_link or api_t.git_repo. |
| Icon | Version-scoped managed iconAssetId | The compiler resolves the asset to an approved absolute public URL and validates scheme, media type, size, and availability. Arbitrary remote icon URLs are not accepted. |
| Semantic version | api_version_t.api_version | This is the business-agent version and must pass the selected publication profile’s SemVer validation. |
These versions are independent and must not be substituted for one another:
api_version_t.api_version business-agent semantic version
A2A-Version wire-protocol version
agt product version runtime software/config compatibility
publication version immutable control-plane generation
model or model-policy version model selection, not agent identity
In particular, agent_definition_t.model_provider identifies the LLM vendor
or model-selection provider. It must never populate the Agent Card provider,
which identifies the organization responsible for the published agent.
Add structured logical authoring records such as:
public_provider_profile_t
host_id
provider_profile_id
provider_name
provider_url
provider_description
aggregate_version
active
agent_public_metadata_t
host_id
agent_def_id # API version ID
provider_profile_id
documentation_url
icon_asset_id
aggregate_version
active
The preferred first implementation also adds a nullable structured
publication_alias column to skill_t, with a normalized partial unique
constraint for active aliases within host_id. Private skills need no alias
until selected for publication. The publication compiler requires one before a
skill can appear in an Agent Card and freezes the alias after its first
successful use. A schema-validated JSON extension must not substitute for this
queryable identity field.
Portal also owns a managed extension registry. A logical first schema is:
a2a_extension_t
host_id
extension_id
extension_uri # exact, versioned public identity
display_name
extension_version
extension_class # DATA, PROFILE, METHOD or STATE_MACHINE
lifecycle_status # EXPERIMENTAL, APPROVED, DEPRECATED or REVOKED
allowed_directions # INBOUND, OUTBOUND or BOTH
required_eligible
allowed_operations
parameter_schema # schema-validated JSONB
parameter_schema_digest
handler_ref
handler_digest
maximum_metadata_bytes
security_review_ref
aggregate_version
active
a2a_extension_dependency_t
host_id
extension_id
dependency_extension_id
required
aggregate_version
active
Portal View manages these records through structured registry forms and lets an
A2A Binding select only active, direction-compatible entries. The URI,
classification, directions, lifecycle, handler, review, and required
eligibility are structured fields; only schema content and schema-validated
extension parameters use JSONB. The binding form does not accept an arbitrary
URI or enable required when the registry record is not required-eligible. The
initial production registry may contain reviewed draft records, but no binding
can activate or advertise them until a later extension profile is qualified.
The host also selects at most one active default public provider profile. Exact
DDL names may follow Portal conventions, but provider profiles must be reusable,
tenant-scoped, versioned, and soft-deletable, while agent public metadata must
be version-scoped through agentDefId. These fields are structured because
Portal View must validate, query, review, and audit them. JSONB remains limited
to schema-validated extension metadata.
The A2A binding references these records and selects disclosure; it does not
duplicate or freely override their values. For REMOTE_A2A, an upstream card
may seed a reviewed draft and remains provenance input, but it does not replace
Portal authority or silently update an active facade card.
Portal View manages name and general description on the API form, semantic version and version description on the API Version form, reusable provider profiles in Host Administration, and provider selection, documentation URL, and icon in an Agent Public Metadata panel. The A2A Binding workflow shows a read-only effective-card preview with the source of each field. Publication fails rather than inventing a value when required metadata is missing, ambiguous, inactive, invalid, or no longer accessible.
Logical A2A Publication
Portal needs one versioned logical publication per exposed agent and environment. The physical schema uses normalized authoring tables plus an immutable versioned JSONB publication aggregate and compiled runtime projections. This is a deliberate hybrid, not a choice between an editable table and an editable JSON document.
The normalized authoring model contains the provider and agent-public-metadata
records above plus a core agent_a2a_binding_t, with
agent_a2a_interface_t, agent_a2a_access_grant_t, and
agent_a2a_disclosure_t child relations for repeating interfaces, grants, and
disclosure selections. agent_a2a_publication_t stores the immutable compiled
manifest and agent_a2a_instance_publication_t records application and rollback
for a target runtime instance. Exact DDL names may be aligned with Portal schema
conventions during implementation, but these logical ownership boundaries are
settled.
The normalized model gives Portal View queryable columns, foreign keys, uniqueness constraints, optimistic aggregate versions, and soft-delete or revocation behavior. JSONB is limited to explicitly schema-validated extension options, validation evidence, source aggregate-version maps, and immutable compiled publication content. Raw credentials and signing-key material are never stored in the binding; only server-owned references are accepted.
The logical contract contains:
hostId,agentDefId/apiVersionId, GatewayinstanceApiId, API version, and environment;- implementation kind (
LIGHT_AGENT,EXTERNAL_SIDECAR, orREMOTE_A2A); - environment-specific agent binding ID and network zone;
- publication ID, version, digest, lifecycle, validity, and revocation epoch;
- unique public path prefix and allowed host names;
- generated coarse policy endpoints for card and invocation admission;
- public and extended visibility rules;
- compiled provider, documentation, icon, business-agent version, input modes, and output modes, with source-record identities and aggregate versions;
- supported binding/version/interface declarations in preference order;
- capability declarations, including streaming and push support;
- exact advertised, inbound, outbound, and required extension selections with registry version, handler/schema digests, dependencies, operations, and metadata limits;
- security schemes and requirements;
- selected public AgentSkill projections and their immutable
publicationAlias, internalskillId, skill-version, and skill-digest mappings; - backend kind, service identity, environment, path, and TLS policy;
- server-owned credential or delegation policy references;
- signing-profile ID, signature policy, signed-card JWS/JWKS metadata, and signing audit reference; and
- source aggregate versions used to compile the publication.
The publication compiler rejects missing required card fields, ambiguous paths,
duplicate skill IDs, unsupported protocol combinations, insecure public
interfaces, unresolved backend identities, and secret material embedded in
card content. It also rejects an inactive or mismatched Gateway Instance API
association, a normalized public host-and-path collision, a generated policy
key collision, or any route whose instanceApiId, apiVersionId, and
agentDefId do not form one consistent binding.
Portal Authoring And Publication Workflow
Portal View exposes an explicit Publish through A2A handoff from Agent registration or Agent detail. The handoff opens the A2A Bindings action or workspace; completing ordinary Agent registration never creates a public A2A route implicitly. Because an agent can have zero or more environment-specific bindings, the primary UI is a table with binding name, environment, implementation kind, deployment mode, public path, selected profiles, target kind, visibility, validation status, publication version/state, and last update. Create and update use structured, conditional forms for interface, backend, security, disclosure, access, and limit fields. Compiled JSON is a read-only preview; only explicitly supported extension fields may use a schema-validated advanced JSON editor.
Create, update, and delete commands modify Draft authoring state through the normal Portal command, CloudEvent, and query-projection path. They do not mutate live runtime configuration. An explicit validate-and-publish workflow:
- reads a consistent set of API, API-version, public-provider, agent-public-metadata, binding, agent, skill, disclosure, policy, backend, and target-instance projections;
- validates references, skill-alias uniqueness and stability, extension registry/direction/dependency/required-eligibility rules, and records every source aggregate version;
- verifies or creates through the authorized deployment workflow the active
Gateway
instance_api_tassociation and its unique public path prefix; - compiles the immutable unsigned Agent Card with final environment-specific public URLs and optional fields, plus the publication manifest, content digest, public-skill mappings, generated deployment-scoped policy endpoint keys, and audience-specific Config Server property sets;
- invokes
light-oauthwith the authorized host, environment, and signing profile;light-oauthvalidates the profile purpose, canonicalizes the card without existing signatures, signs it with the current profile key, and returns the Agent Card signature,kid, and JWKS location; - verifies the returned signature against the selected profile JWKS and stores the complete signed card, signing audit reference, and canonical digest in the immutable publication;
- emits the Config command/events that stage the property sets in
instance_property_tfor the target registered instances; - creates and validates immutable
config_snapshot_tConfig Server snapshots for every target(host, serviceId, envTag); - activates the release manifest and its exact target-to-snapshot mapping;
- asks the Controller to reload each target by
host,serviceId, andenvTag; - each runtime calls
/configs, validates the selected immutable snapshot, and atomically applies it or retains its still-valid last-known-good generation; and - records each applied or rejected snapshot ID and digest for Portal diagnostics.
Portal may update all current pointers atomically in its database, but
independently operated Gateway, light-agent, and light-a2a processes observe
and apply them at different times. A release therefore requires compatible
adjacent generations or an explicit staged protocol, per-target reload and
acknowledgement, and exact-generation rollback. It must not claim instantaneous
cross-service activation.
Retiring a published binding creates and activates a new generation without the route and, when immediate invalidation is required, advances the revocation epoch. Historical publications remain immutable for audit and bounded rollback; Portal does not hard-delete the active runtime contract.
Portal View manages A2A artifact-retention profiles through structured fields
for transient, content, task-visibility, and metadata periods, external-reference
handling, scanning, and memory-promotion posture. A host default may be selected
and an agent binding may select an approved override. The effective runtime
projection retains the selected profileId together with its compiled rules so
admission evidence can identify the authoritative profile without querying the
control plane. Artifact access itself
continues to use the existing fine-grained access-control authoring and policy
projection; the artifact form does not introduce a parallel ACL or special
administrator bypass. Publication validates the combined retention and access
policy, freezes its digest into the runtime generation, and exposes the
effective read-only result in the compiled preview.
Runtime Projections
Portal publishes separate least-privilege projections for the gateway and the
selected runtime. The gateway projection contains only public route admission,
the active instanceApiId, agentDefId/apiVersionId, generated card and
invoke policy endpoints, implementation kind, and registered target-service
routing data. Its combined rule.endpointRules uses those generated policy
endpoints rather than the repeated raw A2A specification paths. For
EXTERNAL_SIDECAR and REMOTE_A2A, the light-a2a projection contains Agent
Cards including accepted publication signatures, fine-grained policy, backend
bindings, trust, credential references, logical signing-profile metadata, and
protocol, artifact-access, and artifact-retention limits. For LIGHT_AGENT,
the existing agent.yml projection retains
prompt, model, agentPolicy.skills,
agentPolicy.catalog.effectiveCatalog, memory, knowledge, tool, and execution
policy and adds the native inbound a2aPolicy, including the same artifact
policy schema, and accepted signed card required to serve that publication.
Portal stages these compiled values in
instance_property_t; snapshot creation copies the candidate values into the
immutable Config Server generation selected by host, service ID, and environment
tag. Property definitions or an editable instance row alone do not constitute a
published runtime contract.
GET /configs?host&serviceId&envTag is the only current-workload configuration
path. The Config Server does not resolve mutable A2A authoring data, does not
serve instance_property_t directly, and does not support an A2A-specific
runtime-config endpoint or a secondary instanceId, productId, or
productVersion lookup mode. Publication alone does not hot-load a process;
the explicit Controller reload causes the runtime to call /configs again.
The selected runtime renders public metadata only from the accepted publication; an environment variable, upstream card refresh, or backend response cannot replace an individual name, description, provider, documentation, icon, or business-version field.
Every runtime audience projection is immutable and digest-bound. The runtime checks:
- audience matches the selected runtime (
agentorlight-a2a); - host, service ID, and environment tag match the running target service;
- schema version and compatibility generation are supported;
- content digest and publication digest match canonical content;
- activation-time, generation, and revocation constraints pass;
- every profile is single-generation, each 0.3 profile carries no extension configuration, and each 1.0 profile’s advertised, inbound, outbound, and required extension sets match the accepted card and registry digests with required entries eligible, implemented, and dependency-complete; and
- every agent route is unique after path normalization.
The Gateway rejects a projection when two routes have the same normalized
public host and path, two Instance API bindings produce the same policy
endpoint, or a route’s policy endpoint is not owned by its declared
instanceApiId. Route resolution supplies these trusted identities to access
control; request headers, query parameters, and bodies cannot override them.
Activated runtime policy remains valid until explicitly revoked or replaced by
an activated publication. Publication failure retains the last-known-good
generation; elapsed time does not expire it. Legacy refreshAfter and expiresAt
fields are accepted and ignored, and new runtime envelopes omit them. Revocation,
identity, digest and generation checks still apply. Request tokens, invocation
deadlines and independently reviewed outbound credentials retain their own expiry
rules. Runtime request handling does not query Portal authoring or projection tables.
Agent Card And Portal Skill Mapping
Portal skills are richer than A2A Agent Skills. Publish a deliberate projection:
| A2A AgentSkill field | Portal source |
|---|---|
id | Stable, tenant-scoped skill_t.publication_alias; never the Portal UUID. |
name | Approved skill_t.name. |
description | Approved public description, not instruction Markdown. |
tags | Policy-filtered tags and selected category paths. |
examples | Explicit reviewed examples; never inferred from private history. |
inputModes | Publication policy or approved capability metadata. |
outputModes | Publication policy or approved capability metadata. |
| security requirements | Publication policy intersection for that skill. |
Only active skills assigned to the exact agent definition and approved for the selected disclosure class are candidates. Assignment alone does not make a skill public. The compiler rejects an absent, duplicate, normalized-colliding, or previously rebound alias. A compatible skill revision keeps its alias while the immutable publication records the new version and digest.
Do not expose:
contentMarkdownor internal prompt instructions;- tool IDs, schemas, backend paths, workflow bindings, or execution placement;
- skill configuration supplied for one tenant or user;
- semantic embeddings or internal ranking scores;
- approval rules, cost limits, or private policy diagnostics; or
- anything obtained from memory, session history, tool output, or model text.
The internal effective catalog remains the source for agent-side progressive disclosure and tool selection. The A2A Agent Card is a smaller interoperability surface, not a replacement catalog.
See MCP Tool Metadata Usage and Centralized Agentic Skill Registry for the internal catalog and execution-placement boundaries.
Skill Runtime Projection And Executable Packages
The public Agent Card skill list and the internal agent skill projection have different purposes. The card contains bounded discovery metadata. The runtime projection contains the assigned instruction content and the independently authorized tool and workflow descriptors required by one Agent publication.
For LIGHT_AGENT, Config Server is the runtime authority. At startup or an
explicit reload, light-agent resolves the current immutable generation into
agent.yml, validates the envelope and content digests, compiles the projected
skill Markdown into its system instructions, and caches the projected effective
catalog locally. It does not call genai-query/getEffectiveAgentCatalog on the
request path or treat live Portal authoring rows as a fallback. At design time,
the template and strict runtime loader exist; completing the Portal compiler,
snapshot activation, acknowledgement, and last-known-good reload path remains
implementation work owned by the publication phases below.
genai-query/getEffectiveAgentCatalog remains useful on the control plane for
Portal View preview, assignment validation, publication compilation,
administrative diagnostics, and semantic ranking. Its current live authoring
result is not a runtime authority. If a future catalog is too large for a
bounded Config Server property, the snapshot may contain an immutable catalog
artifact URI and digest. A future runtime search API must be scoped by
publication ID and content digest and may only rank or narrow entries already in
that immutable manifest; it must not add a capability from newer authoring
state.
A skill is discovery, instructions, and capability composition. Executable behavior belongs to a governed tool, workflow, fixed service, or reviewed skill package. Use the following distribution boundaries:
| Content | Runtime distribution and execution |
|---|---|
| Skill alias, public metadata, bounded instruction Markdown, version, and digest | Immutable Config Server projection. |
| Tool aliases, schemas, stable references, execution placement, and policy bindings | Immutable Config Server projection; runtime intersects them with live Gateway or runner authority. |
| API or MCP implementation | Deployed backend invoked through the governed Gateway path; code is not downloaded by the agent. |
| Workflow | Immutable workflow identity, version, and digest in the projection; execution remains in light-workflow. |
| Python, JavaScript, WASM, plugin, binary, template, or other package assets | Reviewed, scanned, signed, content-addressed artifact storage; a trusted runner verifies and stages them under the selected sandbox policy. |
| Credentials and secrets | Server-owned secret or delegation service; never skill content, Config Server values, Agent Cards, or hybrid-query results. |
Config Server may carry an immutable artifact reference, digest, media type,
entrypoint, supported runtime profile, and sandbox-policy reference. It must not
carry large package bytes or mutable source for in-process evaluation. Existing
tool_t.script_content is an authoring/legacy input only: production publication
packages it as a signed artifact with runner placement instead of delivering it
through Config Server or allowing light-agent to retrieve and execute it from
a live Portal query.
Artifact verification and staging complete before a new configuration generation becomes active. Failure retains only a still-valid last-known-good generation. Sessions and tasks remain pinned to their accepted publication and skill digests so an alias cannot change meaning midway through a turn.
Portal-Native light-agent Integration
Native A2A Server Integration
The current light-agent public interaction is a /chat WebSocket. A2A has
message, task, streaming, lookup, cancellation, version, and error semantics.
Putting those semantics into light-gateway would make the gateway a second
agent runtime and create non-durable state that cannot survive routing changes
or gateway restarts. Putting a light-a2a sidecar beside light-agent would
instead add a redundant network hop and split task and policy ownership across
two managed processes.
light-agent therefore embeds the shared A2A server, card, policy, and task
modules and maps their abstract operations directly onto its durable domain
operations. It does not emulate its browser /chat WebSocket client and does
not call a sidecar. A2A wire and policy code lives in shared A2A crates; durable
Light session, turn, action, memory, and model-loop ownership remains in
light-agent.
Identity Mapping
| A2A concept | Light runtime mapping |
|---|---|
| Published agent | host_id + agent_def_id + definition_version. |
| Authenticated caller | Bound principal and optional user from validated gateway delegation. |
contextId | External alias for a durable agent_session_t row scoped to host, principal, agent, and policy. |
| A2A task ID | External alias for a durable agent turn/job. |
| Message ID | Idempotency key for durable turn admission. |
| Task status | Projection of durable turn/action state into A2A task state. |
| Artifact | Bounded projection of durable result/artifact metadata. |
External identifiers may be opaque gateway-safe IDs rather than raw database UUIDs. Every lookup must include authenticated ownership and agent binding.
State Mapping
Define an explicit, tested mapping rather than string substitution. For example:
| Light state | A2A task state |
|---|---|
| queued/received | submitted |
| running model/action/reconciliation | working |
| waiting approval or additional user input | input-required or auth-required as appropriate |
| completed | completed |
| failed | failed |
| cancelled | canceled |
| policy or agent refusal before/during work | rejected |
| unknown/operator-required with indeterminate outcome | unspecified; never reported as completed or failed without evidence |
For A2A 1.0 these correspond to TASK_STATE_REJECTED and
TASK_STATE_UNSPECIFIED; compatibility profiles map to their version-specific
wire names. The exact vocabulary must be pinned to the selected A2A version.
Lossy mapping preserves the original Light state in internal audit, not in an
ungoverned public extension.
Memory Boundary
The embedded A2A server binds the authenticated session identity before native
light-agent domain admission. light-agent selects and validates the memory
bank, loads session history, recalls relevant memory, and retains accepted
experience according to its immutable memory policy.
Recalled memory remains untrusted context. It cannot change A2A routing, published skills, authorization, destination, credentials, or task ownership. See Hindsight Memory.
Governed Outbound A2A
Outbound A2A requires a catalog binding distinct from public inbound publication. An external agent registration should include:
- stable host-scoped external-agent or API-version identity;
- discovered Agent Card URL and last accepted card digest;
- discovery time, expiry, trust status, reviewer, and revocation state;
- selected protocol binding and version;
- approved destination and redirect policy;
- credential/delegation policy reference;
- allowed calling agents, principals, environments, and data classifications;
- card signature verification state and trust anchor;
- declared skills and a policy-filtered internal search projection;
- declared extensions, required flags, selected allowlist decisions, dependencies, and implementation/schema digests; and
- connection, request, stream, concurrency, and cost limits.
Discovery is an onboarding or refresh workflow, not a request-path operation. The refresh worker fetches the card with restricted egress, validates it, records provenance and signature results, computes a digest, and produces an approved runtime binding. A changed card remains pending until automatic policy or human review accepts the new capabilities and destination. A newly declared required extension, an optional-to-required transition, a changed extension URI or dependency, or a previously approved extension becoming deprecated or revoked always requires review. An unapproved required extension makes the binding non-executable; it is never passed through transparently.
Governed outbound is mandatory for the first production milestone. At runtime
the calling light-agent chooses a stable Portal agentRef from its effective
catalog and sends the call through the published Gateway and light-a2a path.
The model, workflow payload, and caller never select the Agent Card URL, target
service, credential reference, or physical destination. light-a2a resolves
the approved binding, verifies its active trust and revocation state, applies
server-owned credentials or delegation, and enforces the calling principal,
calling agent, target agent, operation, skill, environment, data boundary,
delegation depth, budget, and task/context policy.
Each outbound catalog entry has a Portal-generated UUID catalogToolId. Its
human-readable logical alias may contain dots, but it is a separate field and
must not be substituted for this stable evidence identifier. The shared
delegation-depth wire type is an unsigned 16-bit integer; Portal authoring and
publication therefore accept only integral values from 1 through 65535.
The initial outbound production profile is deliberately bounded to the same selected JSON-RPC message, streaming, task lookup, subscription, and cancellation capabilities qualified for inbound use and supported by the target binding. Arbitrary runtime discovery, model-selected destinations, push notifications, public A2A HTTP+JSON, public A2A gRPC, and custom bindings remain deferred.
Workflow call: a2a Migration
The existing Workflow DSL model has optional author-supplied agentCard and
server fields even though light-workflow currently rejects call: a2a as
unimplemented. Those fields must not become an escape path around this design.
Before enabling execution, add a required stable Portal catalog agentRef and
compile it to an approved light-a2a binding. In governed mode validation must
reject agentCard, raw server URIs, embedded credentials, and any combination
of legacy destination fields with agentRef. If legacy syntax must remain
parseable for schema compatibility, it stays non-executable and produces a
specific migration error. Workflow runtime and model-generated data never
select the physical endpoint.
Authentication, Authorization, And Delegation
Inbound
- Public card access may be anonymous only when publication policy says so.
- Extended cards and A2A operations require the declared authentication scheme.
- Gateway JWT/session authentication binds host, principal, environment, and audience before A2A authorization.
- Gateway route resolution binds the request to an active
instanceApiIdand uses its generatedcardorinvokeendpoint inrule.endpointRulesfor coarse per-agent admission. Portal-generated identities, not caller data, populate the rule context. - The selected
light-agentorlight-a2aruntime independently validates the delegated identity and evaluates the stable agent, skill, abstract operation, tenant, task/context ownership, data-boundary, delegation-depth, budget, and limit policy. - Native
light-agentadmits the authorized operation directly.light-a2amints a remote credential orAuthorizedInvocationbound to the target agent, operation, context/task when known, expiry, environment, and policy digest only for its external integration path. - The external Authorization header is never blindly forwarded when a server-owned backend credential or delegation is required.
Outbound
- The calling agent must have the external agent assigned or otherwise allowed by its immutable policy snapshot.
light-a2aauthorizes both the caller’s agent identity and originating human/workload principal after gateway edge admission.light-a2aresolves credentials from server-owned references and scopes them to the configured destination.- Delegation tokens are short-lived, audience-bound, non-replayable where required, and excluded from logs and Agent Cards.
Task Ownership
Authorization for GetTask, CancelTask, resumption, or push configuration
must prove that the task belongs to the authenticated caller or that the normal
fine-grained policy explicitly grants the requested operation. Artifact
metadata, content, download, export, deletion, and memory promotion apply the
same rule. No Portal role or operator identity receives implicit task or
artifact authority. Possession of a task ID, context ID, artifact ID, or URL is
never sufficient.
Fine-Grained Decision Contract
The A2A policy engine evaluates an explicit tuple:
caller principal and calling agent
x target agent and selected skill
x A2A operation
x host, tenant, environment and network zone
x task/context/artifact ownership
x data classification and boundary
x delegation depth, deadline, budget and rate limits
Do not collapse every operation into one a2a.invoke permission. At minimum,
distinguish card read, extended-card read, message send, message stream, task
read, task list, task cancel, task subscribe, and each push-configuration
operation. Artifact metadata read, content read or download, export, deletion,
and promotion to memory are also distinct operations; an allow decision for
task read does not automatically grant every artifact operation.
This requirement applies to the selected runtime’s authoritative A2A policy.
The Gateway’s generated invoke policy endpoint is only coarse admission to a
particular agent and does not replace or satisfy any operation-specific runtime
decision.
The existing delegation contract is tool- and knowledge-oriented. A2A support must add versioned, operation-specific delegation kinds with target agent, skill, task/context, tenant, policy digest, data-boundary digest, replay, and expiry bindings. A generic path or arbitrary-operation grant is not acceptable.
Security And Privacy
Agent Card Poisoning
- Validate every card field and URI against the selected version and binding.
- Reject credentials, inline secret material, unsupported required extensions, and internal-only addresses in public interfaces.
- Preserve source digest, signature verification, and review provenance.
- Treat descriptions, examples, tags, and extension parameters as untrusted display/model context.
SSRF And Destination Safety
- Resolve only server-owned service identities or approved remote bindings.
- Deny link-local, loopback, private, or metadata-service addresses unless the backend is explicitly admitted for that network zone.
- Revalidate resolved addresses and redirect destinations.
- Apply TLS hostname and trust policy after final target resolution.
- Never let card refresh or runtime fallback select an undeclared interface.
Prompt And Capability Injection
Agent Cards, external agent results, skills, memories, and artifacts are data. They cannot add tools, elevate execution placement, widen network access, change credentials, or override system and policy instructions.
Resource Exhaustion
Enforce independently configurable limits for:
- request, response-inspection, Agent Card, message-part, artifact, and event sizes;
- JSON depth, collection size, extension count, and interface/skill count;
- total and per-principal concurrent requests and streams;
- header bytes and header count;
- request, upstream setup, stream idle, and total task wait time; and
- card cache entries, retained runtime generations, and telemetry label cardinality.
Filtering
Request access control may inspect bounded A2A message metadata and content only when the endpoint policy enables body access. Response filtering applies to bounded JSON responses and individual bounded stream events. The gateway must not buffer an unbounded stream to run a whole-response filter.
Agent Card URL And Signature Rules
Portal-published cards should contain their final public URL before signing.
Public host and scheme come from approved publication configuration, not
untrusted Host or forwarding headers.
Signing Authority And Identity
light-oauth is the first-production signing and JWKS authority for Light
Agent Card publications. This reuses the platform’s existing operational
boundary for asymmetric signing, kid selection, public-key distribution, and
rotation, but it does not reuse an OAuth access-token or long-lived-token key
as an Agent Card key. OAuth JWT issuance and Agent Card publication are
different cryptographic purposes with independent compromise, rotation,
revocation, and audit boundaries.
The first implementation extends light-oauth; it does not introduce a second
network service solely for A2A keys. Reusable signing-profile, KMS/HSM adapter,
JWS, and JWKS lifecycle modules should remain separable so a future general
light-signing service can be extracted if additional platform artifacts adopt
the same authority. That extraction must preserve profile IDs, trust URLs, and
audit semantics and is not an A2A production prerequisite.
The default Agent Card signing identity is the tuple:
host/tenant + environment + publication purpose
The initial purposes are:
| Purpose | Signs | Meaning to a verifier |
|---|---|---|
A2A_CARD_NATIVE | Native LIGHT_AGENT publications | The named Light host and environment approved this native Agent Card publication. |
A2A_CARD_EXTERNAL_FACADE | EXTERNAL_SIDECAR and rewritten REMOTE_A2A facade publications | The named Light host and environment validated and published this governed external facade; it does not claim that the upstream vendor signed the rewritten content. |
The signing profile is the issuer identity, a rotating kid identifies one key
under that profile, the individual agent publication is the signed subject, and
the light-agent or light-a2a process is only the serving runtime. Runtime
fleet, replica, and instance IDs never define the issuer because topology may
change without changing publication ownership. One key per agent publication
is not the default; an independently delegated per-agent profile is an optional
high-assurance override with its own lifecycle and approval.
The native and external-facade profiles use separate key rings even when they
belong to the same host and environment. A caller authorized for one purpose
cannot request the other purpose, select an arbitrary kid, or use an OAuth
provider key. light-oauth chooses the current key after resolving the
authorized profile. If the signature includes jku, it points to the stable
public JWKS for that exact profile; verifiers must trust the profile authority
and must not treat possession of any key published by the same service as
equivalent. Runtime verification pins the expected light-oauth origin and
profile from the accepted control-plane projection. It never follows an
arbitrary card-provided jku; when jku is present, it must equal the projected
profile JWKS URL.
Signing Profile And Key Lifecycle
Portal adds logical signing-profile and signing-key records. Exact DDL names may follow Portal conventions during implementation, but the model contains:
signing profile
hostId, environment, profileId, purpose, algorithm, jwksUrl
rotationPolicy, validity, revocationEpoch, active
signing key
profileId, kid, publicJwk, privateKeyRef = managed:<logical-alias>
state = CURRENT | PREVIOUS | REVOKED
validFrom, validUntil, rotation/revocation audit
privateKeyRef is the server-created managed:<logical-alias> reference. The
Portal form accepts only the logical alias; it never accepts a path, URI, PEM,
or provider credential. light-oauth resolves that alias below its configured
and mounted A2A key root and rejects path traversal, unsafe permissions, and
non-canonical files. Production A2A private keys are not stored as plaintext
Portal values or exposed through Config Server, Agent Cards, logs, or
telemetry. A future KMS/HSM provider may resolve the same logical alias behind
light-oauth; existing OAuth provider keys remain OAuth-owned and are not
repurposed for A2A.
Portal View exposes a structured Signing Profiles table under the host and
environment administration boundary. The form manages purpose, algorithm,
approved managed-key provider or logical alias, rotation policy, validity,
status, and revocation. It displays the platform-derived JWKS URL,
current/previous kid values, and audit history but never private material. A2A
bindings inherit the environment’s default native or external-facade profile;
selecting a non-default or per-agent profile requires an explicit authorized
override. Raw JSON is not the primary editor.
Key generation, scheduled or administrative rotation, retirement, and
emergency revocation use the normal Portal command, event, projection, and
audit path. The light-oauth public JWKS contains the current key and previous
keys only for the documented verification overlap. Normal rotation signs new
publications with the new current key only after that public key is retrievable
from the profile JWKS. It retains old public keys through the maximum publication
validity, card-cache, and rollback windows. Emergency revocation removes the key
from the published JWKS, advances the affected revocation epoch, immediately
blocks local serving of cards that depend exclusively on that key, and requires
a new signed publication before service resumes. External verifiers observe the
removal within the documented JWKS cache bound. Multiple Agent Card signatures
may be assembled by the controlled publication workflow during rollover when
supported by the selected A2A profile.
Signing And JWKS Service Contract
light-oauth exposes a purpose-specific Agent Card signing operation and a
public profile JWKS operation. Exact HTTP paths are finalized with the
light-oauth OpenAPI, but the semantic operations are:
SignAgentCard(profileId, publicationId, cardKind, finalCardWithoutSignatures)
GetSigningProfileJwks(profileId)
SignAgentCard is authenticated and fine-grained-authorized. It resolves host,
environment, purpose, publication authority, algorithm, and current key from
server-owned state; validates that the workload may use the profile for the
publication; binds the authorization to agentDefId, publicationId, purpose,
and canonical content digest. cardKind is required and is exactly
AGENT_CARD or EXTENDED_AGENT_CARD; it selects the corresponding immutable
card in the prepared publication manifest, and the service rejects a payload
that matches the other card kind. The operation requires the signing payload
to omit existing signatures; validates and JCS canonicalizes the final Agent
Card; and returns
the A2A JWS signature, kid, JWKS location, canonical digest, and audit
reference. It is not a generic
arbitrary-byte or caller-selected-key signing endpoint. The public JWKS
operation exposes only verification material and applies bounded cache headers
compatible with rotation and emergency revocation.
An independently verified upstream or rollover signature is retained outside
the SignAgentCard request and may be assembled into the final signatures
array only when it covers the same canonical digest. It is never treated as
input authority for selecting the Light signing profile or key.
The normal publication workflow calls this operation after all public URLs and
optional fields are final, once for AGENT_CARD and, when present, once for
EXTENDED_AGENT_CARD. It verifies each result and projects the complete signed
cards. light-agent and light-a2a validate the signature and digest when
activating a projection and then serve the accepted immutable card without a
request-path light-oauth call. An explicitly approved activation-time fallback
may call light-oauth with a logical profile ID when the final card can only be
constructed at that boundary; it caches the result and still never receives a
key or key reference. light-gateway never calls the signing operation.
For transparent proxy cards:
- rewrite only an interface URL whose origin and complete path match the configured backend agent base;
- preserve an approved relative subpath and query according to binding rules;
- reject interface URLs that escape the backend binding;
- never partially match path segments;
- remove stale
Content-Length, ETag, and upstream signature after mutation; - generate a new ETag from final canonical content, disclosure class, authorization-policy digest, and revocation epoch; and
- obtain a new signature only through the configured
A2A_CARD_EXTERNAL_FACADElight-oauthprofile.
Legacy top-level url and current supportedInterfaces are handled by separate
version profiles. A malformed hybrid card is rejected rather than guessed.
Errors And Failure Mapping
Use stable internal error codes plus binding-correct A2A errors. At minimum, distinguish:
- unsupported A2A version, binding, operation, or media type;
- invalid or unavailable explicitly activated extension;
- missing client activation of a published required extension in the 1.0
profile as
ExtensionSupportRequiredError; - invalid JSON-RPC envelope or A2A payload;
- unknown or unauthorized agent, context, or task;
- task not cancelable;
- public or extended card unavailable;
- card signature or publication validation failure;
- request or response size/depth limit exceeded;
- backend unavailable, timeout, protocol violation, or invalid agent response;
- access-control denial and response-filter denial; and
- replaced or revoked publication.
Do not translate a JSON-RPC error returned with HTTP 200 into success telemetry. Conversely, do not expose internal topology, database state, policy expressions, credential identifiers, or parsing details in public error messages.
Observability
Traces And Logs
Record bounded, low-cardinality fields such as:
a2a.versionanda2a.binding;a2a.operationand safe JSON-RPC method;- publication ID/version/digest prefix;
- signing profile purpose, safe profile identifier,
kid, and signing or verification outcome; - stable agent definition or external-agent reference;
- route and backend kind;
- response outcome and A2A error class/code;
- result kind and task state;
- hashed or otherwise policy-safe context/task correlation;
- streaming/non-streaming mode;
- config generation; and
- request, upstream, first-event, and total duration.
Do not log prompts, message parts, artifacts, memory content, credentials, complete JWTs, or high-cardinality raw task/context IDs by default.
Metrics
Provide counters and histograms for:
- requests by operation, version, binding, route, and outcome;
- card serves, upstream fetches, cache hits, validation failures, and signature outcomes;
- active streams, stream setup/idle failures, events, and disconnects;
- task create/get/cancel outcomes;
- backend latency and protocol violations;
- authorization and filtering denials;
- rejected config reloads and last-known-good retention; and
- admission, concurrency, body-size, and timeout rejections.
Agent ID, task ID, context ID, user ID, and arbitrary skill names must not become unbounded metric labels.
Reload, Caching, And Availability
- Compile and validate a complete generation off the request path.
- Keep every target snapshot internally complete and atomic. Coordinate the cross-service release through one exact target-to-snapshot manifest, compatible adjacent generations, explicit reload, per-target acknowledgement, and exact-generation rollback.
- Keep the previous valid generation when a refresh is malformed.
- Fail closed when a publication is revoked; do not expire activated runtime policy.
- Cache final public cards by publication digest, disclosure class, authorization-policy digest, and revocation epoch.
- Cache each pinned signing-profile JWKS only for its bounded cache lifetime.
Refresh on an unknown
kidduring projection activation; if the trusted JWKS is unavailable or the key/profile check fails, reject the new generation and retain only a still-valid last-known-good generation. Agent Card requests do not perform JWKS network fetches. - Use ETag/conditional requests without allowing stale authenticated disclosure after authorization or revocation changes.
- Preserve only bounded edge correlation in gateway memory.
- Store external adapter correlation durably in
light-a2awhen required. Nativelight-agentstores its A2A context/task aliases with its durable session and turn state; the selected agent runtime remains authoritative for business task state. - A gateway restart must not change task ownership or make a durable task permanently unreachable.
Implementation Phases
These phases describe implementation order, not optional production scope. Inbound paths may be enabled first in development and canary environments, but the first production release requires the applicable Phase 0 through Phase 5 exit gates. Phase 6 remains opt-in: only the independently qualified external-sidecar profile described below may be activated. Release evidence must include at least one governed inbound native-agent path, one governed inbound external-integration path, and one governed outbound remote-agent path.
Phase 0: Contract And Threat Model
Deliver:
- A2A 1.0.1 tag
v1.0.1at commit3303592588e388e62e0f69f701af531d2f4e3991, A2A 0.3 compatibility tagv0.3.0at commit210f03d426e2f2fa92000e14ef0de3b7ba15aee5, and A2A TCK version 1.0.0 at commit5996b79f9cefa6fc390980e383e358a66fb9e49e; the implementation must not depend on an unversionedlatestspecification page; - canonical internal operation and error models;
- request/response and Agent Card size/depth limits;
- inbound and outbound route, task-ownership, signing, SSRF, confused-deputy, delegation, data-exfiltration, and disclosure threat model;
- task-artifact ownership, fine-grained operation, retention, legal-hold, deletion-evidence, external-reference, and memory-promotion contracts;
- private
light-a2a-backend/v1OpenAPI/JSON Schema authority, signed-context, loopback HTTP/JSON, SSE, restart-reconciliation, SDK, and conformance contracts; - purpose-separated native and external-facade signing-profile,
light-oauthsigning/JWKS, rotation, revocation, and audit contracts; - extension registry, exact-URI/versioning, optional-ignore, required-error, dependency, metadata-isolation, and runtime-handler contracts;
- profile-scoped extension configuration, single-generation profile validation, and 1.0/0.3 isolation rules for the runtime projection schema;
- compatibility matrix and explicit deferred features; and
- handler/config/module contracts.
The baseline record must cite the official
A2A specification repository
and changelog, then
freeze the exact revision used by generated models and conformance fixtures.
The machine-readable baseline, shared canonical projection fixture, and
existing-code inventory are maintained under contracts/a2a/phase0/.
Exit gates:
- all selected normative fixture shapes parse and canonicalize deterministically;
- the Portal Java compiler and Rust
a2a-coreimplementation produce the same digest for the shared Phase 0 golden projection; - malformed, oversized, ambiguous-version, invalid-activated-extension, and destination-escape fixtures fail closed, while an unknown optional 1.0 extension remains inactive and isolated;
- a missing published required extension in the 1.0 profile returns
ExtensionSupportRequiredError, the 0.3 profile rejects every extension declaration or activation during publication and projection compilation, and an unapproved required extension cannot enter an active publication; - extension configuration is expressible only per profile, a multi-generation profile is rejected, and a 1.0 profile’s extension set cannot reach an agent bound to a 0.3 profile in the same projection;
- card mutation tests prove no stale signature is preserved;
- no implementation phase starts with unresolved ownership or retention of durable tasks and task artifacts; and
- the public A2A binding and private sidecar backend protocol are represented as distinct contracts and cannot be enabled or versioned through each other’s configuration.
Phase 1: Shared A2A Foundation And Transparent Federation
Implementation evidence: scripts/run-a2a-phase1-gates.sh verifies the shared
Rust protocol, client, policy-envelope, task, artifact, native-runtime,
integration-runtime, and Gateway route contracts together with the Java Portal
projection compiler. The Gateway mints a short-lived body-bound authorization
context only for an authenticated principal and an exact Instance API route;
the selected native or integration runtime reclassifies the A2A operation and
enforces the signed binding, publication, policy digest, direction, operation,
audience, and tenant host. Remote federation uses a redirect-free bounded
client that resolves and pins public destinations before connecting.
Deliver:
- shared A2A protocol, runtime, client, policy, policy-envelope, and artifact lifecycle contracts;
- registered
apps/light-a2aservice onlight-axumandlight-runtime; - minimal gateway
a2a-routeredge module and implementation-kind service routing contract, including public-route-to-Instance-API resolution and exact generatedcardandinvokepolicy endpoints; - JSON-RPC 1.0 plus explicit 0.3 compatibility;
- well-known card proxying and safe URL rewriting;
- gateway edge authentication plus shared fine-grained A2A policy integration;
- shared 1.0 parsing and negotiation for
A2A-Extensions, with profile-scoped extension configuration, empty advertised, inbound, outbound, and required sets in every production profile, and compile-time rejection of extension configuration in the 0.3 profile; - immutable
light-a2aprojection loading, validation, last-known-good reload, replacement, and revocation; - streaming pass-through; and
- A2A telemetry.
Exit gates:
- unit and integration matrices cover both card generations, complete path matching, public host rules, JSON-RPC errors with HTTP 200, streaming, body limits, timeout, client disconnect, and reload;
- unsupported methods and versions never fall through to a generic proxy;
- unknown optional extension metadata cannot reach runtime policy, models, backends, artifacts, logs, or telemetry dimensions;
- two-node gateway and
light-a2atests prove no accidental affinity to one edge process and correct registered-service routing; - two agents with identical raw A2A specification endpoints retain distinct
routes and coarse authorization decisions without
endpointRulesoverwrite or permission leakage; and - soak testing shows bounded memory and stream cleanup.
Phase 2: Portal-Published Agent Cards
Implementation evidence: scripts/run-a2a-phase2-gates.sh verifies the
normalized Portal DB authoring schema, frozen CloudEvent creation/replay
inventories, deterministic Java publication compiler, Hybrid command/query
surfaces, Portal View production build, purpose-separated light-oauth
signing/JWKS contracts, strict Rust activation, and import-ready Config Server
metadata events. A publication remains PREPARED until its exact card is
signed, becomes STAGED when immutable audience properties are written, and
becomes ACTIVE only when those exact properties belong to the activated
Config Server snapshot. Runtime reload continues to use the existing snapshot
and control-plane lifecycle; Phase 2 does not add a mutable request-path query.
Deliver:
- reuse of the Agent registration publication foundation for native
LIGHT_AGENT; Phase 2 must not create a second Agent compiler, snapshot lifecycle, current pointer, reload protocol, or acknowledgement store; - Portal A2A publication authoring and validation;
- managed extension registry and dependency records, structured Portal View forms, and binding selectors that ship with no extension eligible for initial production activation;
- structured Portal View artifact-retention profiles, host defaults, agent overrides, fine-grained access-control policy linkage, effective preview, and immutable Config Server compilation;
- reusable public provider profiles and version-scoped Agent Public Metadata authoring, validation, source attribution, and effective-card preview;
- public disclosure projection; extended disclosure remains Phase 6 work;
- structured, tenant-scoped public skill aliases managed on the Skill form, frozen after first publication, and mapped immutably to internal skill UUID, version, and digest;
- safe AgentSkill mapping that never exposes Portal UUIDs, instructions, executable source, or internal tool/workflow placement;
- normalized Portal authoring tables, immutable versioned JSONB publication
aggregates, and separate compiled Gateway and selected
light-agentorlight-a2aConfig Server projections; - publication of bounded
agentPolicy.skillsand effective-catalog values or immutable catalog/package references into instance properties and activated Config Server snapshots; - active Gateway Instance API association, unique public path prefix, and
deployment-scoped
rule.endpointRulescompilation; - Portal View A2A Bindings table, structured forms, compiled preview, validation, explicit publication, revocation, and history diagnostics;
- Portal View host/environment Signing Profiles administration with inherited native and external-facade defaults and authorized overrides;
- purpose-separated A2A signing profiles and key lifecycle in
light-oauth, including authenticated Agent Card signing, public profile JWKS, rotation, revocation, and signing audit; - publication-time signed-card generation and verification, plus
light-agentandlight-a2aactivation-time signature validation, ETag, cache, expiry, and revocation; and - Portal UI/API visibility and publication diagnostics.
Exit gates:
- the generic Agent publisher can activate and reload a non-A2A Agent before
the optional native
a2aPolicyoverlay is enabled, and adding the overlay produces one new combined Agent snapshot generation rather than a separately activated A2A generation; - aggregate-version and publication-digest changes are deterministic;
- effective name, description, provider, documentation URL, icon, and semantic version follow the pinned precedence rules and retain source provenance;
- public skill aliases are normalized, unique within the host, stable across compatible revisions, reproducible from the publication, never expose a Portal UUID, and never become rebound to a different internal skill;
- inactive/unassigned/private skills never leak into public cards;
- runtime skill and catalog loading succeeds from the activated Config Server generation while Portal query is unavailable, and mutable authoring changes have no effect until a new publication is activated;
- artifact-retention defaults and agent overrides compile deterministically, load without Portal availability, and reference the existing fine-grained access-control policy without creating a parallel ACL;
- first-production cards and runtime projections contain empty extension sets in every profile, and arbitrary binding JSON cannot introduce or require an extension or move extension configuration outside its profile;
- native and external-facade cards verify against different host/environment signing profiles after final public URL generation;
- OAuth token keys, a wrong-environment profile, a wrong-purpose profile, a
caller-selected
kid, and an unauthorized signing workload all fail closed; - rotation proves new cards use the current
kid, cached and rollback cards remain verifiable only for the bounded overlap, and emergency revocation blocks affected cards until a newly signed publication is active; - no publication using a new
kidactivates before that key is available from the pinned profile JWKS, and JWKS failure rejects the new generation without breaking a still-valid last-known-good card; - revocation removes or denies the card within the documented propagation bound; and
- gateway,
light-agent, andlight-a2aoperate without request-path access to Portal authoring or projection tables.
Phase 3: External Business Agent Sidecar
Deliver:
EXTERNAL_SIDECARPortal implementation and binding model;- shared and sidecar deployment profiles for the same
light-a2abinary; - versioned
AgentBackendSDK and signedAuthorizedInvocationcontract; - canonical
contracts/a2a-backend/v1OpenAPI/JSON Schema contract, Rust reference adapter, golden vectors, and language-neutral TCK; - fixed-loopback HTTP/JSON backend transport with bounded unary operations and SSE for declared streaming backends;
- production Python, Java, and TypeScript/Node.js SDKs that expose only the business callbacks and own the transport/security plumbing;
- task, context, cancellation, status reconciliation, idempotency, and streaming adaptation;
- managed sidecar artifact validation, scanning, tenant-scoped storage, fine-grained access, expiry, legal hold, and verified deletion;
- sidecar Controller registration metadata; and
- developer templates containing business logic only.
Exit gates:
- the business backend receives neither raw caller tokens nor Portal policy;
- Python, Java, and TypeScript reference backends pass the same contract,
signed-context, unary, streaming, status, cancellation, artifact, deadline,
error, and restart TCK manifest and report the same compiled
light-a2abuild digest; - forged, expired, replayed, wrong-audience, wrong-agent, wrong-skill, wrong-operation, wrong-task, and wrong-context invocation attempts fail closed;
- the sidecar cannot call an unconfigured destination or cross tenant/network boundaries;
- sidecar/backend restarts reconcile detached work by signed task and backend operation identity without duplicating effects or guessing a terminal state;
- task-owner-only artifact access and explicit policy denials, expiry, legal hold, verified deletion, and deletion tombstones survive sidecar/backend restarts; non-owner sharing remains disabled until a separately versioned artifact grant API is implemented and qualified; and
- a reference external agent passes conformance, security, reload, audit, and telemetry gates without implementing platform plumbing.
Phase 4: Portal-Native light-agent Integration
Deliver:
- shared A2A server, card, policy, and task modules embedded in
light-agent; - direct Gateway routing to the registered
light-agentforLIGHT_AGENT; - native inbound
a2aPolicycompiled into the immutableagent.ymlprojection, with nolight-a2asidecar; - reuse of Config Server-projected
agentPolicy.skillsandagentPolicy.catalog.effectiveCatalog, with no live Portal catalog query as runtime authority; - durable context/session and task/turn mapping;
- message idempotency, task lookup, cancellation, and streaming event mapping;
- native task-artifact persistence through the shared lifecycle, independently governed from session history and Hindsight memory;
- authenticated task and artifact ownership and fine-grained operation checks, with normal access-control audit; and
- memory and effective-catalog integration through existing
light-agentboundaries.
Exit gates:
- gateway and
light-agentrestarts preserve native task lookup and ownership without requiring alight-a2aprocess; - duplicate message IDs do not create duplicate turns or effects;
- cross-principal and cross-agent context/task probes fail closed;
- cross-principal artifact probes, guessed download URLs, and implicit Portal administrator access fail closed unless the normal fine-grained policy explicitly grants the requested operation;
- cancellation and terminal-state races are deterministic;
- artifact expiry does not delete chat history or Hindsight memory, and memory retention does not keep an otherwise expired artifact retrievable;
- memory failure does not retry an accepted effectful action;
- a task pinned to one publication retains the same public-alias-to-skill digest mapping across a compatible skill update; and
- A2A and existing
/chatpaths produce equivalent governed agent behavior for an agreed test corpus.
Implementation evidence: scripts/run-a2a-phase4-gates.sh verifies the shared
strict A2A server models, native light-agent embedding, immutable Portal
projection fields, Portal compiler tests, import-ready Config Property events,
the full Rust workspace, and operational-bundle 1.6.0 parity across
portal-config-loc/all-in-lt, portal-config-dev, and
light-portal-install. scripts/run-a2a-phase4-operational-gates.sh runs the
native durability test against a disposable PostgreSQL store. That test proves
restart-safe task/context lookup, client-message idempotency, cross-principal
and cross-agent denial, idempotent cancellation, immutable public-skill mapping
evidence, independent artifact expiry, and survival of the underlying Agent
turn. Native A2A execution enters the same durable admission, fair dispatcher,
governed model/tool loop, effective catalog, memory, and terminal-turn paths as
/chat; it does not introduce a second Agent reasoning runtime or require a
light-a2a process.
Phase 5: Required Governed Outbound A2A
Implementation evidence: scripts/run-a2a-phase5-gates.sh verifies immutable
Portal trust review and changed-card quarantine, Instance-API-scoped outbound
Gateway routing, light-agent catalog tools, Workflow agentRef validation and
snapshot reload, server-owned destination and credential handling, bounded
delegation envelopes, incremental SSE governance, and managed-versus-ephemeral
artifact behavior. Portal View exposes the structured reviewed-card,
assignment, credential-reference, data-boundary, budget, and artifact controls;
the command service recomputes the canonical card digest and rejects destination,
review, signature, extension, assignment, or policy-scope inconsistencies. The
gate also verifies operational bundle parity across
portal-config-loc/all-in-lt, portal-config-dev, and
light-portal-install. scripts/run-a2a-phase5-operational-gates.sh runs the
outbound durability contract against a disposable PostgreSQL 17 store and
proves restart-safe local-to-remote task/context correlation, ownership denial,
idempotency, replay rejection, cancellation, artifact metadata, and audit
outbox persistence. Import-ready Config Property events publish the Agent and
Workflow outbound binding properties; they remain data files until explicitly
imported. A deployed multi-node canary remains release-qualification evidence,
not authority to weaken any of these code gates.
Deliver:
- external-agent onboarding and card refresh;
- end-to-end
light-agentoutbound invocation through the published Gateway route andlight-a2abinding; - Workflow
call: a2aagentRefcontract, validator, and migration errors for legacyagentCardorserverdestinations; - trust, signature, review, digest, and revocation lifecycle;
- remote-card extension review that permits core calls without activating unapproved optional extensions and blocks any unapproved required extension;
- policy-filtered external-agent discovery for Light agents;
- server-owned destination and credential resolution;
- managed import or explicitly ephemeral handling of remote task artifacts, without treating an upstream URI as a durable platform object;
- outbound operation, skill, task/context, delegation-depth, budget, loop-prevention, and data-boundary enforcement;
- bidirectional principal, calling-agent, target-agent, task, and policy-snapshot audit correlation; and
- changed-card review workflow.
Exit gates:
- arbitrary URL, redirect, DNS-rebinding, private-address, and credential substitution tests fail closed;
- a production-profile
light-agentcompletes an outbound message, stream, task lookup, and cancellation through a multi-node Gateway andlight-a2adeployment without accepting caller- or model-selected destination data; - an agent can call only assigned external agents and approved operations;
- delegation loops, excessive depth, expired budgets, cross-environment calls, and disallowed data classifications fail closed;
- changed or revoked cards cannot silently widen capabilities;
- a retained outbound artifact is available from managed tenant storage after the remote reference disappears, while a binding configured for ephemeral references makes no local-retention promise;
- a newly required extension, optional-to-required transition, changed extension URI/dependency, or revoked registry entry quarantines the remote binding; and
- outbound audit correlates human/workload principal, calling agent, target agent, task, and policy snapshot without exposing credentials.
Phase 6: Additional Bindings, Extensions, And Push
Add public A2A HTTP+JSON, public A2A gRPC, extended cards, push notifications, custom bindings, sidecar private-network mTLS or gRPC, additional backend SDK languages, or individual extensions only as independent profiles with conformance fixtures and operational qualification. Start with optional data-only extensions. A profile, method, state-machine, transport, SDK language, or required extension needs explicit threat model, compatibility, dependency, handler, error-mapping, downgrade, and rollback evidence before activation.
Push notification delivery additionally requires approved callback registration, callback ownership verification, SSRF controls, HMAC or mTLS, replay protection, retry budgets, dead-letter handling, and durable delivery state outside the gateway process.
Implemented Phase 6 Profile
The first Phase 6 profile is deliberately narrow. It is selectable only for an
EXTERNAL_SIDECAR binding using the A2A 1.0 JSON-RPC profile. Portal compiles
the selected extended-card profile, reviewed optional data extensions, and
push profile into the immutable light-a2a projection. It signs the public and
extended cards through the purpose-scoped light-oauth signing workflow.
An extension handler must exactly match the published URI, schema digest, and operation allowlist. Required extensions, dependencies, non-data behavior, and A2A 0.3 extension activation fail publication. Extended-card authorization is evaluated before conditional-cache handling, and its ETag includes the authorization-policy digest and revocation epoch.
Push configuration accepts only a Portal-approved registration ID and its
fixed credential-free HTTPS URL. The caller cannot supply a token, signing key,
or arbitrary callback. light-a2a validates task ownership, stores the
configuration and delivery outbox in a2a_ops, re-resolves the destination,
rejects redirects and non-global addresses, signs each attempt with a
server-owned HMAC key, and applies bounded retry, leases, and dead-letter state.
The frozen fixture and threat/rollback contract are under
contracts/a2a/phase6/. The public HTTP+JSON and gRPC bindings, custom public
bindings, private-backend mTLS/gRPC, required extensions, extra SDK languages,
and a native light-agent Phase 6 profile are still disabled. Each requires an
independent projection, persistence boundary, conformance suite, and rollback
gate; none may inherit activation from this external-sidecar profile.
Phase 7: Production Qualification And Rollback
Phase 7 adds no protocol method, transport, extension, or SDK. It turns the implemented profiles into a releasable unit by making deployment drift, readiness, multi-worker delivery, immutable evidence, canary, and rollback requirements executable.
The runtime exposes /_a2a/ready for traffic admission. It returns unavailable
when the active projection is expired. When push is enabled, it also returns
unavailable until the delivery worker has successfully polled the operational
store and whenever that success becomes older than the selected profile’s
maximum lease plus a small completion margin. /health remains liveness only.
Each worker claims one callback delivery at a time. The callback timeout must
leave at least five seconds in its lease, preventing a second replica from
normally reclaiming an in-flight request before the first can persist its
outcome. An expired lease is reclaimable by another worker; the previous owner
cannot complete or retry it. Retry exhaustion persists DEAD_LETTER rather
than dropping the event.
The deployment profiles mount both operational and artifact-store credentials,
persist the artifact root, use the worker-aware readiness endpoint, and load
runtime authority only from the immutable Config Server snapshot identified by
host, serviceId, and envTag. Checked-in local values contain bootstrap and
store bindings only; they contain no runtimePolicy, agent binding, instance
UUID, product fallback, or fabricated Agent Card.
The machine-readable evidence contract is under contracts/a2a/phase7/.
Automated qualification proves source/config parity, Compose rendering, bundle
checksums, fresh/upgrade schema parity, both traffic directions, ownership,
lease takeover, stale-worker readiness, durable dead-letter state, and restart
recovery. A production decision additionally requires immutable image and
snapshot digests, a 24-hour canary, alert review, security approval, and a
successful one-generation rollback exercise. CI must leave the evidence
template NOT_QUALIFIED; only evidence from the target environment may change
that decision.
Testing Strategy
Unit And Property Tests
- version and binding parsing;
- JSON-RPC envelope and A2A model validation;
- Agent Card canonicalization, mapping, URL rewriting, and signing;
- host/environment/purpose signing-profile selection, JWS/JWKS validation, key-overlap, and revocation behavior;
- public metadata precedence, SemVer, provider-profile, documentation-URL, and managed-icon validation;
- public skill-alias normalization, host uniqueness, first-publication freeze, immutable UUID/version/digest mapping, and no-rebinding validation;
- extension URI/version matching, direction and operation selection, dependency closure, required eligibility, optional-ignore behavior, required-error behavior, schema validation, and activated-response echoing;
- deterministic policy-endpoint generation and route-to-Instance-API binding validation;
- full-segment path matching and host/scheme construction;
- state and error mapping;
light-a2a-backend/v1OpenAPI and JSON Schema validation, canonical golden vectors, unknown-field behavior, and public-A2A/private-backend version isolation;- signed backend invocation validation for unary, streaming, status, and cancel, including exact task, context, idempotency, and backend-operation bindings;
- artifact-retention profile resolution and admission-time freezing;
- artifact operation authorization, digest and reference validation, managed import, expiry, legal-hold, deletion retry, verified absence, and tombstones;
- chat-reference and explicit memory-promotion lineage without payload duplication or implicit retention coupling;
- body, depth, collection, header, and event limits; and
- redaction and telemetry-cardinality guards.
Use property/fuzz tests for URI rewriting, JSON nesting, JSON-RPC IDs, extension lists, streaming event framing, and malformed cards.
Integration Tests
- A2A 1.0 client to external backend through
light-gateway; - A2A 0.3 compatibility route isolated from 1.0-only routes;
- public Agent Card access, plus authenticated extended-card access only when its independently authorized Phase 6 profile is enabled;
- JWT, endpoint authorization, delegation, and response filtering;
- two agent APIs with identical raw A2A endpoints, different public prefixes, and different roles proving isolated allow and deny decisions;
- service discovery, TLS, backend failures, and reload;
- Portal publication through
light-oauthsigning and JWKS verification for both native and external-facade profiles, including key rotation; - Portal skill assignment through instance-property staging, immutable snapshot
activation,
agent.ymlloading, and runtime acknowledgement, including Portal-query unavailability after activation; - empty first-production extension profiles, unknown optional extension isolation, activated-extension response negotiation, missing-required errors, and remote-card required-extension quarantine;
- activation of one runtime projection serving both a 1.0 profile and a 0.3 profile, proving profile-scoped extension isolation, rejection of 0.3 extension configuration, and rejection of a multi-generation profile;
- direct
light-gatewayto nativelight-agentmessage, stream, lookup, cancel, duplicate, and subscription/reconnection; - fixed-loopback
light-a2a-backend/v1unary, SSE streaming, status reconciliation, and cancellation, including sidecar and backend restarts; - Python, Java, and TypeScript/Node.js reference business backends passing the
same language-neutral TCK manifest and reporting one
light-a2abuild digest; - task-owner-only artifact metadata and lifecycle access, with Gateway policy able to narrow but never widen ownership; non-owner metadata, content, download, export, deletion, and memory-promotion grants require a future separately versioned grant API and are denied in this profile;
- Config Server artifact-policy activation, 24-hour chunk cleanup, 30-day task and content expiry, longer metadata retention, legal hold, and verified object-store deletion without a live Portal query;
- independent chat, artifact, and Hindsight retention, including explicit artifact-to-memory promotion with provenance and privacy-erasure lineage;
- governed
light-agentoutbound message, stream, lookup, and cancellation viaagentRef, including target revocation during an active task; - multi-gateway routing during a long-running durable task; and
- Workflow
call: a2aalias resolution plus rejection of rawagentCardandserverdestinations.
Security Tests
- cross-host/principal/agent/task access;
- guessed artifact IDs and download URLs, cross-artifact operation escalation, task-read-to-content-read escalation, implicit administrator access, expired artifact probes, and object-reference leakage;
- Agent Card prompt injection and secret scanning;
- Portal UUID, private instruction, executable source, and artifact credential leakage through public AgentSkill projections;
- extension-header count/size abuse, URI dereference/SSRF, untrusted parameter or metadata injection, dependency cycles, handler/schema substitution, downgrade, optional-to-required transition, and required-flag bypass;
- signed-card mutation and downgrade attempts;
- cross-host, cross-environment, cross-purpose, OAuth-key reuse, unauthorized signing-call, caller-selected-key, stale-key, and revoked-key attempts;
- arbitrary destination and DNS/redirect SSRF attempts;
- oversized/deep JSON and streaming resource exhaustion;
- replayed message/delegation/push requests;
- outbound confused-deputy, delegation-loop, cross-environment, model-selected-destination, and disallowed-data-boundary attempts;
- caller attempts to forge
instanceApiId,agentDefId, policy endpoint, or target service through headers, parameters, or request bodies; - forged or replayed
AuthorizedInvocation, raw-token leakage, sidecar bypass, unauthorized local backend access, wrong-task cancellation, wrong-operation status lookup, non-loopback or wildcard backend origins, redirects, proxy environment substitution, and SDK validation bypass in each supported language; - memory or skill content attempting to alter runtime authority; and
- mutable hybrid-query results, changed artifact references, digest mismatch, unsigned packages, and alias-rebinding attempts altering an active runtime.
Rollout And Compatibility
- Ship the A2A handler disabled until a validated profile and route exist.
- Enable one internal transparent-proxy canary before Portal-published cards.
- Enable and qualify a governed outbound canary after the inbound contracts are stable; do not declare production readiness until both directions pass their release gates.
- Keep A2A 0.3 and 1.0 as separate route profiles and metrics dimensions.
- Version the private
light-a2a-backend/v1contract independently from public A2A versions; a public-protocol upgrade must not silently change a business backend callback or SDK wire model. - Do not automatically upgrade an external registration to a new major version.
- Do not automatically activate a new extension, required flag, dependency, or extension version discovered during card refresh.
- Publish compatibility and deprecation windows through Portal.
- Retain a one-generation rollback target while it remains valid and unrevoked.
- Roll back by publication generation, not by editing live card JSON.
Resolved Design Decisions
- Portal uses normalized A2A binding authoring tables, an immutable versioned JSONB publication aggregate, and compiled audience-specific Config Server projections. Portal View manages bindings through a table and structured forms; raw compiled JSON is a read-only preview. CRUD changes Draft state, while an explicit publication workflow validates, snapshots, activates, and records runtime acknowledgement.
- A
LIGHT_AGENTbinding routes directly to native A2A modules embedded inlight-agent; it never deploys alight-a2asidecar. Thelight-a2asidecar is exclusively for external business agents that do not implement Light-Fabric platform concerns. Shared-servicelight-a2aremains the governed federation boundary for remote A2A servers. - Each logical agent remains an API/API-version asset, with
agentDefId == apiVersionId. Anagtproduct version describes deployable runtime compatibility and does not replace the logical agent identity. - Publishing an agent through a Gateway requires an active
instance_api_tassociation and unique public path prefix. Portal compiles opaque policy endpoint keys scoped byinstanceApiId; Gateway uses them for coarse card or invocation admission, while the selected runtime retains fine-grained A2A operation and skill authorization. - Agent Card public metadata uses deterministic Portal sources: name from the API, version-specific description with API-description fallback, semantic version from the API version, provider from an approved public provider profile, and documentation/icon from version-scoped agent public metadata. The immutable publication is runtime-authoritative; bindings, model-provider fields, runtime overrides, and unreviewed upstream cards are not.
- The first production milestone requires both governed inbound publication and governed outbound invocation. Inbound may land first for development and canary qualification, but Phase 5 catalog resolution, trust, credentials, delegation, data-boundary enforcement, loop/budget controls, and correlated audit are production gates. Phase 6 features remain disabled unless the independently qualified external-sidecar A2A 1.0 profile is selected.
light-oauthis the first-production Agent Card signing and JWKS authority. Native and external-facade publications use separate key rings scoped by host, environment, and purpose. The agent publication is the signed subject; a runtime fleet or instance is not the issuer. Portal manages structured signing profiles and lifecycle, Config Server projects only the signed card and logical profile metadata, and no OAuth token key, private key, or backing KMS/HSM reference is projected tolight-agent,light-a2a, or Gateway.- Public A2A skill IDs are stable, tenant-scoped publication aliases managed as structured Portal skill data; Portal UUIDs remain internal. Every immutable agent publication records the alias-to-UUID/version/digest mapping. Config Server delivers bounded skill instructions and capability descriptors as runtime authority, while executable packages live in signed, content-addressed artifact storage. Live Portal hybrid queries support authoring, validation, compilation, and diagnostics; any runtime search is bounded to an immutable publication and cannot hot-load mutable skills or executable authority.
- The first production profiles advertise and activate no A2A extensions. All
future activated extensions use exact, versioned URIs from a Portal-managed
allowlist; required extensions are initially prohibited and later require
explicit required eligibility, implementation, schema, dependency, security,
and conformance approval. Unknown optional requests remain inactive and
isolated. Extension configuration is profile-scoped rather than
instance-global: every profile is single-generation, a 1.0 profile’s sets
bind only the agents selecting it, and one runtime projection may serve both
generations without leaking extension policy between them. In the 1.0
profile, missing support for a published required extension returns
ExtensionSupportRequiredError; the 0.3 profile rejects every extension declaration or activation during publication and projection compilation. An unapproved required extension in a remote card prevents onboarding or activation. - A2A task artifacts have retention and visibility independent from chat
history and Hindsight memory.
TASK_OWNERis the default, while every additional metadata, content, download, export, deletion, or promotion to memory uses the existing fine-grained access-control policy; there is no implicit administrator authority or separate break-glass path. The selected runtime owns tenant-scoped managed storage and freezes the approved retention profile at task admission. Initial defaults are 24 hours for residual chunks, 30 days for managed content and external task visibility, and 365 days for non-content metadata and deletion evidence, with legal holds and approved compliance overrides. Config Server distributes only immutable rules. Chat stores bounded references, Hindsight ingestion is an explicitly authorized derived operation, and external URIs are imported or declared ephemeral rather than assumed durable. - The first external-developer release standardizes one private
light-a2a-backend/v1contract: HTTP/1.1 with JSON on a fixed loopback origin, plus SSE when the backend declares streaming. Python, Java, and TypeScript/Node.js are supported production SDKs; the checked-in Rust adapter is the reference implementation and conformance oracle. Status reconciliation is part of v1 so detached work can survive sidecar restarts. Unix-domain-socket HTTP is an optional non-blocking Linux hardening profile. Private-network mTLS, gRPC, WebSocket, stdio, FFI/plugins, Go, .NET, and other SDK languages are deferred and require independent qualification. The private backend contract is versioned independently and never activates a public A2A HTTP+JSON binding. - Runtime configuration identity is exactly
(host, serviceId, envTag)and current configuration is loaded only through/configs. PortalinstanceIdandinstanceApiIdvalues remain internal association, publication-target, and audit identifiers; no A2A runtime projection or Config Server query treats them as workload identity. Product ID and product version describe compatibility and never select runtime configuration. - Agent registration, native runtime linking, and A2A exposure are separate
lifecycle decisions. A native Agent uses one immutable Agent audience
snapshot containing the base policy and optional
a2aPolicyoverlay. The Gateway and any external-integrationlight-a2aruntime receive their own least-privilege snapshots under one coordinated release manifest. Current pointers may change in one control-plane transaction, but runtime application is completed only through explicit reload and per-target acknowledgement. - Phase 7 is a qualification boundary, not a capability bundle. Deployment
readiness combines immutable projection validity with push-worker/store
progress, callback timeout remains shorter than its fenced lease, and local
values cannot replace Config Server runtime authority. Automated evidence
is necessary but not sufficient: the target Host/environment must bind a
24-hour canary and rollback exercise to immutable image and publication
digests before its decision becomes
QUALIFIED.
Completion Criteria
The A2A Gateway feature is complete only when:
- selected A2A profiles pass protocol and negative conformance fixtures;
- Portal can publish and revoke a versioned Agent Card without direct gateway or runtime database access;
- every Gateway,
light-agent, andlight-a2atarget loads its current immutable snapshot from/configsusing only(host, serviceId, envTag), and mutable authoring or staged properties have no runtime effect before snapshot activation and explicit reload; - a native Agent can be registered, linked, and activated without A2A, while an
explicit later A2A publication adds the overlay through one new combined
Agent snapshot rather than parallel independently active Agent/A2A
generations or confused native and Gateway
instanceApiIdassociations; - public card content is demonstrably smaller and less privileged than the internal effective agent catalog;
- public skill IDs are stable aliases, never Portal UUIDs, and every active publication can reproduce their internal skill/version/digest mappings;
- every public metadata field is reproducibly compiled from its documented Portal source and cannot be silently replaced by runtime or upstream data;
- assigned skill instructions and capability descriptors load from an activated immutable Config Server generation without a live Portal authoring query;
- first-production Agent Cards and runtime projections advertise and activate no
extensions in any profile; the 1.0 negotiation layer correctly isolates
unknown optional metadata and returns
ExtensionSupportRequiredErrorfor missing published requirements, while the 0.3 profile rejects extension configuration before publication or activation; - extension policy is expressible only per profile, every profile is single-generation, and a projection serving both generations keeps each profile’s extension policy isolated;
- every later extension is reproducibly compiled from an active registry entry, exact versioned URI, approved direction/operations, dependency closure, handler/schema digests, metadata limits, and required-eligibility decision;
- executable skill packages are content-addressed, signed, scanned, verified, sandboxed, and never transported as Agent Card, Config Server, or live hybrid-query source content;
- every served Light-signed card verifies against the correct host,
environment, and native or external-facade
light-oauthprofile, while OAuth-token, wrong-purpose, wrong-environment, and revoked keys are rejected; - every request resolves an approved agent identity and backend without caller-selected destination data;
- first-production release evidence includes governed inbound native and external-integration calls plus governed outbound remote-agent calls, with end-to-end authorization and audit correlation;
- multiple published agents may share the same raw A2A specification endpoints without sharing, overwriting, or ambiguously matching Gateway authorization;
- durable Portal-native tasks survive gateway and
light-agentrestarts without a sidecar and can be looked up or canceled only by an authorized caller; - A2A artifacts use the existing fine-grained authorization system for every protected operation, honor the frozen retention profile and legal holds, delete managed content with verified evidence, and remain independently governed from chat history and Hindsight memory;
- an external developer can deploy a conformant business agent without implementing A2A, raw platform authentication, Portal policy lookup, Controller registration, audit, metrics, or tracing;
- the Python, Java, and TypeScript/Node.js SDKs pass identical unary, SSE,
status, cancellation, restart, signed-context, artifact, error, and negative
conformance cases against the same
light-a2a-backend/v1contract andlight-a2abuild; - shared-service and sidecar deployments enforce the same policy-decision and delegation contracts;
- skills remain descriptive discovery metadata and memory remains runtime-owned untrusted context;
- security, filtering, telemetry, reload, scale, and rollback gates pass in a deployed multi-node environment;
- Phase 7 evidence binds the Host/environment, image digests, active and prior publication generations, policy/content digests, bundle version, canary, alerts, lease takeover, dead-letter recovery, and rollback approval; and
- unsupported bindings and capabilities are documented and rejected rather than partially emulated.
Deploy Native
This page describes the recommended VM deployment model for the Rust
light-gateway native binary.
Use this model when a customer wants to run light-gateway as a microgateway on
a VM to protect backend MCP servers. The gateway starts from a small local
bootstrap config, downloads runtime config from config-server, then registers
itself with controller.
Recommended Model
Deliver a versioned install bundle, not an ad hoc runtime script.
The bundle should contain:
light-gatewaynative binary.- Minimal bootstrap config files.
- A
systemdunit. - An install script for filesystem setup.
- A root-owned environment file for secrets.
The install script can create users, directories, symlinks, permissions, and the
systemd unit. It should not be the long-running process wrapper, and it should
not pass secrets as command-line arguments.
Use systemd to run the service:
- It restarts the process on failure.
- It keeps logs in the host journal.
- It avoids shell-history and process-list leakage from command-line secrets.
- It gives the customer a standard operational surface:
start,stop,restart,status, andjournalctl.
Runtime Layout
light-gateway uses relative runtime paths:
configconfig-cache
The systemd service should therefore set WorkingDirectory to the installed
application directory.
Recommended VM layout:
/opt/light-gateway/
light-gateway
config -> /etc/light-gateway
config-cache -> /var/lib/light-gateway/config-cache
/etc/light-gateway/
startup.yml
server.yml
portal-registry.yml
client.yml
values.yml
ca.pem
light-gateway.env
/var/lib/light-gateway/
config-cache/
The local config directory contains only bootstrap-time files. Runtime config
downloaded from config-server is written to config-cache before Pingora starts.
Keep config-cache writable by the light-gateway service user.
Build Artifact
Build a release binary from light-fabric:
cargo build --release -p light-gateway
The artifact is:
target/release/light-gateway
Build on a compatible Linux distribution for the customer VM. If the customer
fleet has mixed Linux versions, prefer a static or target-compatible build so
the binary does not fail on an older glibc.
Package with a versioned filename:
light-gateway-<version>-linux-amd64.tar.gz
For customers with package-management standards, wrap the same layout in a
.deb or .rpm later. Start with tar.gz until the runtime contract is stable.
Bootstrap Config
The local bootstrap config only needs enough information to reach config-server, identify the gateway instance, and trust TLS.
Example values.yml:
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
light-config-server-uri: https://config-server.customer.example.com:8435
server.serviceId: com.customer.mcp-gateway-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: mcp-gateway-01.customer.example.com
server.httpPort: 8080
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.customer.example.com:8438
server.advertisedAddress must be a stable address that controller and clients
can use to reach the VM gateway. Do not advertise 127.0.0.1 or 0.0.0.0.
Example startup.yml:
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
Example server.yml:
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8080}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
tlsCertPath: ${server.tlsCertPath:}
tlsKeyPath: ${server.tlsKeyPath:}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
Example portal-registry.yml:
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
Example client.yml should include the customer CA path and hostname
verification policy for outbound HTTPS calls:
tls:
caCertPath: ${client.caCertPath:config/ca.pem}
verifyHostname: ${client.verifyHostname:true}
Keep the full gateway behavior, including MCP routing, authentication, rule configuration, and downstream MCP targets, in config-server. The VM should not need local edits for normal policy or route changes.
Secrets
Keep secrets in a root-owned environment file or in the customer’s secret manager. Do not pass secrets in command-line arguments.
Example /etc/light-gateway/light-gateway.env:
LIGHT_PORTAL_AUTHORIZATION=Bearer <token>
light_4j_config_password=<config-password-if-needed>
RUST_LOG=info
Permissions:
chown root:light-gateway /etc/light-gateway/light-gateway.env
chmod 0640 /etc/light-gateway/light-gateway.env
LIGHT_PORTAL_AUTHORIZATION is used for config-server bootstrap. The same token
is also used by portal registry startup when portal-registry.yml resolves
portalToken from light_portal_authorization.
Systemd Unit
Example /etc/systemd/system/light-gateway.service:
[Unit]
Description=Light Gateway
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=light-gateway
Group=light-gateway
WorkingDirectory=/opt/light-gateway
EnvironmentFile=/etc/light-gateway/light-gateway.env
ExecStart=/opt/light-gateway/light-gateway
Restart=on-failure
RestartSec=5
LimitNOFILE=65535
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=full
ProtectHome=true
ReadWritePaths=/var/lib/light-gateway/config-cache
[Install]
WantedBy=multi-user.target
Install and start:
systemctl daemon-reload
systemctl enable light-gateway
systemctl start light-gateway
systemctl status light-gateway
View logs:
journalctl -u light-gateway -f
Install Script Scope
An install script is useful, but keep it deterministic and small.
It should:
- Create the
light-gatewayuser and group. - Create
/opt/light-gateway,/etc/light-gateway, and/var/lib/light-gateway/config-cache. - Install the binary with executable permissions.
- Install bootstrap config files.
- Install or update the
systemdunit. - Set file ownership and permissions.
- Print the next operator steps for adding secrets and starting the service.
It should not:
- Embed bearer tokens.
- Pass tokens to
ExecStart. - Rewrite customer config-server state.
- Start the process before secrets and CA files are installed.
Startup Flow
The expected runtime flow is:
systemd
-> /opt/light-gateway/light-gateway
-> read local config/values.yml and startup.yml
-> call config-server with LIGHT_PORTAL_AUTHORIZATION
-> write downloaded config and files into config-cache
-> start Pingora with resolved runtime config
-> register gateway to controller using portalRegistry.portalUrl
-> route protected MCP traffic to downstream MCP servers
When startup.yml configures config-server, the runtime tries to download the
latest values.yml before starting. If that download fails for any reason, the
runtime continues startup with the available local and cached config, including
config-cache/values.yml when present.
Upgrade And Rollback
Use versioned binary releases:
/opt/light-gateway/releases/2.2.1/light-gateway
/opt/light-gateway/releases/2.2.2/light-gateway
/opt/light-gateway/light-gateway -> releases/2.2.2/light-gateway
Upgrade:
systemctl stop light-gateway
ln -sfn /opt/light-gateway/releases/2.2.2/light-gateway /opt/light-gateway/light-gateway
systemctl start light-gateway
Rollback:
systemctl stop light-gateway
ln -sfn /opt/light-gateway/releases/2.2.1/light-gateway /opt/light-gateway/light-gateway
systemctl start light-gateway
Do not delete config-cache during a normal binary rollback. It is the local
cache of the config-server-delivered runtime state.
Validation Checklist
Before handing the VM to the customer:
systemctl status light-gatewayis active.journalctl -u light-gatewayshows successful config-server bootstrap.journalctl -u light-gatewayshows successful controller registration.- The controller shows the gateway registered with the expected service id, environment, address, and port.
- The gateway health endpoint responds from the VM network.
- An MCP
tools/listcall reaches the gateway. - An MCP
tools/callcall reaches the configured backend MCP server. - Restarting the VM starts the gateway automatically.
Security Checklist
- Store bearer tokens and config passwords outside the install bundle.
- Use a customer CA file instead of disabling TLS verification in production.
- Use a stable DNS name for
server.advertisedAddress. - Restrict inbound VM firewall rules to required gateway ports.
- Restrict outbound VM firewall rules to config-server, controller, and backend MCP server addresses.
- Run as the dedicated
light-gatewayuser. - Keep
/etc/light-gateway/light-gateway.envreadable only by root and the service group. - Rotate
LIGHT_PORTAL_AUTHORIZATIONthrough the customer secret process.
Deploy Kubernetes
This page describes the recommended Kubernetes deployment model for the Rust
light-gateway image from light-fabric/apps/light-gateway.
Use this model when light-gateway runs as a microgateway in front of backend
MCP servers. The pod starts from local bootstrap config, downloads runtime
config from config-server into config-cache, starts Pingora, and registers the
gateway with controller.
Recommended Model
Deploy the gateway as a normal single-container Kubernetes workload:
Deploymentfor the gateway pod.Servicefor stable in-cluster access.ConfigMapfor bootstrap config and non-secret values.Secretfor bearer tokens and config passwords.emptyDirorPersistentVolumeClaimforconfig-cache.- Optional
Ingress,Gateway API,NodePort, orLoadBalancerfor external client access.
Keep gateway behavior such as MCP route definitions, access-control rules, backend MCP targets, and runtime TLS files in config-server. The Kubernetes bootstrap config should only contain enough information for startup, trust, and registration.
Image
Build the image from the workspace root:
./apps/light-gateway/build.sh 2.2.1
For local testing without pushing:
./apps/light-gateway/build.sh 2.2.1 --local
Use immutable tags in Kubernetes. Avoid latest for customer deployments.
The runtime image uses:
/app/light-gateway
/app/config -> /config
/app/config-cache
The process runs as the image user gateway. Mount /config for bootstrap
config and make /app/config-cache writable.
Runtime Paths
Recommended container layout:
/config/
startup.yml
server.yml
portal-registry.yml
client.yml
values.yml
ca.pem
/app/config-cache/
values.yml
downloaded certs and files
Use a read-only ConfigMap for /config. Use a writable volume for
/app/config-cache.
For most deployments, use emptyDir for config-cache. This gives each pod a
fresh cache and avoids accidentally keeping stale config across pod replacement.
Use a PersistentVolumeClaim only when the customer explicitly wants the
gateway to restart from the last downloaded config during a config-server
download outage. On each startup, the gateway tries to download the latest
values.yml before starting.
Registration Address
In Kubernetes, do not register the pod IP. Pod IPs are ephemeral.
If controller and callers are inside the same cluster, advertise the Service DNS name:
server.advertisedAddress: ai-microgateway.light-gateway
The pattern is:
<service-name>.<namespace>
The port is still registered separately from the host/address.
If controller or callers are outside the cluster, advertise the externally reachable DNS name instead, such as the Ingress or LoadBalancer hostname:
server.advertisedAddress: mcp-gateway.customer.example.com
For the Rust gateway, this is configured with server.advertisedAddress. The
Java gateway template uses STATUS_HOST_IP; that is a light-4j-specific hook
and is not the Rust gateway contract.
Bootstrap Config
Example values.yml for an in-cluster controller and config-server:
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.customer.mcp-gateway-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: ai-microgateway.light-gateway
server.httpPort: 8080
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
client.caCertPath: config/ca.pem
client.verifyHostname: true
Example startup.yml:
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
Example server.yml:
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8080}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
tlsCertPath: ${server.tlsCertPath:}
tlsKeyPath: ${server.tlsKeyPath:}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
Example portal-registry.yml:
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
Example client.yml:
tls:
caCertPath: ${client.caCertPath:config/ca.pem}
verifyHostname: ${client.verifyHostname:true}
Use the customer CA in ca.pem. Do not disable hostname verification in
production to work around certificate SAN problems.
Secrets
Store the portal bearer token and optional config password in a Kubernetes
Secret.
Example:
apiVersion: v1
kind: Secret
metadata:
name: light-gateway-secret
namespace: light-gateway
type: Opaque
stringData:
LIGHT_PORTAL_AUTHORIZATION: "Bearer <token>"
light_4j_config_password: "<config-password-if-needed>"
LIGHT_PORTAL_AUTHORIZATION is used for config-server bootstrap. It is also
used by portal registry startup when portal-registry.yml resolves
portalToken from light_portal_authorization.
Do not store real bearer tokens in Git, ConfigMaps, Helm values committed to the repo, or rendered deployment examples.
Example Manifests
Create the namespace separately:
kubectl create namespace light-gateway
If deploying through light-deployer, keep Namespace out of the rendered
bundle because deployer policy may block cluster-scoped resources.
Example bootstrap ConfigMap:
apiVersion: v1
kind: ConfigMap
metadata:
name: light-gateway-bootstrap
namespace: light-gateway
data:
values.yml: |
startup.host: customer.example.com
startup.timeout: 3000
startup.connectTimeout: 3000
startup.bootstrapCaCertPath: config/ca.pem
light-config-server-uri: https://config-server.lightapi.svc.cluster.local:8435
server.serviceId: com.customer.mcp-gateway-1.0.0
server.environment: prod
server.ip: 0.0.0.0
server.advertisedAddress: ai-microgateway.light-gateway
server.httpPort: 8080
server.enableHttp: true
server.httpsPort: 8443
server.enableHttps: false
server.enableRegistry: true
server.startOnRegistryFailure: true
portalRegistry.portalUrl: https://controller.lightapi.svc.cluster.local:8438
client.caCertPath: config/ca.pem
client.verifyHostname: true
startup.yml: |
host: ${startup.host:dev.lightapi.net}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
envTag: ${server.environment:dev}
acceptHeader: application/yaml
timeout: ${startup.timeout:3000}
connectTimeout: ${startup.connectTimeout:3000}
configServerUri: ${light-config-server-uri:https://local.localhost}
authorization: ${light_portal_authorization:}
bootstrapCaCertPath: ${startup.bootstrapCaCertPath:config/ca.pem}
server.yml: |
ip: ${server.ip:0.0.0.0}
advertisedAddress: ${server.advertisedAddress:127.0.0.1}
httpPort: ${server.httpPort:8080}
enableHttp: ${server.enableHttp:true}
httpsPort: ${server.httpsPort:8443}
enableHttps: ${server.enableHttps:false}
tlsCertPath: ${server.tlsCertPath:}
tlsKeyPath: ${server.tlsKeyPath:}
serviceId: ${server.serviceId:com.networknt.light-gateway-1.0.0}
enableRegistry: ${server.enableRegistry:true}
startOnRegistryFailure: ${server.startOnRegistryFailure:true}
dynamicPort: ${server.dynamicPort:false}
environment: ${server.environment:dev}
shutdownGracefulPeriod: ${server.shutdownGracefulPeriod:2000}
portal-registry.yml: |
portalUrl: ${portalRegistry.portalUrl:https://localhost:8438}
portalToken: ${light_portal_authorization:}
controllerDiscoveryToken: ${portalRegistry.controllerDiscoveryToken:}
client.yml: |
tls:
caCertPath: ${client.caCertPath:config/ca.pem}
verifyHostname: ${client.verifyHostname:true}
ca.pem: |
-----BEGIN CERTIFICATE-----
<customer-ca-certificate>
-----END CERTIFICATE-----
Example Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: ai-microgateway
namespace: light-gateway
labels:
app: ai-microgateway
spec:
replicas: 2
selector:
matchLabels:
app: ai-microgateway
template:
metadata:
labels:
app: ai-microgateway
spec:
securityContext:
fsGroup: 999
fsGroupChangePolicy: OnRootMismatch
containers:
- name: light-gateway
image: networknt/light-gateway:2.2.1
imagePullPolicy: IfNotPresent
env:
- name: LIGHT_PORTAL_AUTHORIZATION
valueFrom:
secretKeyRef:
name: light-gateway-secret
key: LIGHT_PORTAL_AUTHORIZATION
- name: light_4j_config_password
valueFrom:
secretKeyRef:
name: light-gateway-secret
key: light_4j_config_password
optional: true
- name: RUST_LOG
value: info
ports:
- name: http
containerPort: 8080
- name: https
containerPort: 8443
readinessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /health
port: http
initialDelaySeconds: 30
periodSeconds: 30
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: "1"
memory: 512Mi
volumeMounts:
- name: bootstrap-config
mountPath: /config
readOnly: true
- name: config-cache
mountPath: /app/config-cache
volumes:
- name: bootstrap-config
configMap:
name: light-gateway-bootstrap
- name: config-cache
emptyDir: {}
The example uses fsGroup: 999, which matches the default gateway group in the
current image. Adjust it if the image user or group changes.
If HTTP is disabled and only HTTPS is enabled, change the probes to an HTTPS probe or a TCP probe.
Example Service:
apiVersion: v1
kind: Service
metadata:
name: ai-microgateway
namespace: light-gateway
spec:
type: ClusterIP
selector:
app: ai-microgateway
ports:
- name: http
port: 8080
targetPort: http
- name: https
port: 8443
targetPort: https
For external access, add an Ingress, Gateway API route, NodePort, or
LoadBalancer according to the customer cluster standard. If external clients
or controller use that external path, set server.advertisedAddress to the same
externally reachable DNS name.
Apply With Kubectl
Apply manifests in this order:
kubectl apply -f namespace.yml
kubectl apply -f secret.yml
kubectl apply -f configmap.yml
kubectl apply -f deployment.yml
kubectl apply -f service.yml
Check rollout:
kubectl -n light-gateway rollout status deploy/ai-microgateway
kubectl -n light-gateway get pods -l app=ai-microgateway
kubectl -n light-gateway logs deploy/ai-microgateway
For local testing with a ClusterIP Service:
kubectl -n light-gateway port-forward svc/ai-microgateway 8080:8080 8443:8443
Deploy Through Light-Deployer
When light-deployer runs outside the cluster and has
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR set, repoUrl: "local" can point to local
templates.
When light-deployer runs inside Kubernetes, use a real Git URL:
{
"template": {
"repoUrl": "https://github.com/networknt/light-fabric.git",
"ref": "main",
"path": "apps/light-gateway/k8s/light-gateway"
}
}
Do not use repoUrl: "local" for an in-cluster deployer unless the template
repo is mounted into the deployer container and
LIGHT_DEPLOYER_TEMPLATE_BASE_DIR points to it.
The in-cluster deployer checks out repoUrl at ref and reads manifests from
template.path.
Keep Namespace out of templates rendered by light-deployer if the deployer
policy blocks cluster-scoped resources. Create the namespace separately:
kubectl create namespace light-gateway
Config-Server Requirements
Before deploying the gateway pod, config-server should already have config for the tuple used by startup:
host = startup.host
serviceId = server.serviceId
envTag = server.environment
At minimum, config-server should return runtime config for:
handler.ymlmcp-router.ymlaccess-control.ymlandrule.ymlwhen MCP authorization is enabled.security.yml,unified-security.yml, or other active auth config.websocket-router.ymlwhen WebSocket MCP/BFF routing is enabled.- Any downstream client, token, or registry config required by the selected handlers.
The pod bootstrap files should stay small and stable. Normal route, policy, and backend changes should go through config-server and controller reload flows.
Startup Flow
Expected runtime flow:
Kubernetes starts pod
-> /app/light-gateway
-> read /app/config -> /config bootstrap files
-> call config-server with LIGHT_PORTAL_AUTHORIZATION
-> write downloaded config and files into /app/config-cache
-> start Pingora with resolved runtime config
-> register gateway to controller using portalRegistry.portalUrl
-> advertise server.advertisedAddress and configured port
-> route protected MCP traffic to backend MCP servers
When startup.yml configures config-server, the runtime tries to download the
latest values.yml before starting. If that download fails for any reason, the
runtime continues startup with the available local and cached config, including
/app/config-cache/values.yml when present.
Upgrade And Rollback
Use Kubernetes rolling updates with immutable image tags:
kubectl -n light-gateway set image deploy/ai-microgateway \
light-gateway=networknt/light-gateway:2.2.2
kubectl -n light-gateway rollout status deploy/ai-microgateway
Rollback:
kubectl -n light-gateway rollout undo deploy/ai-microgateway
For production, prefer changing only one variable at a time: either image tag or config-server runtime config, not both in the same rollout.
Validation Checklist
After deployment:
kubectl -n light-gateway rollout status deploy/ai-microgatewaysucceeds.- Pods are ready and restart count is stable.
- Logs show successful config-server bootstrap.
- Logs show successful controller registration.
- Controller shows the gateway registered with the expected service id, environment, host, and port.
server.advertisedAddressis reachable from the controller.- The Service responds on
/health. - MCP
tools/listreaches the gateway. - MCP
tools/callreaches the backend MCP server. - A pod restart still starts cleanly with the selected cache policy.
Security Checklist
- Keep bearer tokens in Kubernetes
Secret, notConfigMap. - Use customer CA trust and keep
client.verifyHostname: truein production. - Use immutable image tags and image pull credentials from Kubernetes secrets when the registry is private.
- Run as the non-root image user.
- Make
/configread-only. - Make only
/app/config-cachewritable. - Restrict ingress traffic to required gateway ports.
- Restrict egress traffic to config-server, controller, token/key services, and backend MCP servers.
- Rotate
LIGHT_PORTAL_AUTHORIZATIONthrough the customer secret process.
Kubernetes Gateway API Design
Status
Proposal.
This page captures how the current light-gateway work can be reused for
Kubernetes Gateway API without turning the microgateway product into a
catch-all Kubernetes control plane. The recommended direction is a separate
light-k8s-gateway product built on light-pingora for north/south ingress,
with a later sidecar or mesh product for transparent east/west traffic.
Context
The current Kubernetes deployment model runs light-gateway as a normal
Deployment with a ClusterIP Service. Runtime behavior comes from local
bootstrap config, config-server downloaded files in config-cache, and the
Pingora data plane built by light-pingora.
The current gateway already has useful data-plane pieces:
- HTTP and HTTPS proxying through Pingora.
- Static upstreams from
proxy.yml. - Service-aware routing from
router.yml. - Direct registry, controller-backed discovery, and static service targets.
- Handler chains for security, header mutation, CORS, rate limits, token handling, MCP, WebSocket, static resources, and config reload.
- Live config managers and reloaders for route and handler modules.
Gateway API adds a Kubernetes-native control plane. For ingress, users create
GatewayClass, Gateway, and route resources such as HTTPRoute. For service
mesh, the GAMMA model attaches route resources directly to Kubernetes
Service objects instead of using Gateway and GatewayClass.
Product Boundary
Keep the product line split by operational role:
light-pingorais the shared data-plane framework.light-gatewayremains the microgateway, sidecar, BFF, API, agent, MCP, and LLM gateway product configured through Light runtime, config-server, controller-rs, and local config.light-k8s-gatewayis the proposed Kubernetes Gateway API product for north/south ingress. It should reuselight-pingoraand lift reusablelight-gatewaymodules where appropriate, but it should own Kubernetes watches, Gateway API status, RBAC, listener translation, TLS Secret handling, and EndpointSlice routing.light-k8s-gateway-controllerandlight-k8s-gateway-proxyshould be separate deployments from the first implementation. The controller owns Kubernetes RBAC and status writes. The proxy owns untrusted client traffic and should not need Kubernetes API permissions.- A future
light-meshorlight-sidecarproduct should own transparent east/west Service Mesh behavior if we pursue GAMMA conformance. It should share the Gateway API route compiler andlight-pingoradata-plane modules, but its deployment model is sidecar or node-local interception, not ingress.
This avoids giving ordinary microgateway deployments broad Kubernetes RBAC and keeps config-server/controller-rs routing separate from portable Gateway API routing intent.
Goals
- Let operators install
light-k8s-gatewayas a Gateway API implementation with a controller name such asnetworknt.com/light-k8s-gateway. - Support north/south ingress with
GatewayClass,Gateway,HTTPRoute, KubernetesService,EndpointSlice,Secret, andReferenceGrant. - Separate Kubernetes reconciliation from request proxying so control-plane RBAC is never granted to the public traffic data plane.
- Provide a migration path from NGINX or Traefik by running side by side with a distinct GatewayClass, then moving routes class by class or host by host.
- Reuse the existing Pingora proxy, handler chain, service discovery, metrics, and config reload model instead of creating a separate proxy stack.
- Use Gateway API policy attachment for Light-specific Kubernetes policy CRDs instead of annotations or out-of-band route policy.
- Support east/west traffic using Gateway API mesh semantics where
HTTPRoute.parentRefscan point at aService. - Keep Light-specific policies available without forcing them into portable Gateway API fields. Gateway API should configure routing; Light config and future policy CRDs should configure Light-specific behavior.
- Build toward Gateway API conformance tests for both Gateway and Mesh feature sets.
Non-Goals
- Do not remove existing config-server, direct registry, portal registry, or static route support.
- Do not require every
light-gatewaydeployment to watch Kubernetes. Gateway API support should be disabled unless explicitly configured. - Do not run the Kubernetes controller reconciler inside public data-plane pods with broad Kubernetes RBAC.
- Do not claim immediate support for every Gateway API route type. Start with
HTTPRoute; addGRPCRoute,TLSRoute,TCPRoute, andUDPRoutein later milestones. - Do not make transparent east/west interception a hidden side effect of the ingress deployment. Mesh mode needs an explicit data-plane deployment model.
- Do not treat a non-transparent egress gateway as fully GAMMA-compliant mesh support.
Target API Versions
The north/south MVP targets the Gateway API v1
Standard Channel
resources:
GatewayClassGatewayHTTPRouteReferenceGrant
Experimental or later milestones must be labeled explicitly in docs, manifests,
and conformance reports. This includes GAMMA mesh behavior and route kinds such
as GRPCRoute, TLSRoute, TCPRoute, and UDPRoute when those features rely
on non-Standard channels in the installed Gateway API version.
North/South Ingress Model
For ingress replacement, light-k8s-gateway should run as two cooperating
pieces:
light-k8s-gateway-controller: watches Kubernetes resources, validates attachment and policy, updates status, performs leader election, and produces a compiled routing snapshot.light-k8s-gateway-proxy: consumes signed or mTLS-protected snapshots and serves client traffic through Pingora. It has no Kubernetes watch or status permissions and can scale independently with an HPA.
The split is mandatory from day 1. It prevents a proxy vulnerability in the
public data plane from becoming a Kubernetes control-plane compromise. The
controller can run as an HA deployment with Kubernetes Lease leader election;
only the leader reconciles resources and writes status. Non-leader controller
replicas stay warm and can take over quickly.
Snapshot delivery can start as a lightweight internal gRPC stream and evolve
toward an xDS-like API if we need richer incremental updates. The proxy should
apply the received GatewayApiSnapshot through the same kind of ConfigManager
swap used by the current Pingora modules.
Typical installation:
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
name: light-k8s-gateway
spec:
controllerName: networknt.com/light-k8s-gateway
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: public
namespace: gateway-system
spec:
gatewayClassName: light-k8s-gateway
listeners:
- name: http
protocol: HTTP
port: 80
allowedRoutes:
namespaces:
from: All
- name: https
protocol: HTTPS
port: 443
hostname: api.example.com
tls:
mode: Terminate
certificateRefs:
- kind: Secret
name: api-example-com
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: petstore
namespace: apps
spec:
parentRefs:
- name: public
namespace: gateway-system
sectionName: https
hostnames:
- api.example.com
rules:
- matches:
- path:
type: PathPrefix
value: /pets
backendRefs:
- name: petstore
port: 8080
The controller resolves this into a runtime route table:
Gateway listener
-> accepted HTTPRoutes
-> host/path/header/method/query matches
-> filters supported by light-k8s-gateway
-> backend Service
-> EndpointSlice addresses
-> Pingora ProxyTarget set
The existing proxy.yml and router.yml paths remain useful for legacy and
non-Kubernetes deployments. Kubernetes Gateway API routes should not depend on
service_id headers or pathPrefixService.yml; they should route from the
compiled Gateway API table directly to Kubernetes endpoints.
Required Ingress Patches
Add a Kubernetes Gateway API module:
k8sGatewayApi:
enabled: ${k8sGatewayApi.enabled:false}
mode: ${k8sGatewayApi.mode:ingress}
controllerName: ${k8sGatewayApi.controllerName:networknt.com/light-k8s-gateway}
gatewayClassName: ${k8sGatewayApi.gatewayClassName:light-k8s-gateway}
watchNamespaces: ${k8sGatewayApi.watchNamespaces:[]}
statusAddress: ${k8sGatewayApi.statusAddress:}
Implementation changes:
- Create
apps/light-k8s-gateway-controllerandapps/light-k8s-gateway-proxy. - Add Gateway API and Kubernetes clients, likely behind a Cargo feature such
as
k8s-gateway-api, usingkube,kube-runtime,k8s-openapi, and generated Gateway API resource types. - Watch
GatewayClass,Gateway,HTTPRoute,ReferenceGrant,Service,EndpointSlice,Secret, andNamespace. - Compile watched objects into a deterministic
GatewayApiSnapshot. - Push the compiled snapshot to proxy pods over an authenticated internal channel.
- Store the received snapshot in a proxy-side
ConfigManager, similar to the current proxy and router reload model. - Add a
light-pingoraGateway API route-table module that can select a backend before falling back to existing proxy/router behavior. - Update Kubernetes status conditions for
GatewayClass,Gateway, listeners, and routes. Status must clearly report unsupported route types, listener conflicts, missing TLS secrets, rejected cross-namespace references, empty backends, and unsupported filters. - Add Kubernetes
Leaseleader election so only one controller replica writes status and publishes snapshots. - Add controller RBAC for read watches, Secret reads where allowed, Lease writes, and status updates. Secret read permissions should be namespace-scoped where possible.
- Give proxy pods no Kubernetes RBAC by default.
- Add install manifests for separate controller and proxy
ServiceAccount,ClusterRole,ClusterRoleBinding,Deployment,Service, and a sampleGatewayClass.
The transport also needs a listener model. Today PingoraTransport binds the
single server.httpPort and single server.httpsPort from server.yml. That
is enough for the first 80/443 ingress path, but full Gateway API support
needs multiple listeners with independent protocol, port, hostname, and TLS
settings.
Suggested runtime patch:
server:
listeners:
- name: http
protocol: HTTP
ip: 0.0.0.0
port: 80
- name: https-api
protocol: HTTPS
ip: 0.0.0.0
port: 443
hostname: api.example.com
tlsCertPath: /var/run/light-k8s-gateway/tls/api/tls.crt
tlsKeyPath: /var/run/light-k8s-gateway/tls/api/tls.key
Keep server.httpPort, server.enableHttp, server.httpsPort, and
server.enableHttps as backward-compatible shorthand.
HTTPRoute Support Plan
Start with the common ingress subset:
GatewayClassacceptance fornetworknt.com/light-k8s-gateway.Gatewaylisteners forHTTPand terminatedHTTPS.HTTPRouteattachment byparentRefs,sectionName, listener hostname, listener namespace policy, and route hostname.HTTPRoutematches for path prefix, exact path, method, headers, and query parameters.backendRefsto KubernetesServicebackends, including weights.ReferenceGrantfor cross-namespace backend references.- Endpoint resolution from
EndpointSlice, with Service DNS as a fallback only when endpoint watching is unavailable. - TLS Secret loading for terminated HTTPS.
- Request header modification and URL rewrite where existing Pingora handlers already provide equivalent behavior.
Later milestones:
- Request redirect, response header modification, request mirroring, retries, and timeouts.
GRPCRouteover HTTP/2.TLSRoutefor SNI routing and passthrough.TCPRouteandUDPRoutefor L4 ingress if Pingora transport support is added.- Backend TLS policy and mTLS to upstream services.
Light Policy Attachment
Kubernetes-native deployments should use the Gateway API Policy Attachment pattern from GEP-713 for Light-specific behavior. Do not use annotations for core behavior, and do not require config-server-owned route policy for the Kubernetes Gateway API path.
Add Light policy CRDs with targetRefs that point at Gateway API resources:
apiVersion: gateway.lightapi.net/v1alpha1
kind: LightAuthPolicy
metadata:
name: petstore-auth
namespace: apps
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: petstore
jwt:
issuer: https://issuer.example.com
audience:
- petstore
apiVersion: gateway.lightapi.net/v1alpha1
kind: LightRateLimitPolicy
metadata:
name: petstore-ratelimit
namespace: apps
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: petstore
limits:
- name: default
requests: 1000
window: 60s
The controller should resolve effective policy for supported target kinds such
as Gateway, listener section, HTTPRoute, route rule, and eventually
Service for mesh. Policy status should report Accepted, Programmed, and
conflict conditions so resource owners can tell whether a policy is active.
Config-server remains valid for non-Kubernetes light-gateway deployments and
for migration bridges. For light-k8s-gateway, Kubernetes resources should be
the source of routing and policy intent.
TLS Secret Handling
TLS Secret material must not be written to persistent disk or normal
config-cache.
Preferred handling:
- The controller reads referenced TLS
Secretobjects, validates references andReferenceGrantrequirements, and distributes certificate material to proxies through the authenticated snapshot channel. - Proxies hold certificate material in memory and update Pingora TLS state without persisting private keys.
- If Pingora integration requires file paths for an early milestone, write
temporary files only to an
emptyDirmounted withmedium: Memory, under a path such as/var/run/light-k8s-gateway/tls.
Never copy TLS private keys into config-server, config-cache, persistent
volumes, image layers, or logs.
Endpoint Abstraction
light-pingora should not need to know whether endpoints came from Kubernetes,
direct-registry.yml, controller-rs discovery, or a static config file. Add a
shared endpoint abstraction such as:
UpstreamCluster
name
protocol
tls settings
load-balancing policy
EndpointSet
endpoint address
port
health/ready state
metadata
light-k8s-gateway-controller translates Service and EndpointSlice objects
into this shape. Existing Light discovery paths can translate direct registry
and portal-registry results into the same shape. The Pingora route-table module
then selects an UpstreamCluster without carrying Kubernetes-specific logic.
East/West Mesh Model
Gateway API mesh support uses a different binding model. Routes attach directly
to Service resources:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: petstore-policy
namespace: apps
spec:
parentRefs:
- group: core
kind: Service
name: petstore
port: 8080
rules:
- matches:
- path:
type: PathPrefix
value: /v1
backendRefs:
- name: petstore-v2
port: 8080
weight: 10
- name: petstore-v1
port: 8080
weight: 90
The runtime semantics are:
- If no route attaches to a Service, default mesh behavior forwards to the Service backend.
- If routes attach and the request matches at least one route, the selected route backendRefs determine the destination.
- If routes attach and no route matches, reject the request.
- Same-namespace routes are producer routes and affect all clients.
- Different-namespace routes are consumer routes and affect clients in the route namespace.
The current light-gateway can proxy service-to-service calls explicitly, but
it does not transparently intercept traffic to Kubernetes Service frontends.
That means a real mesh implementation needs a data-plane attachment model, not
only a route compiler.
Recommended mesh milestones:
- Mesh M0: compile Service-attached
HTTPRouteresources and expose the effective route table through logs, module registry, and status. This proves the control-plane model without traffic interception. - Mesh M1: support an explicit in-cluster egress gateway mode. Workloads call
light-gatewaydirectly or through a configured HTTP proxy. This is useful operationally, but not advertised as transparent GAMMA conformance. - Mesh M2: add sidecar mode. Inject a lightweight
light-gatewaysidecar, or preferably a smallerlight-sidecarorlight-meshbinary using the samelight-pingoraroute-table module. Redirect outbound HTTP traffic to the sidecar, identify the original Service destination, apply Service-attached routes, then proxy to selected endpoints. - Mesh M3: add node-local or ambient mode. Use a DaemonSet plus CNI or eBPF redirection to intercept Service traffic without per-pod sidecars. This has a larger operational surface and should follow sidecar validation.
Sidecar mode is the shortest path because the current light-gateway already
has sidecar concepts such as sidecar.egressIngressIndicator, token handling
for outbound calls, and service discovery. The production packaging should
still be a dedicated sidecar or mesh product if the target is transparent
east/west traffic. The missing pieces are transparent redirect, original
destination detection, and a Service-oriented route table.
Mesh Data-Plane Requirements
To proxy east/west traffic with GAMMA semantics, add:
- A mesh route compiler that watches
HTTPRoute,Service,EndpointSlice,ReferenceGrant, and namespaces. - A Service frontend index keyed by namespace, Service name, port, DNS name, ClusterIP, and possibly original destination socket address.
- Producer and consumer route merge logic that follows Gateway API mesh rules.
- Request matching and rejection behavior for Services with attached routes.
- Backend endpoint selection from the selected route’s backendRefs.
- A sidecar or node-local interception mechanism that can recover the original destination Service before the request is proxied.
- Policy hooks for Light security, token, and observability handlers.
- Mesh conformance test wiring with
--supported-features=Mesh.
Do not map GAMMA Service routes to Gateway listeners. In mesh mode, the
Service is the parent object, and GatewayClass/Gateway are intentionally not
part of the route binding.
Coexistence With Existing Light Runtime
Keep these layers distinct:
- Gateway API resources express portable Kubernetes routing intent.
light-pingoraroute tables execute the selected routing intent.handler.ymland Light module config apply Light-specific behavior.light-gatewaycontinues to serve the current microgateway, sidecar, BFF, API, agent, MCP, and LLM provider use cases.light-k8s-gatewayowns Kubernetes Gateway API ingress behavior.portal-registryanddirect-registry.ymlremain available for non-Kubernetes targets and existing Light service discovery.- Config-server remains the source for non-Kubernetes
light-gatewaypolicy and migration bridges. Kubernetes-nativelight-k8s-gatewayrouting and policy intent should come from Gateway API resources and Light policy CRDs.
For ingress, Kubernetes Service and EndpointSlice should be the primary
backend source. For non-Kubernetes or hybrid targets, add an explicit
implementation-specific backend policy instead of overloading portable
backendRefs.
Status And Conformance
Gateway API users rely on status. The controller must update:
GatewayClass.status.conditions.Gateway.status.addresses, listener conditions, and supported features.HTTPRoute.status.parentsfor every parentRef.- Light policy CRD status, including
Accepted,Programmed, and conflict conditions.
Only the active leader should update Kubernetes status. Controller replicas use
Kubernetes Lease leader election to avoid API-server write races and status
flapping.
Minimum conformance gates:
go test ./conformance -run TestConformance -args \
--gateway-class=light-k8s-gateway \
--supported-features=Gateway,HTTPRoute
Mesh conformance gate:
go test ./conformance -run TestConformance -args \
--supported-features=Mesh
When ingress and mesh are both enabled:
go test ./conformance -run TestConformance -args \
--gateway-class=light-k8s-gateway \
--supported-features=Mesh,Gateway,HTTPRoute
Observability And Telemetry
light-k8s-gateway must be operable as a primary ingress controller. Provide
Prometheus metrics, OpenTelemetry traces, and structured logs from day 1.
Proxy metrics:
- Request count tagged by Gateway, listener, route namespace,
HTTPRoute, backend Service, status code, and status class. - Request duration and upstream duration histograms.
- Active connections and in-flight requests.
- Upstream connection errors, retries, timeouts, and circuit-breaker opens.
- Snapshot version, snapshot age, and snapshot apply errors.
Controller metrics:
- Reconcile count, duration, and error count by resource kind.
- Kubernetes watch reconnect count and API-server request errors.
- Status update count and conflict count.
- Leader-election state.
- Snapshot generation count, size, and publish errors.
Tracing:
- Propagate W3C
traceparentand existing Light correlation IDs. - Create ingress spans tagged with Gateway API resource identity:
gateway.namespace,gateway.name,listener.name,route.namespace,route.name,route.rule,backend.service.namespace, andbackend.service.name. - Record upstream selection, retries, and policy decisions as span events without logging tokens, private keys, or sensitive headers.
Migration From NGINX Or Traefik
Recommended customer migration:
- Install
light-k8s-gatewaywith a newGatewayClassnamedlight-k8s-gateway. - Keep NGINX or Traefik running for existing
Ingressor Gateway API classes. - Create equivalent
GatewayandHTTPRouteresources for one host. - Validate status, route behavior, TLS, logs, metrics, and backend health.
- Move DNS or load balancer traffic for that host to
light-k8s-gateway. - Repeat host by host.
- Remove the old ingress controller only after route parity and operational dashboards are in place.
An optional Ingress-to-HTTPRoute converter can help customers migrate, but it should be a tool, not part of the runtime request path.
Open Questions
- What is the first supported east/west deployment model: current
light-gatewayas explicit egress gateway, a dedicated sidecar, or ambient? - How much of the current
server.ymllistener contract should remain inlight-runtimeversus moving Gateway API listener binding intolight-pingora? - Should the controller-to-proxy snapshot protocol stay as a small internal gRPC API, or should it adopt an xDS-compatible model early?
- Which Light policy CRDs are required for the MVP: auth, rate limit, header policy, request size, token, or a generic handler-chain policy?
- What is the exact
UpstreamClusterhealth model shared by Kubernetes EndpointSlice, controller-rs discovery, and direct registry sources?
Suggested Implementation Order
- Create
apps/light-k8s-gateway-controllerandapps/light-k8s-gateway-proxywith separate ServiceAccounts and RBAC. - Add controller leader election with Kubernetes
Leaseobjects. - Define
GatewayApiSnapshot,UpstreamCluster,EndpointSet, and the authenticated controller-to-proxy snapshot stream. - Implement proxy-side snapshot loading through
ConfigManager. - Implement
GatewayClass,Gateway,HTTPRoute,ReferenceGrant,Service,EndpointSlice,Secret, andNamespacewatches. - Implement attachment validation, policy validation, status updates, and snapshot publishing.
- Add a
light-pingoraGateway API route table and route HTTP traffic to Kubernetes Service endpoints. - Add memory-only TLS Secret handling and terminated HTTPS listener support
for the common
80/443ingress case. - Add initial Light policy CRDs using Gateway API policy attachment.
- Add Prometheus metrics, OpenTelemetry tracing, and structured logs for the controller and proxy.
- Run HTTPRoute Gateway conformance and close gaps.
- Add multi-listener runtime support.
- Add mesh route compilation for Service-attached
HTTPRouteresources. - Add explicit egress gateway mode for early east/west use.
- Add sidecar interception and run mesh conformance.
- Evaluate ambient/node-local mode after sidecar behavior is proven.
Light-Gateway IPv6 Support
light-gateway can run in IPv4-only, IPv6-only, and dual-stack networks. The
gateway uses light-pingora for the inbound HTTP and HTTPS listener, and uses
the same routing model for IPv4 and IPv6 upstream services.
Configuration
The inbound bind address is controlled by server.ip and projected into
server.yml:
ip: ${server.ip:0.0.0.0}
The default remains IPv4 wildcard binding:
server.ip: 0.0.0.0
Use IPv6 wildcard binding when the host or container network should accept IPv6 connections:
server.ip: "::"
Use a specific IPv6 address when the gateway should bind only to one interface:
server.ip: "fdd0:0:0:1::10"
server.advertisedAddress is separate. It is the address registered with the
controller and shown to peers. Do not set it to 0.0.0.0 or ::; use a stable
DNS name or a reachable address for the deployment:
server.advertisedAddress: ai-microgateway.light-gateway
Listener Behavior
light-gateway validates server.ip as an IP address before starting the
Pingora listener. It then builds the listener socket with the parsed IP and the
configured HTTP or HTTPS port.
Examples:
server.ip: 0.0.0.0, server.httpsPort: 8443 -> 0.0.0.0:8443
server.ip: "::", server.httpsPort: 8443 -> [::]:8443
This avoids the invalid IPv6 address form that results from concatenating the IP and port as strings.
Upstream Routing
Gateway upstream routes can use DNS names, IPv4 literals, or bracketed IPv6 literals.
For proxy.hosts:
hosts: https://[fdd0:0:0:1::20]:8443
For direct-registry.directUrls:
directUrls:
com.example.orders-1.0.0: https://[fdd0:0:0:1::21]:8443
Discovery responses may also contain IPv6 addresses. The router and websocket router bracket IPv6 discovery addresses before constructing the upstream authority.
When an upstream is referenced by DNS name, the selected address family depends on DNS resolution and the connector behavior. In dual-stack Docker or Kubernetes networks, make sure the backend listens on the address family that DNS returns first, or use service discovery/configuration that points to a reachable address.
Native Deployment
For native host deployment, keep the bind address aligned with the host network:
server.ip: "::"
server.advertisedAddress: gateway.example.com
server.httpsPort: 8443
server.enableHttps: true
Verify the host firewall and TLS certificate cover the advertised hostname.
Kubernetes Deployment
For Kubernetes, use IPv6 binding only when the cluster and service are intended to expose IPv6 traffic:
server.ip: "::"
server.advertisedAddress: ai-microgateway.light-gateway
The Service, pod network, DNS policy, and any ingress or Gateway API resources must also support IPv6. The gateway bind address alone does not make the cluster dual-stack.
Verification
Inside the same network namespace or from a peer pod/container:
getent ahosts <gateway-service-name>
curl -k -g https://[<gateway-ipv6>]:8443/health
curl -k -v https://<gateway-service-name>:8443/health
For an upstream backend reached through light-gateway, confirm both sides:
getent ahosts <backend-service-name>
curl -k -v https://<gateway-host>/<gateway-route>
If the gateway log shows connection refused to an IPv6 upstream address, the backend service is likely not listening on IPv6 or the network does not route that address family.
How to Call an MCP Server with Curl
To talk to an HTTP-based Model Context Protocol (MCP) server using curl, you must follow the strict JSON-RPC 2.0 lifecycle defined by the spec. This includes initiating a handshake, completing an initialization confirmation, and executing the actual tool call.
Here is the exact multi-step process required to interact with a streamable HTTP or Server-Sent Events (SSE) MCP server.
1. Initialize the Connection
Every MCP interaction requires a handshake. You must send an initialize method to create your session.
curl -s -i -X POST "https://your-mcp-server.example.com/mcp" \
-H "Content-Type: application/json" \
-d '{"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {}}'
Action: Extract the mcp-session-id from the response headers and export it (e.g., export SESSION_ID="...").
2. Confirm Initialization
Send an initialized notification to finalize setup.
curl -s -X POST "https://your-mcp-server.example.com/mcp" \
-H "Content-Type: application/json" \
-H "Mcp-Session-Id: $SESSION_ID" \
-d '{"jsonrpc": "2.0", "method": "initialized", "params": {}}'
3. List and Call Tools
Use tools/list to find available tools, and tools/call to execute them, ensuring arguments are structured correctly.
List Tools
curl -s -X POST "https://your-mcp-server.example.com/mcp" \
-H "Mcp-Session-Id: $SESSION_ID" \
-d '{"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {}}'
Call Tool
curl -s -X POST "https://your-mcp-server.example.com/mcp" \
-H "Mcp-Session-Id: $SESSION_ID" \
-d '{"jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": {"name": "...", "arguments": {}}}'
Tips
- Auth: Add
-H "Authorization: Bearer $TOKEN"for protected servers. - Streaming: Use
curl -Nfor SSE endpoints.
MCP Tools Access Control
light-gateway can enforce fine-grained access control for MCP tools exposed
through the MCP router. The router uses the shared access-control runtime from
light-pingora, so MCP tools use the same access-control.yml and rule.yml
policy files as HTTP API access control.
For MCP traffic, rules apply only to tools/call requests:
req-accrules run before the downstream HTTP or MCP tool is called.res-filrules run after the downstream response is converted to an MCP result and before the JSON-RPC response is returned to the agent.
tools/list, initialize, notifications/initialized, and session
management requests are handled by the MCP router and are not authorized as
individual business tools.
MCP Router And access-control.enabled
The access-control.enabled flag is the top-level switch for MCP tool access
control:
enabled: true
accessRuleLogic: any
defaultDeny: true
defaultInclude: false
skipPathPrefixes: []
When access-control.enabled is true, the MCP router evaluates configured
req-acc rules before invoking a tool. If the tool call is allowed and the
matching endpoint has res-fil rules, the router also applies response row or
column filters before returning the MCP result.
skipPathPrefixes bypasses the same two phases for matching MCP tool names or
matching endpoint keys. For example, if skipPathPrefixes contains
local_mcp, then tool names such as local_mcp_echo are allowed without
req-acc evaluation and their results are returned without res-fil filtering,
even when the policy endpoint key is something else, such as accounts@call.
Endpoint-key prefixes still work. If skipPathPrefixes contains accounts,
then an accounts@call endpoint is also allowed without req-acc evaluation
and returned without res-fil filtering.
When access-control.enabled is false, the MCP router bypasses both phases:
req-accrules do not deny MCP tool calls.res-filrules do not alter MCP tool results.
This bypass applies even when rule.yml is still present and contains matching
endpoint rules. The rules can remain loaded for later re-enable or reload, but
the disabled access-control switch prevents the MCP router from enforcing
authorization or response filtering.
This setting is independent from mcp-router.enabled. Set
mcp-router.enabled: false to disable the MCP endpoint itself. Set
access-control.enabled: false only when the MCP endpoint should continue to
serve tools without access-control enforcement.
Endpoint Rules
Each MCP tool maps to a stable endpoint key. If the tool config contains an
explicit endpoint, that value is used. Otherwise, the router derives a key
from the tool name and method, such as accounts@call.
tools:
- name: accounts
description: List accounts
targetHost: http://account-api:8080
path: /accounts
method: GET
endpoint: accounts@call
apiType: http
The same endpoint key is referenced from rule.yml:
endpointRules:
accounts@call:
req-acc:
- allow-account-reader
res-fil:
- filter-account-rows
- filter-account-columns
permission:
roles: teller manager
row:
role:
teller:
- colName: accountType
operator: "="
colValue: C
col:
role:
teller: accountNo,accountType,balance
Request Authorization
A req-acc rule decides whether the MCP tool call can proceed. When
defaultDeny is true, a tool call with no matching endpoint rule or no
req-acc rules is denied.
defaultDeny only controls the fallback behavior when the MCP router cannot
find request access rules for a tool endpoint. It does not disable
access-control globally and it does not bypass configured rules.
When defaultDeny is true, the MCP router fails closed:
- If the tool endpoint has no entry in
rule.endpointRules, the tool call is denied. - If the tool endpoint has an
endpointRulesentry but noreq-accrule IDs, the tool call is denied. - If
req-accrules are configured, the rule result decides whether the tool call is allowed.
When defaultDeny is false, the MCP router permits tool calls that do not
have request access rules:
- If the tool endpoint has no entry in
rule.endpointRules, the tool call is allowed. - If the tool endpoint has an
endpointRulesentry but noreq-accrule IDs, the tool call is allowed. - If
req-accrules are configured, the rule result still decides whether the tool call is allowed.
Use defaultDeny: false when the gateway should expose protected MCP tools
without fine-grained authorization rules for every tool endpoint. This avoids
creating no-op req-acc rules only to make unrouted tools callable. Keep
defaultDeny: true when every MCP tool must have an explicit access-control
policy.
For example, this configuration keeps access-control enabled but allows MCP
tool calls that have no matching rule.endpointRules entry:
enabled: true
accessRuleLogic: any
defaultDeny: false
defaultInclude: false
skipPathPrefixes: []
ruleBodies:
allow-account-reader:
common: Y
ruleId: allow-account-reader
ruleName: Allow account reader
ruleType: req-acc
conditionLanguage: cel
conditionSecurityProfile: strict
expression: >
auditInfo.subject_claims.ClaimsMap.role in ["teller", "manager"]
actions:
- actionClassName: com.networknt.rule.RoleBasedAccessControlAction
Response Filtering
A res-fil rule transforms the MCP tool result after the downstream call
succeeds. Row filters and column filters operate on the JSON payload carried in
the MCP result structuredContent and mirrored text content.
ruleBodies:
filter-account-rows:
common: Y
ruleId: filter-account-rows
ruleName: Filter account rows
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: "true"
actions:
- actionClassName: com.networknt.rule.ResponseRowFilterAction
filter-account-columns:
common: Y
ruleId: filter-account-columns
ruleName: Filter account columns
ruleType: res-fil
conditionLanguage: cel
conditionSecurityProfile: strict
expression: "true"
actions:
- actionClassName: com.networknt.rule.ResponseColumnFilterAction
defaultInclude affects row filtering when no caller claim matches a configured
row-filter entry. Keep it false to fail closed and return no rows. Set it to
true only when the desired compatibility behavior is to keep all rows.
Response Filtering And Call Caches
The current MCP router has a tools/list visibility cache but no gateway
tools/call response cache. If call-result caching is added later, access
control wraps the cache rather than being cached with the result:
- Authenticate and run
req-accbefore using a cache entry. - Cache only the normalized backend MCP result before
res-fil. - Run the current caller’s
res-filrules on every hit and miss. - Never place a post-filter caller view in a shared call-result cache.
The pre-filter cache is sensitive because it may contain rows or columns that the current caller cannot receive. It must remain tenant/principal scoped by default, use a key covering every downstream request dimension, be bounded and short-lived, and be unavailable to tenant code or diagnostics that expose payloads. Cross-principal reuse requires explicit proof that the backend result is principal-invariant. If the gateway cannot rerun current authorization and filtering on a hit, caching is disabled for that tool.
Portal-Managed MCP Tool Access Control
Status
- Decision state: Proposed
- Owners: Light Portal and Light Gateway maintainers
- Design date: 2026-08-23
- Scope: Caller authorization and response filtering for Portal-managed MCP
tools published to
light-gateway
Purpose
Portal users can manage access rules and permissions for API endpoints from
API Admin, but a standalone tool created in Tool Admin has no equivalent
access-control workflow. This forces operators to either maintain gateway-local
configuration or add the tool name to access-control.skipPathPrefixes.
Neither is an acceptable steady state. Local files override config-server snapshots, and a skipped MCP tool bypasses both request authorization and response filtering.
This design adds first-class Tool Access Control to Light Portal while keeping the existing Light Gateway policy format and enforcement behavior.
Current Contracts
Gateway Tool Identity
Every published tool has an authorization endpoint key. An explicit tool
endpoint is used when present; otherwise the gateway derives the key as
{path}@{method}. It does not derive the key from the tool name. Because
path defaults to an empty string, a workflow-backed tool with neither
endpoint nor path resolves to @call; multiple such tools therefore
collapse onto the same policy key. Managed tools must not rely on that default.
Explicit logical identities may use names such as:
customer_360@call
workflow_mcp_smoke@call
The same endpoint key is used for:
tools/callrequest authorization;- MCP response filtering;
- optional
tools/listvisibility filtering; - policy logs, metrics, and audit context.
The endpoint key, not toolId or stableToolRef, is the runtime lookup key in
rule.endpointRules.
Three adjacent tool identities must remain visible and distinct:
endpointKeyselects the access-control endpoint rule;authorizationToolName, when configured intoolMetadata, is the tool name used by call authorization, response filtering, and CEL list authorization;endpointNameis the downstream MCP backend operation name.
The public tool name remains the lookup name for tools/call and is currently
used by permission-mode tools/list authorization. Portal preview must display
all four values so an operator can see any divergence.
Gateway Policy Format
The gateway already accepts tool policies through the standard access-control snapshot:
rule.endpointRules:
customer_360@call:
req-acc:
- req-access-light-portal.lightapi.net
permission:
roles: admin customer-service-agent
No new gateway policy document or tool-specific runtime handler is required.
tools/call remains the final enforcement point even when tools/list
visibility filtering is enabled.
Portal API Access Management
Endpoint Access Overview manages rules, permissions, and response filters for
an api_endpoint_t record. The current permission and filter projections are
therefore API-endpoint-centric. API access publication compiles those records
into rule.endpointRules and rule.ruleBodies for a target gateway instance.
Portal Tool Management
A standalone Tool Admin record can have no endpointId. Workflow-backed tools
created directly in Tool Admin commonly have this shape. They can be published
into mcp-router.tools, but they cannot enter Endpoint Access Overview and
cannot contribute access rules to the gateway snapshot.
Workflow Tool Access is a separate security boundary. It grants a workflow definition permission to use a pinned dependency. It does not authorize an end user or agent to invoke a gateway-published MCP tool.
Current Storage Reality
rule.endpointRules and rule.ruleBodies are registered as map properties.
Portal currently contains more than one resolution implementation, so the
physical table layout alone does not determine effective behavior.
The canonical portal-db create_snapshot procedure copies only active
instance, instance-API, instance-app, instance-app-API, association, and lower
inheritance rows. It combines the four instance-scoped sources into one
InstancePool; map properties are merged with jsonb_object_agg and recorded
as source level instance_merged. The runtime-config query used by the deployed
Portal service likewise reads active rows from all four instance scopes, and
the config-server assembler merges map values before generating values.yml.
ConfigPersistenceImpl.insertEffectiveConfigSnapshot, however, contains an
alternate Java resolver that omits active filters and applies
MAX(effective_value) across per-API groups. That path would select one whole
JSON value lexicographically rather than perform the canonical cross-API merge.
It must be removed, delegated to create_snapshot, or brought under the same
contract before carrier migration is considered portable across deployments.
Its presence is a resolver-parity defect; it is not evidence that the current
loc gateway is dropping 41 API policies.
On 2026-08-23, the running gateway’s host, service ID, and loc environment
tag resolved to portal-bff-loc. That instance had 42 active instance-API rows
for each rule property. The per-API endpoint maps contributed 831 distinct keys
and the instance row contributed one more. Both the running gateway cache and
the latest database snapshot contained all 832 keys; the snapshot reported
source level instance_merged. The loc deployment therefore has multiple
contributors but is not currently suffering the Java MAX winner loss.
The same deployment explicitly sets access-control.enabled: true and
defaultDeny: true, then skips workflow_mcp_smoke and customer_360. It is a
protected instance with two local bypasses, not an instance masked by the
shipped disabled default.
The current API access compiler is scoped to one (hostId, apiId, apiVersion)
and replaces the complete endpoint-rules and rule-bodies values on its selected
instance_api_property_t rows. Whole-property replacement is therefore a
hazard within one contributor row, while the canonical snapshot/runtime merge
retains unrelated API contributors. It still lacks explicit endpoint ownership,
duplicate-key conflict rejection, and a publication event describing which
source owns each merged key. Those are pre-existing API publication gaps, not
hazards introduced only by Tool Access.
The canonical database can mechanically merge instance-pool maps, but that is
not sufficient ownership control: it cannot reject two sources claiming the
same endpoint key, and a standalone tool has no natural instance-API carrier.
The map branch also calls jsonb_object_agg without ORDER BY. PostgreSQL
keeps one value for a duplicate key, so two contributors claiming the same
endpoint can produce an order-dependent winner. Portal must reject that
collision before it reaches the resolver and compute the publication digest
from its own canonical, key-sorted pre-merge; digest determinism must not be
inherited from database input order.
As defense in depth, the database merge should use a documented total order.
Ordering only by update_ts is insufficient when timestamps tie; the order
must also include stable source rank and source identity. That deterministic
tie-break still does not establish ownership or make duplicate keys valid.
Portal must therefore merge every active API and tool contribution with
explicit ownership before writing one canonical property value.
The existing mcp-router.tools publisher is the precedent: it performs an
identity-keyed mergeExistingTools, writes the authoritative result to
instance_property_t, and retires legacy per-API property rows.
Problem Statement
The control plane has two independently working halves:
- Tool publication creates the MCP catalog entry.
- API access publication creates endpoint access rules.
There is no Portal-owned connection between a standalone tool and the access rules for its published endpoint key. As a result:
- Tool Admin cannot assign roles, groups, positions, attributes, or users;
- Tool Admin cannot attach
req-accorres-filrules; - policy publication cannot prove that a protected tool has a matching rule;
- tool lifecycle changes can leave manually maintained policy keys stale;
- operators may use
skipPathPrefixes, which bypasses enforcement; tools/listvisibility andtools/callauthorization can drift.
Goals
- Add an Access Control action to Tool Admin.
- Reuse the established Endpoint Access Overview user experience.
- Give API endpoints and tools one logical access-target contract.
- Publish tool rules into the existing
rule.endpointRulesandrule.ruleBodiesproperties. - Use the exact endpoint key produced by tool publication.
- Keep
defaultDeny: trueas the safe fallback. - Support roles, groups, positions, attributes, users, request rules, response filters, and list visibility.
- Make create, update, retirement, replay, and republish deterministic.
- Keep tenant, instance, environment, and publication ownership explicit.
- Remove the need for gateway-local tool bypasses.
Non-Goals
- Replacing Light Gateway’s access-control runtime.
- Combining caller authorization with Workflow Tool Access grants.
- Changing workflow binding, digest, or environment validation.
- Making MCP authentication optional.
- Creating a second tool-only gateway policy format.
- Treating
tools/listfiltering as a substitute fortools/callauthorization. - Requiring users to model every standalone tool as a real HTTP API.
Proposed Control-Plane Model
Access Target
Introduce a Portal access-target abstraction:
AccessTarget
hostId
accessTargetId
targetType API_ENDPOINT | TOOL
targetId endpointId | toolId
endpointKey exact Light Gateway policy key
sourceVersion source aggregate version
active
An access target is a control-plane identity. The gateway continues to receive
only endpoint-keyed rules and does not need to understand targetType or
targetId.
For an API endpoint:
targetType = API_ENDPOINT
targetId = api_endpoint_t.endpoint_id
endpointKey = the published path-or-logical-operation key
For a tool:
targetType = TOOL
targetId = tool_t.tool_id
endpointKey = the authorization endpoint from the compiled tool publication
accessTargetId should be stable and host-scoped. It must not change when the
display name or description changes.
Endpoint-Key Authority
The tool publication compiler is authoritative for endpointKey. Tool Access
Control must consume the compiled authorization endpoint rather than
independently reconstructing it from a mutable display name.
Portal should require a stable explicit endpoint key before a tool becomes
access-policy-ready. Migration must read the endpoint key from the compiled
tool publication; it must never reconstruct it from a name. Portal may offer
to pin a non-empty compiled {path}@{method} value after showing it in preview.
It must not auto-pin @call when the path is empty, because that value is not
tool-unique.
Publication must reject:
- an empty endpoint key;
- an endpoint key without the expected operation suffix;
- exact duplicate keys in one gateway instance;
- a prefix or path-template rule that would shadow the endpoint key;
- a tool and API endpoint claiming the same key with different policy owners;
- a policy whose source version does not match the selected tool publication.
stableToolRef remains the immutable tool identity used by workflow bindings
and grants. It must not be substituted for the endpoint key in gateway policy.
Permissions and Rules
Access-target assignments should represent the existing principal dimensions:
- roles;
- groups;
- positions;
- attributes and attribute values;
- users.
Rule bindings should support:
req-accrequest authorization;res-filresponse row filtering;res-filresponse column filtering;- list visibility derived from the call permission block for protected tools; and
- one canonical, audited allow-all
req-accrule for public tools.
An explicit visibility block short-circuits claim evaluation in both
tools/list and the early claim-only stage of tools/call. The first release
therefore must derive visibility from permission rather than accept an
independently broader value. If independent visibility is added later, the
publisher must prove visibility is a subset of permission for every
principal dimension under the same claim mappings.
The logical model should be shared by API endpoints and tools. A staged schema migration may preserve the existing endpoint-specific tables while generic access-target projections are introduced, but new Tool Admin behavior must not create a second gateway policy representation.
Portal User Experience
Tool Admin
Add an Access Control row action for active, publishable tools. Keep it visually and semantically separate from Workflow Access:
| Action | Security boundary |
|---|---|
| Access Control | Which users and agents may discover or invoke this published tool |
| Workflow Access | Which workflow definitions may use this tool as a pinned dependency |
The Tool Access Overview header should display:
- tool name and
toolId; stableToolRef;- execution placement;
- explicit authorization endpoint key;
- effective
authorizationToolName; - downstream
endpointName; - current tool version and aggregate version;
- selected gateway instance and environment;
- policy readiness and publication status;
- last published snapshot revision.
Reused Access Panels
Reuse the existing access-management panels with an AccessTargetContext
instead of an API-only route context:
hostId
targetType
targetId
endpointKey
instanceId
environment
The overview should expose:
- request rules;
- role permissions;
- group permissions;
- position permissions;
- attribute permissions;
- user permissions;
- row filters;
- column filters;
- tools-list visibility;
- preview and publication status.
The components may retain API-specific adapters during migration, but commands and queries should use the generic access-target identity at their boundary.
Readiness States
Tool Admin should show one of these states:
| State | Meaning |
|---|---|
UNCONFIGURED | No access target or request rule exists |
CONFIGURED | Policy exists but is not published to the selected instance |
PUBLISHED | Active snapshot has the exact owned rule and matching tool and policy source versions |
STALE | Desired property values differ from the current snapshot, or the reviewed base values changed before apply |
CONFLICT | Another owner or runtime pattern collides with or shadows the endpoint key |
RETIRED | Tool or access target is inactive and its policy is being removed |
State precedence is RETIRED, CONFLICT, STALE, UNCONFIGURED,
CONFIGURED, then PUBLISHED; the UI should retain all secondary reasons.
Portal must not infer that an unconfigured tool is non-callable from
defaultDeny alone: the gateway may fall back to a prefix or path-template
rule. A protected publication requires an exact owned endpoint rule and must
block every conflict, stale source, and shadow match.
Commands, Queries, and Events
The exact Portal service names can follow existing naming conventions, but the contract should cover these operations:
Queries
- get access overview by target;
- list rules and principal assignments;
- preview compiled endpoint rules;
- list target gateway instances and environments;
- compare source versions with the active snapshot;
- report endpoint-key ownership conflicts.
Commands
- create or reactivate an access target;
- set or retire a stable endpoint key;
- assign or remove principals;
- attach or detach request and response rules;
- preview list visibility derived from permission;
- publish access policy to a gateway instance;
- retire tool-owned policy from an instance.
Commands must use expected aggregate versions. Events must carry hostId,
targetType, targetId, accessTargetId, and the source aggregate version so
replay cannot attach a policy to the wrong tenant or stale tool revision.
Publication Design
Inputs
A tool access publication candidate contains:
- the selected tool publication and binding;
- the compiled tool endpoint key;
- the effective tool name,
authorizationToolName, andendpointName; - tool ID, stable reference, version, and aggregate version;
- target instance and environment;
- access-target rules and principal assignments;
- current
rule.endpointRulesandrule.ruleBodiesownership metadata; - the resolved property IDs, active flags, and
value_typeregistrations; - the current snapshot ID and digest used as the review baseline; and
- the existing carrier values, digests, and aggregate versions used for the compare-and-apply guard.
Compilation
Compile a tool access target into the existing gateway shape:
rule.endpointRules:
customer_360@call:
req-acc:
- req-access-light-portal.lightapi.net
permission:
roles: admin customer-service-agent
groups: customer-operations
Rule bodies remain deduplicated by rule ID in rule.ruleBodies. Endpoint maps
and rule ID lists must be deterministically ordered before digesting or
publishing. Protected-tool list visibility is derived from the permission block
in the first release; it is not a separately editable allow set. A public tool
instead binds the canonical public rule described below.
Ownership and Merge Rules
Portal must merge all active contributors into one physical carrier per target
gateway instance. The authoritative homes are single instance_property_t
rows for rule.endpointRules and rule.ruleBodies, keyed by
(hostId, instanceId, propertyId). API and tool publishers must not write
independent complete values at different source levels and expect the config
server to merge them.
Per-source ownership belongs in Portal publication and binding records, not in separate config-property rows. Each compiled endpoint contribution records:
sourceType API_ENDPOINT | TOOL
sourceTargetId
sourceVersion
endpointKey
publicationId
instanceId
ruleBodyIds
ruleBodyDigests
The instance policy compiler must:
- load every active API-endpoint and tool contribution for the instance;
- replace only the selected source’s desired contribution in Portal state;
- reject exact ownership collisions and runtime pattern shadowing;
- retain unrelated API and tool contributions;
- remove only endpoint keys owned by a retired source;
- compute the union of referenced rule IDs and reject one rule ID with different bodies or digests;
- retain each rule body while at least one active endpoint references it and retire it only after its last reference disappears;
- deterministically rebuild the complete endpoint and rule-body maps;
- assert that both selected config properties are active maps with the expected property IDs;
- in one publication transaction, write both canonical instance values and deactivate every legacy instance-API and instance-app-API row for those properties against the accepted target revision.
This follows the identity-keyed merge shape already used by
mergeExistingTools. Resolver convergence is a strict predecessor to carrier
migration, not parallel work in the same rollout. Every supported snapshot and
runtime-config path must honor inactive rows and the canonical instance-pool
merge. The deployed create_snapshot procedure and runtime-config query meet
that requirement; the alternate Java MAX resolver does not. It must delegate
to the canonical procedure, implement identical semantics, or be proven
unreachable and removed. Physical deletion is not an acceptable compatibility
workaround because it discards the audit trail.
The migration transaction absorbs all active legacy contributions, writes the
merged instance carriers, and deactivates the higher-priority legacy rows as
one publication. Only the snapshot generated after that transaction may be
used for verification. Readiness permanently asserts that no active
instance_api_property_t or instance_app_api_property_t row exists for
rule.endpointRules or rule.ruleBodies on the target instance. The legacy
API publisher must be changed in the same phase so it cannot recreate one.
The tool publisher’s apiVersionId remains selection scope and provenance; it
does not make an API property row the owner of standalone Tool Access.
Runtime Matching and Conflict Analysis
Gateway request lookup tries an exact endpoint key first, then prefix and path
template matches, choosing the longest matching pattern. A missing @method
is treated as @call. Portal conflict analysis must reproduce those rules,
not merely compare strings.
For managed Tool Access, an exact rule owned by that access target is required
for PUBLISHED. Pattern fallback is legacy runtime compatibility and cannot
satisfy readiness. Preview must show every exact collision, ancestor-prefix
match, path-template match, implicit @call match, and the rule that the
gateway would select. This avoids disagreement with the gateway’s exact-only
validate_request_policy readiness check.
Effective Policy Context and Digest
Preview must load these effective instance settings from the target snapshot:
access-control.enabled,defaultDeny,accessRuleLogic, anddefaultInclude;skipPathPrefixesand its prefix relation to the endpoint key, public tool name, andauthorizationToolName;toolsListAccessControl.mode,unknownRuleFallback,maxCelEvaluations, and claim mappings.
The preview digest must include those values, compiled tool identities, all source owners and versions, the endpoint shadow report, referenced rule-body digests, property IDs and value types, legacy-row inventory, property actions, and accepted config revision. Identical property JSON is not an identical authorization outcome when the instance-global settings or winning source level differ. Portal computes this digest before persistence from a canonical, key-sorted ownership merge, then requires the generated snapshot to reproduce the same JSON and digest. The database aggregate is a verification target, not the digest authority.
Desired State and Promotion Behavior
Tool and access-policy desired state share one reviewed Portal workflow. The
review page compares the existing instance_property_t values with the proposed
canonical values and shows additions, removals, changed rules, public/protected
mode changes, and unrelated entries that will be retained. Applying the review
writes the catalog and policy carrier values with their expected aggregate
versions; a concurrent change rejects the apply and requires the comparison to
be refreshed.
Writing instance_property_t changes desired state but does not change a
running instance. Portal then creates and validates a candidate snapshot from
that desired state. Only an explicit promotion that makes the candidate current
activates the change. If property persistence, snapshot generation, validation,
or promotion fails, the previous current snapshot remains authoritative. Portal
reports the tool policy as PUBLISHED only after the current snapshot contains
the reviewed tool and policy values. The control plane must never temporarily
add the tool to skipPathPrefixes to bridge these stages.
Gateway Runtime Behavior
Protected tools reuse the existing endpoint-keyed authorization algorithm. The
gateway needs two bounded extensions: recognize the canonical public rule as
visible in permission-mode tools/list, and resolve the restricted nested
response path before applying the existing row or column filter actions. The
normal tools/call rule engine executes the public rule’s constant-true CEL
condition. Neither extension introduces a second policy source or bypasses
response filtering.
For tools/call:
- Resolve the configured tool.
- Resolve its authorization endpoint key.
- Apply
access-control.enabledandskipPathPrefixesgates. - Look up the exact endpoint rule, then any compatible prefix or path-template rule using longest-pattern precedence.
- Evaluate
req-accwith the authenticated principal, permission metadata, headers, and tool arguments. - Invoke the tool only when allowed.
- Apply configured
res-filrules to the normalized result.
For tools/list, the gateway supports none, permission, and cel modes.
Protected deployments should use permission or CEL mode and must show the
effective choice in Portal. CEL evaluates no more than maxCelEvaluations;
tools after that bound are hidden. Call authorization remains mandatory because
list visibility can be stale or argument-insensitive.
skipPathPrefixes is a prefix test applied to both the endpoint key and a tool
name. Today permission-mode list filtering passes the public tool name, while
CEL list filtering, call authorization, and response filtering pass
authorizationToolName. Portal must test every skip prefix against all three
values and block publication on any match. The gateway should separately align
permission-mode list filtering to authorizationToolName; until then,
acceptance tests must cover both list and call paths when the two names differ.
The loc value customer_360 demonstrates why this is a prefix check, not an
equality check: it bypasses customer_360, customer_360_v2, and every other
endpoint key or authorization tool name beginning with that string.
For response filtering, any unfilterable MCP result, missing rule body, rejected rule, rule-execution error, missing filtered body, serialization/application failure, or top-level row denial must suppress the original payload and return a bounded MCP access-control error. No error path may return the unfiltered backend result.
Security Requirements
- Protected gateway instances use
access-control.enabled: true. - Protected gateway instances use
defaultDeny: true. - Preview and readiness include
accessRuleLogic,defaultInclude, list mode, unknown-rule fallback, CEL limit, and effective claim mappings. - Portal publication never creates
skipPathPrefixesfor managed tools. - A prefix match against the endpoint key, public tool name, or
authorizationToolNameis reported as a policy-readiness error. - Every mutation and publication is host-scoped and aggregate-versioned.
- Cross-tenant target IDs and endpoint ownership are rejected.
- An explicit access policy cannot silently fall back to another endpoint key.
- Missing exact rules, unknown rule bodies, stale source versions, ownership
collisions, and pattern shadowing fail publication closed. At runtime a
missing request-rule body denies a call; permission-mode list visibility may
apply
unknownRuleFallback, so list visibility alone is never proof of call authorization. tools/callalways reauthorizes regardless oftools/listvisibility.- Response filtering uses the same endpoint key and principal as request authorization.
- Audit records identify the user, target, endpoint key, source versions, instance, environment, and publication digest.
- Secrets and bearer tokens are never persisted in access-target events or policy snapshots.
Lifecycle Behavior
Tool Update
Description and metadata changes do not change the endpoint key. A tool version or access-relevant schema change marks the publication stale until republished.
Tool Rename
A display-name change must not implicitly rename a managed endpoint key. An endpoint-key change is an explicit migration that publishes the new denied or protected key before retiring the old key.
Tool Retirement
Retiring a tool deactivates its access target and removes only its owned endpoint-rule contribution from selected gateway instances. Shared rule bodies are retained while another active endpoint references them. Publication must reject the same rule ID supplied with different body content; retirement removes a body only after its reference set becomes empty.
Replay
Projection replay must produce the same access target, ownership records, compiled endpoint map, and digest. Replayed stale events cannot overwrite a newer aggregate version or active publication.
Migration Plan
Phase 0: Demonstration Bridge
For an immediate demonstration, operators may create an internal API version
with CALL endpoints whose endpoint strings exactly match the tool keys, then
use Endpoint Access Overview and the current API access publisher. This is a
temporary authorization projection, not the final Tool Admin experience.
This bridge is destructive on a shared property row today: the API publisher
compiles one API version and replaces the complete rule.endpointRules and
rule.ruleBodies values. Phase 0 is permitted only on an isolated demonstration
instance where those property rows have no other owner. It must not be used on
an instance with another API or tool policy contributor. Shared-instance
rollout waits for the instance-level merger in Phase 1B.
Do not remove a tool from skipPathPrefixes until the effective snapshot has
its exact owned req-acc and permission entry and no runtime shadow conflict.
Phase 1A: Snapshot Resolver Convergence
- Make every supported snapshot and runtime-config resolver use the canonical
active-row and instance-pool merge semantics. Prefer one implementation by
delegating Java snapshot creation to
create_snapshot. - Extend
config_snapshot_empty_collection_setup.sqland its snapshot tests with active and inactive cases for every override priority, including instance-app-API, instance-API, instance-app, instance, product version, environment, product, and default registrations. Cover fallthrough to the next priority and preservation of intentional empty maps and lists. - Add a parity gate that feeds the same contributors to every remaining resolver and requires identical canonical JSON, source level, and digest.
- Give the database map aggregate an explicit total order by update timestamp, stable source rank, and source identity; add a duplicate-key fixture proving that the resolver is reproducible even though Portal rejects the publication.
- Deploy and prove that convergence before enabling any carrier-migration command on a deployment where the alternate Java path is reachable.
Phase 1B: Authoritative Instance Merger and Access-Target Read Model
- Establish authoritative
instance_property_tcarrier rows for endpoint rules and rule bodies. - Assert active
mapregistrations for both properties on the target host. - Load all API and tool contributors, merge by owned identity, and atomically write the carriers while deactivating legacy higher-priority rows.
- Prevent the API publisher from recreating per-API rows for these properties.
- Add rule-body reference tracking and conflicting-body detection.
- Add the generic access-target identity and query contract.
- Backfill API endpoint access targets.
- Create tool targets from compiled gateway publication bindings.
- Detect exact collisions, pattern shadowing, and stale source versions.
- Keep existing API access commands working through adapters.
Phase 2: Tool Admin Access Control
- Add the Tool Admin Access Control action and overview.
- Reuse permission, rule, and filter panels.
- Add the bounded, schema-validated nested response target path to row- and column-filter editing and preview.
- Keep Workflow Access as a separate action.
- Add readiness and publication status.
Phase 3: Tool Policy Publication
- Compile tool targets into
rule.endpointRules. - Use the Phase 1B ownership and deterministic instance merger.
- Compare proposed carrier values with current
instance_property_t, apply them with expected aggregate versions, create and validate a candidate snapshot, and explicitly promote it to current. - Compile and digest nested response target paths with their output-schema dependency.
- Support retire and replay without deleting unrelated API rules.
Phase 4: Protected Rollout
- Require Phase 1A resolver qualification before any protected multi-API
publication. If a deployment can reach the alternate Java
MAXresolver, freeze further API access publication on an already-enabled instance and do not enable a new instance until resolver convergence and carrier migration succeed. A qualified canonicalinstance_mergeddeployment, including the checked loc instance, is not subject to that winner-swap freeze. - Even on a qualified instance, do not publish Tool Access or remove a tool bypass until Phase 1B carrier ownership and exact-rule validation succeed.
- Configure policies for the demonstration tools.
- Confirm the effective snapshot contains their endpoint keys.
- Remove their
skipPathPrefixesbypasses. - Remove the gateway-local
access-control.ymloverride when all remaining settings are config-server-owned. The gateway loads that file as the complete access-control config and consultsvalues.ymlonly when no local file is found; the two sources are not merged. - Enable permission-based
tools/listfiltering where discovery must match call permissions.
Validation Plan
Persistence and Replay
- API endpoint and tool access targets are host-isolated.
- Duplicate active endpoint keys and runtime pattern shadows are rejected per gateway instance.
- Principal and rule mutations enforce aggregate versions.
- Replay is idempotent and preserves source ownership.
- Tool retirement removes only tool-owned contributions.
- A shared rule body survives until its last reference is retired; conflicting content for one rule ID is rejected.
Publication
- Tool access preview shows the exact endpoint key, public name,
authorizationToolName,endpointName, compiled policy, effective global settings, and shadow analysis. - Unrelated API and tool endpoint rules survive republish.
- Missing exact
req-acc, stale tool versions, ownership conflicts, shadow matches, and skip-prefix matches block a protected publication. - Deterministic input produces a deterministic snapshot digest.
- Two contributors claiming one endpoint key produce
CONFLICTbefore any database write, regardless of whether their compiled rule values are equal. - Repeated snapshot generation of a deliberately colliding resolver fixture is byte-stable under the database total order, while publication of that fixture remains forbidden.
- Concurrent publication rejects a review whose base property value, digest, or aggregate version is stale.
- Carrier write and legacy-row deactivation share one publication transaction; verification uses only a snapshot generated afterward.
- Readiness rejects any active instance-API or instance-app-API row for either property on the target instance.
- Every supported resolver ignores inactive association and property rows and produces the same instance-pool merge.
- Snapshot tests cover active/inactive behavior and priority fallthrough at every override level, including empty map and list values.
- Resolver parity tests require identical canonical JSON and digest; the Java
MAXimplementation cannot remain as an alternate result. - Both properties resolve to active
mapregistrations for the selected host; changing a registration blocks publication until reviewed. - After migration, both effective properties have the canonical
instance_mergedsource level with only the instance carrier contributing, and a later API publication cannot recreate a per-API carrier.
Portal UI
- Standalone tools expose Access Control without requiring
endpointId. - Workflow Access and caller Access Control are clearly distinguished.
- Existing API Endpoint Access Overview behavior is unchanged.
- Readiness states explain the exact missing or stale prerequisite.
- Preview and publish actions show the selected instance and environment.
Gateway
- An authorized caller can list and call the tool.
- An unauthorized caller cannot call the tool.
- Permission-mode
tools/listhides unauthorized tools. - The canonical public rule makes a public tool visible and callable for every caller that reaches MCP routing, while route-level authentication remains authoritative.
- A missing, modified, or mixed public rule fails publication; a missing runtime rule body hides the tool and denies the call.
- CEL-mode
tools/listhides entries beyondmaxCelEvaluations. - A direct
tools/callremains denied even if list visibility is stale. - Row and column filters apply to top-level and configured nested workflow-backed structured content, and missing or mistyped nested targets fail closed.
- Every response-filter failure returns an MCP error without exposing the unfiltered result.
- Missing endpoint rules deny when
defaultDenyis true. - No skip prefix matches a managed endpoint key, public tool name, or
authorizationToolName. - An unqualified multi-API resolver cannot transition from disabled to protected, or accept further protected API publication, until resolver parity passes. Tool Access additionally requires the carrier and exact-rule gates.
Demonstration Acceptance
For customer_360@call and workflow_mcp_smoke@call:
- the config snapshot contains request rules and permissions for both keys;
- no
skipPathPrefixesvalue is a prefix of either endpoint key, public tool name, orauthorizationToolName; - an allowed principal receives a successful result;
- a denied principal receives an access-control error before workflow start;
tools/listbehavior matches the configured visibility mode;- restart and snapshot regeneration preserve the same behavior without local
mcp-router.ymloraccess-control.ymloverrides.
Alternatives Considered
Keep Gateway-Local access-control.yml
Rejected. It overrides config-server ownership, is deployment-specific, and cannot provide Portal audit, preview, lifecycle, or replay behavior.
Keep skipPathPrefixes
Rejected except as a temporary development bypass. It disables both request authorization and response filtering.
Require Users to Create APIs Manually
Useful as a short-term bridge, but rejected as the final user experience. It creates duplicate lifecycle ownership and makes a standalone tool appear to be an API solely to reach permission screens.
Add a Tool-Specific Gateway Policy File
Rejected. The gateway already has the necessary endpoint-keyed rule format. A second format would create precedence, reload, audit, and migration problems.
Reuse Workflow Tool Grants
Rejected. Workflow grants authorize definition dependencies and pin versions, digests, capabilities, and environments. Caller permissions authorize users and agents invoking an exposed tool. Combining them would weaken both models.
Resolved Design Questions
Access-Target Storage Migration
Do not replace the endpoint-specific tables in the first release. Introduce generic access-target storage and projection adapters, migrate readers and writers incrementally, and retire the old tables only after API Endpoint Admin and Tool Admin both use the generic contracts.
The eventual replacement surface is the 16 tables whose authorization data is
keyed directly by endpoint_id:
api_endpoint_rule_t;role_permission_t,group_permission_t,position_permission_t,attribute_permission_t, anduser_permission_t;- the role, group, position, attribute, and user variants of
*_row_filter_t; and - the role, group, position, attribute, and user variants of
*_col_filter_t.
The generic model can collapse those into access_target_rule_t, a
principal-typed access_target_permission_t, access_target_row_filter_t, and
access_target_col_filter_t, all referencing access_target_t. It must retain
attribute values and the optional user validity interval. api_endpoint_t
remains the API catalog identity and api_endpoint_scope_t remains the OAuth
scope projection; neither is replaced by this migration.
This is a high-impact migration because the 16 tables participate in command, query, event-replay, snapshot/export, cascade-lifecycle, and Portal UI paths. Projection adapters keep the tool feature from requiring a flag-day conversion and provide a period in which generic and legacy query results can be compared.
Publication Command
Tool and policy publication use one command and one accepted source revision.
The preview presents Access Control as a separate, explicitly approved section
and compares the complete proposed carrier values with the current
instance_property_t values. The operator reviews the endpoint key,
public/protected mode, principals, filters, removals, retained unrelated entries,
and every overwrite before accepting the desired-state update. Snapshot creation
and current-snapshot promotion remain explicit subsequent gates in the same
workflow.
Review Freshness and Promotion
There is no separate field-by-field “policy publication invalidation” action. The authoritative review artifact is the comparison between the existing and proposed complete property values, plus their canonical digest and expected aggregate versions. Any desired access-relevant change appears in that diff. A display-name or description-only edit that does not alter the compiled values does not create a policy change.
If either carrier value or its aggregate version changes after the review is
rendered, applying that review is rejected as stale and the user must review a
fresh comparison. Once desired state is written, the current snapshot continues
to govern runtime behavior until a reviewed candidate snapshot is explicitly
made current. The STALE readiness state means desired/current snapshot drift,
not automatic revocation of the active policy. Retirement and emergency
disablement remain explicit, audited desired-state changes followed by snapshot
promotion.
Public Tools
Support an explicit PUBLIC access mode in Portal, compiled as a canonical
shared allow-all request rule rather than a new endpoint marker:
rule.endpointRules:
public_tool@call:
req-acc:
- allow-public-access
rule.ruleBodies:
allow-public-access:
common: Y
ruleId: allow-public-access
ruleName: Allow public access
ruleType: req-acc
accessControlEffect: public
conditionLanguage: cel
conditionSecurityProfile: strict
expression: "true"
The compiler owns the rule ID and exact body digest; users cannot edit or
substitute it. A missing or modified body fails publication and runtime access
closed. Permission-mode tools/list needs a small gateway change to recognize
only the canonical rule ID and complete rule shape as visible. It must not infer
public access from arbitrary constant expressions, rule names, or an
accessControlEffect value alone. A public endpoint cannot combine this rule
with principal-specific request rules or permission metadata. The normal call
path still executes the rule, so request auditing and configured response
filters remain active.
Public means no principal-specific per-tool restriction; it does not bypass
authentication enforced before MCP routing. If the MCP route permits anonymous
access, the tool is callable anonymously. Public access must not use
skipPathPrefixes, an empty permission block, or an independently broad
visibility block. Publication requires explicit public-access approval and
records the approver and reason. Moving between PUBLIC and PROTECTED changes
the compared property values and therefore requires a new review and snapshot
promotion.
Permission-Mode Tool Listing
Permission mode is a discovery filter. For each tool, tools/list looks up its
endpoint rule and compares the rule’s permission metadata with the caller’s
normalized role, group, position, attribute, and user claims. It does not run
arbitrary CEL or argument-dependent rules. The subsequent tools/call remains
independently authorized, so list visibility is not an authorization grant.
Use permission mode by default for newly protected instances after the gateway
uses authorizationToolName consistently for list and call and recognizes the
canonical public rule. Keep unknown rules hidden. Existing instances migrate by
previewed opt-in because changing from none can remove tools from clients’
discovery results. CEL mode remains an explicit option for deployments that
accept its evaluation cost and argument-insensitive list semantics.
Nested structuredContent Filtering
The gateway already filters a top-level object, an array of objects, and an
object containing an items array. It does not generically address an array or
object deeper in a composed workflow result. Add a restricted, schema-checked
response target path rather than recursively filtering every object with a
matching field name.
The first nested-filter version should:
- use a bounded path grammar, such as JSON Pointer plus one explicit array-item selector, rather than unrestricted JSONPath;
- validate the path against the published output schema and show the selected object or array in Portal preview;
- apply row filters to the selected array and column filters to its object elements, or apply column filters to one selected object;
- fail closed on a missing path, wrong node type, traversal-limit breach, or schema mismatch;
- bound path depth, selected node count, and filtered response size; and
- continue regenerating textual MCP content from the filtered
structuredContent, as the current gateway does.
The JSON traversal itself is modest. The work is medium-sized because it spans gateway filter semantics and tests, output-schema validation, generic filter storage, Portal editing and preview, publication digests, and migration of existing filters. A fixed pointer to one nested object or array is a reasonable first increment; full recursive or general JSONPath filtering should remain out of scope until its ambiguity and resource limits have a separate design.
Recommended Decisions
- Adopt a first-class Access Target abstraction in Portal.
- Add Access Control directly to Tool Admin.
- Keep Workflow Access separate.
- Treat the compiled tool endpoint as the policy-key authority.
- Require an explicit endpoint for managed tools; permit an operator-confirmed
pin of a non-empty compiled path-derived key, but never auto-pin
@call. - Reuse the
rule.endpointRulesandrule.ruleBodiesproperty format. Compile public access as one canonical shared allow-all request rule rather than introducing an endpoint marker, a second policy file, or a skip prefix. - Store their canonical merged values in one
instance_property_tcarrier per gateway instance and track source ownership in Portal publication records. - Require active
mapregistrations and atomically retire all higher-priority per-API carriers when the instance carrier is first written. - Converge all resolvers on the deployed
create_snapshotactive-row and instance-pool semantics before exposing the carrier-migration command; never substitute physical deletion. - Permanently reject active legacy rule carriers during readiness.
- Define source policy by host and access target, but compile, publish, and assess readiness per gateway instance. Environment is selection metadata, not another merge layer; any future override must be explicit.
- Treat exact endpoint ownership as required readiness and gateway pattern fallback as legacy compatibility only.
- Derive list visibility from permission until Portal can enforce a formal subset invariant.
- Include effective instance-global access settings and runtime shadow analysis in preview and digest.
- Make Portal’s canonical ownership merge the digest authority; treat database ordering only as deterministic defense in depth and reject every duplicate endpoint owner before persistence.
- Keep
defaultDeny: trueand prohibit generated tool bypasses. - Freeze protected publication only on deployments that cannot prove canonical resolver semantics; require carrier migration and exact-rule readiness before publishing Tool Access or removing its bypasses.
- Use a manually managed internal API only on an isolated transition instance with no other rule-property contributor.
- Remove local access-control overrides after Portal-published policies are verified in the effective snapshot.
MCP Tools List Access Control
This document describes a design for filtering MCP tools/list results by the
same access-control policy model that protects MCP tools/call.
The goal is to avoid showing an agent tools that the current user cannot call,
while keeping tools/call authorization as the final enforcement point.
Background
The MCP router exposes tools from multiple backends through one gateway MCP endpoint. A tool can represent a downstream MCP operation:
{
"name": "local_mcp_echo",
"apiType": "mcp",
"endpoint": "echo@call",
"serviceId": "com.networknt.local.mcp-1.0.0"
}
or a downstream OpenAPI endpoint:
{
"name": "demo_offer_decision_api_search_offers",
"apiType": "openapi",
"endpoint": "/offers@get",
"serviceId": "com.networknt.offer.decision-1.0.0"
}
The access-control policy uses the tool endpoint key to enforce req-acc when
the tool is called:
{
"echo@call": {
"req-acc": ["allow-role-based-access-control.lightapi.net"],
"permission": {
"roles": "account-manager",
"groups": "portal.w"
}
},
"/offers@get": {
"req-acc": ["allow-scp-claim-group-access-control.lightapi.net"],
"res-fil": [
"res-column-filter-jwt-claims.lightapi.net",
"res-row-filter-jwt-claims.lightapi.net"
],
"permission": {
"roles": "account-manager teller",
"groups": "portal.w"
}
}
}
With the current runtime behavior, tools/list returns configured tools and
tools/call enforces access-control. This is secure for invocation, but it can
expose unusable tools to an agent.
Problem
The direct way to filter tools/list is to run each tool’s req-acc rules
before returning the list. That is simple but has drawbacks:
tools/listmay run many rule evaluations for every agent discovery request.- Some
req-accrules can depend ontoolArguments, buttools/listhas no call arguments. - Evaluating a call-time rule with empty arguments can hide tools that would be callable with valid arguments, or show tools that later fail for a specific argument value.
tools/callmust still run authorization, so list filtering cannot replace invocation enforcement.
For these reasons, tools/list filtering should be treated as a visibility
optimization, not the authoritative authorization decision.
Design
Add an optional MCP tools-list visibility filter that uses access-control configuration to decide which tools are visible to the current principal.
The configuration belongs in access-control.yml because it controls how the
access-control runtime affects MCP discovery. The MCP router reads the runtime
decision, but the policy switch should live with the rest of the
access-control settings:
enabled: true
accessRuleLogic: any
defaultDeny: true
defaultInclude: false
skipPathPrefixes: []
claimMappings: {}
toolsListAccessControl:
mode: permission
unknownRuleFallback: hidden
maxCelEvaluations: 100
maxCacheEntries: 2000
The filter should support three modes:
| Mode | Behavior |
|---|---|
none | Current behavior. tools/list returns all configured tools after query filtering. |
permission | Recommended default for protected gateways. Use endpoint rules, permission metadata, and JWT claims to cheaply decide tool visibility. |
cel | Optional strict mode. Evaluate configured req-acc rules for each listed tool with empty toolArguments. This is best-effort and must be documented as argument-insensitive. |
tools/call still evaluates req-acc in all modes except when access-control
is globally disabled or skipped by skipPathPrefixes.
Permission Mode
permission mode should use the same endpoint key that tools/call uses:
- Get the tool endpoint key from
tool.endpoint, such asecho@callor/offers@get. - Apply the access-control global gates.
- Look up
rule.endpointRules[endpoint]. - Evaluate the endpoint permission metadata against the authenticated principal claims.
- Return only visible tools.
This mode does not execute arbitrary CEL. It recognizes the common permission shape already used by the MCP router policies:
{
"permission": {
"roles": "account-manager teller",
"groups": "portal.w"
}
}
An endpoint can also provide a dedicated list visibility block. This is the
preferred shape when call-time req-acc is complex or argument-dependent:
endpointRules:
accounts@call:
req-acc:
- allow-complex-financial-check
visibility:
roles: manager teller
groups: portal.w
When visibility is present, tools/list uses it instead of deriving
visibility from permission and known req-acc rule IDs. tools/call still
uses the configured req-acc rules.
The visibility check should normalize both permission values and JWT claim values as string sets. It should accept either space-separated strings or arrays. For example, the following should be treated as equivalent:
{ "roles": "account-manager teller" }
{ "roles": ["account-manager", "teller"] }
The standard dimensions are:
| Permission key | JWT claim keys |
|---|---|
roles | role, roles |
groups | scp, grp, group, groups |
positions | pos, position, positions |
attributes | att, attribute, attributes |
users | uid, user_id, sub |
Claim lookup is against the same normalized claims map used by req-acc CEL:
auditInfo.subject_claims.ClaimsMap
For example, the visibility checker resolves roles by reading
auditInfo.subject_claims.ClaimsMap.role or
auditInfo.subject_claims.ClaimsMap.roles, and resolves groups by reading
auditInfo.subject_claims.ClaimsMap.scp,
auditInfo.subject_claims.ClaimsMap.grp,
auditInfo.subject_claims.ClaimsMap.group, or
auditInfo.subject_claims.ClaimsMap.groups.
If a deployment uses non-standard claim names, add an access-control-wide claim mapping:
claimMappings:
roles:
- custom_roles
groups:
- custom_scope
When a mapping is present for a permission key, the mapped claim names are used
instead of the standard aliases for that key. Keys without a mapping continue to
use the standard aliases. The mapping also applies to built-in request access
and response filters. Existing toolsListAccessControl.claimMappings values are
retained as a compatibility fallback when the corresponding top-level mapping
is absent.
This covers the sample policy where:
local_mcp_echois visible toaccount-manager.local_mcp_get_random_numberis visible tocategory-admin.- OpenAPI tools protected by
allow-scp-claim-group-access-control.lightapi.netare visible when the caller hasportal.winscp.
Rule Awareness
The visibility filter should inspect the endpoint’s req-acc rule IDs and use
accessRuleLogic to combine known checks:
{
"req-acc": [
"allow-role-based-access-control.lightapi.net",
"allow-scp-claim-group-access-control.lightapi.net"
]
}
For known generic rules:
allow-role-based-access-control.lightapi.netmaps topermission.rolesagainst role claims.allow-scp-claim-group-access-control.lightapi.netmaps topermission.groupsagainst group or scope claims.
If accessRuleLogic is any, one known rule match makes the tool visible. If
it is all, every known rule must match.
Unknown custom req-acc rules need a configured fallback. The safer default is
to hide the tool in permission mode unless explicit visibility metadata is
present:
endpointRules:
accounts@call:
req-acc:
- allow-custom-account-access
visibility:
groups: portal.w
This avoids accidentally exposing tools protected by custom call-time logic.
Rules that do not authorize access should be marked so list visibility can ignore them:
ruleBodies:
request-correlation-logger:
ruleId: request-correlation-logger
ruleType: req-acc
accessControlEffect: telemetry
The list visibility checker should ignore rules whose
accessControlEffect is telemetry or none. Rules without an explicit effect
are treated as authorizing rules.
Default Deny And Missing Rules
The list visibility fallback should mirror tools/call fallback behavior:
| Policy state | defaultDeny: true | defaultDeny: false |
|---|---|---|
| No endpoint rule | Hidden | Visible |
Endpoint rule with no req-acc | Hidden | Visible |
Endpoint rule with known req-acc | Visible only when permission matches | Visible only when permission matches |
Endpoint rule with unknown req-acc | Hidden unless list-specific metadata allows it | Hidden unless list-specific metadata allows it |
This keeps issue-165 behavior consistent: defaultDeny: false can expose tools
without requiring no-op rules, while configured access rules still control
tools that have policy.
If an endpoint has explicit visibility metadata, that metadata decides list
visibility regardless of defaultDeny. defaultDeny only applies when no
endpoint rule or no request-access/list-visibility policy is available.
skipPathPrefixes
The MCP router already treats skipPathPrefixes as matching either the MCP
tool name or the endpoint key. tools/list should use the same behavior:
skipPathPrefixes:
- local_mcp
With this configuration, tools such as local_mcp_echo and
local_mcp_get_random_number are visible and callable without access-control
evaluation, even when their endpoint keys are echo@call and
getRandomNumber@call.
CEL Mode
cel mode can be useful when an operator wants the list to follow the exact
configured req-acc expressions and accepts the cost.
In this mode, the router evaluates req-acc for each candidate tool with:
{
"toolArguments": {}
}
This mode should be documented as argument-insensitive. Rules that require
specific toolArguments are not reliable for list visibility. tools/call
remains authoritative.
CEL mode must fail closed. If a req-acc rule fails to evaluate during
tools/list, the tool is hidden and the gateway logs a debug or warn event with
the rule ID, endpoint, and tool name. This makes argument-dependent rules
visible to operators without exposing tools whose list-time authorization could
not be proven.
CEL mode also needs a scale guard. The router should stop evaluating list-time
CEL after maxCelEvaluations candidate tools and hide the remaining unevaluated
tools, or reject the tools/list request with a clear configuration error. The
preferred default is to hide unevaluated tools and log a warning.
Query Filtering Order
The router already supports tools/list query filtering. The recommended order
is:
- Start from configured tools.
- Apply the query or intent filter.
- Apply list visibility.
- Return the filtered MCP tools array.
Applying the query first reduces the number of authorization checks without changing the response semantics, because hidden tools are never returned.
The query value used for filtering and caching must be normalized before it is included in a cache key. At minimum, trim whitespace and lowercase the query. If the router later accepts structured query parameters, sort the parameter names and normalize repeated whitespace before hashing.
Caching
The visibility result is cached per gateway process when
toolsListAccessControl.mode is not none and maxCacheEntries is greater
than zero. The cache key includes:
- Authenticated principal identity, such as
uid,sub, orclient_id. - A stable hash of the normalized claims map used by visibility checks, or the token signature when available. Do not key only by user ID, because a user’s roles or scopes can change between tokens.
- A stable hash of normalized request headers, because CEL mode can inspect
headers as part of
req-accevaluation. - Normalized query string.
The cache is a size-bounded LRU cache. The default maximum is 2000 entries per
gateway process and can be changed with maxCacheEntries. Setting
maxCacheEntries: 0 disables the cache. The MCP router runtime does not carry
this cache across reloads, so MCP router, access-control, and rule reloads
naturally invalidate cached visibility results.
The cache implementation must protect the gateway from high-cardinality query strings generated by agents. The LRU bound limits memory growth; highly unique queries will evict older entries instead of growing the cache without bound.
Security Notes
tools/list filtering improves agent ergonomics and reduces accidental tool
selection. It must not be treated as the security boundary.
The security boundary remains tools/call:
req-accruns before every downstream tool call.res-filruns after eligible downstream responses.- Argument-dependent authorization belongs in
tools/call, nottools/list.
The design should therefore prefer fast, conservative list filtering and keep full rule evaluation on invocation.
MCP Tool Metadata Usage
This document describes how light-gateway MCP tool metadata should be used
for tool search, progressive disclosure, deterministic routing, policy
enforcement, and diagnostics.
The main principle is simple: metadata can help an agent find the right tool,
but tools/call remains the execution and authorization boundary.
Background
The MCP router can expose tools backed by downstream MCP servers and tools backed by OpenAPI endpoints through the same gateway MCP endpoint.
A downstream MCP tool can be represented as:
- name: local_mcp_echo
path: /mcp
method: call
apiType: mcp
endpoint: echo@call
endpointName: echo
protocol: http
productId: gtw
serviceId: com.networknt.local.mcp-1.0.0
endpointId: 019ec75c-72c5-702e-8e42-59dcf1e68cc2
description: Echoes back the input
inputSchema:
type: object
properties:
message:
type: string
required:
- message
toolMetadata:
routing:
domain: MCP0002
semanticNamespace: MCP0002
semanticDescription: Echoes back the input
semanticKeywords:
- echo
- Echoes back the input
semanticWeight: 1.0
sensitivityTier: internal
sourceProtocol: mcp
safety:
read_only: false
idempotent: false
destructive: false
humanApprovalRequired: false
lifecycle:
version: 1.0.0
status: active
read_only: false
destructive: false
An OpenAPI-backed tool can be represented as:
- name: demo_customer_profile_api_get_customer_preferences
path: /customers/{customerId}/preferences
envTag: dev
method: get
apiType: openapi
endpoint: /customers/{customerId}/preferences@get
protocol: http
productId: gtw
serviceId: com.networknt.customer.profile-1.0.0
endpointId: 019e621b-3a4c-78f4-82f5-16ed24f5ba58
description: Get customer preferences
inputSchema:
type: object
properties:
customerId:
type: string
description: Customer identifier.
channel:
type: string
default: portal
description: Requested channel context.
required:
- customerId
toolMetadata:
routing:
domain: Customers
semanticNamespace: API0004
semanticDescription: Get customer preferences
semanticKeywords:
- Customers
- getCustomerPreferences
- Get customer preferences
semanticWeight: 1.0
sensitivityTier: internal
sourceProtocol: openapi
parameters:
customerId: path
channel: query
safety:
read_only: true
idempotent: true
destructive: false
humanApprovalRequired: false
runtime:
cacheTtlSeconds: 60
costTier: low
estimatedLatencyMs: 100
lifecycle:
version: 1.0.0
status: active
read_only: true
destructive: false
Imported catalog data may store inputSchema and toolMetadata as escaped JSON
strings. The router accepts that shape, but hand-authored config should prefer
structured YAML or JSON. Structured metadata is easier to validate, diff, index,
and review.
Current Runtime Boundary
At runtime, the MCP tool config includes:
| Field | Purpose |
|---|---|
name | Gateway-facing tool name exposed to agents. |
endpointName | Backend MCP operation name used when forwarding tools/call to a downstream MCP server. |
description | Human and model-facing summary. |
protocol | Discovery and direct-registry protocol selector. |
serviceId | Service identity used for portal-registry or direct-registry lookup. |
envTag | Optional environment discriminator for service lookup. |
targetHost | Direct base URL override. |
path | HTTP path or MCP endpoint path. |
method | HTTP method, or call for backend MCP calls. |
endpoint | Stable policy endpoint key, such as echo@call or /offers@get. |
apiType | mcp or openapi. |
inputSchema | JSON Schema used for model tool parameters and argument validation. |
toolMetadata | Structured routing, semantic, safety, and governance metadata. |
The gateway tools/list response should stay compact. It exposes the fields
needed by model tool calling: name, description, and inputSchema.
Richer metadata belongs in the catalog/search layer and gateway runtime config. This avoids flooding the model context with operational fields while still making the data available for ranking, policy, routing, diagnostics, and audit.
The current stateful MCP router also accepts params.query or params.intent
on tools/list. This is a case-insensitive substring filter, not a scored
semantic or vector search. It matches the tool name, description, endpoint ID,
selected routing fields, routing.semanticKeywords, and direct values in the
safety and lifecycle objects. The router applies tools-list access control
after this query filter and returns every remaining match without ranking them.
routing.semanticWeight does not affect gateway tools/list matching,
ordering, or visibility. The stateless 2026-07-28 profile does not accept the
legacy query or intent parameters. Rich semantic ranking, including the
weight, belongs to portal catalog search and the agent’s per-turn selection.
Metadata Responsibilities
Use metadata in three layers:
| Layer | Uses metadata for | Should not use metadata for |
|---|---|---|
| Portal or catalog search | Ranking, assignment, filtering, disclosure, governance preview. | Direct backend execution. |
| Agent runtime | Progressive disclosure, placement-aware availability, and per-turn schema selection. | Bypassing the final gateway, runner-lease, workflow, or fixed-service policy for the selected placement. |
light-gateway MCP router | Deterministic routing, argument mapping, access control, response filtering, audit, diagnostics. | Letting model text decide target URLs or service routing. |
The agent can use metadata to decide which tools to offer to the model. The
gateway uses config metadata to decide how an accepted tools/call is executed.
Tool Source And Execution Placement
Gateway discovery is only one tool source. Every effective catalog entry must carry a server-owned execution placement and stable internal tool reference, for example:
gateway: remote API or MCP tool executed throughlight-gateway;runner: shell, filesystem, browser, local MCP, or other capability exposed by an active runner runtime;workflow: typed durable workflow start/status/cancel operation;fixed-service: typed high-value action such as branch, publish, or sign.
Do not intersect the whole catalog with gateway tools/list. Apply an
independent live-availability intersection for each placement:
gateway tools = assigned gateway catalog entries
intersect gateway tools/list and toolsListAccessControl
runner tools = assigned runner catalog entries
intersect execution-profile policy
intersect lease allowedTools
intersect approved runtime capability manifest
intersect live worker/local-MCP enumeration where applicable
effective model tools = authorized union of each placement-specific set
The model-facing tool definition is bound to its internal tool reference, placement, schema digest, and policy snapshot. A returned tool call is dispatched only through that bound placement; the model cannot turn a gateway tool into a local command or vice versa. Model-facing name collisions across placements fail closed or are resolved by deterministic server-owned aliases recorded in the snapshot. Never rely on an unqualified name alone.
A local MCP server uses its sandbox-local tools/list under the runner lease;
it is not expected to appear in light-gateway tools/list. The model broker,
runner control socket, and credential broker are infrastructure channels and
must never be advertised as local tools.
The current long-lived light-agent exposes gateway tools only, so its existing
catalog-to-gateway intersection remains correct. Placement-aware union is
required before enabling coding, browser, filesystem, or personal-edge tools.
Search And Progressive Disclosure
Agents should not send every configured tool to the model. Instead, they should search the effective agent catalog, select a small set of likely tools, then apply the live availability check for each candidate’s execution placement.
Recommended flow:
user prompt
-> load assigned effective agent catalog
-> search metadata and schema text
-> apply safety and policy disclosure filters
-> select top tools for the turn
-> partition candidates by server-owned placement
-> intersect gateway candidates with gateway tools/list
-> intersect runner candidates with lease/runtime/local capability manifests
-> union the independently authorized, collision-free tool definitions
-> send only selected schemas to the model
-> dispatch each selected tool only through its bound placement
The search index should include:
| Metadata | Search use |
|---|---|
name | Exact and alias matching. |
endpointName | Backend operation matching. |
description | General keyword matching. |
routing.semanticDescription | Agent-oriented capability description. |
routing.semanticKeywords | High-value domain and operation terms. |
routing.domain | Business-domain filtering, such as Customers or Offers. |
routing.semanticNamespace | Product, API, or catalog namespace filtering. |
routing.sourceProtocol | Protocol-aware selection between MCP and OpenAPI tools. |
routing.sensitivityTier | Disclosure and governance filtering. |
routing.semanticWeight | Score multiplier for preferred or higher-quality tools. |
inputSchema.properties | Parameter-intent matching, such as customerId, state, or category. |
inputSchema.required | Completeness checks before exposing or calling a tool. |
safety.read_only | Prefer safe read tools when the prompt is informational. |
safety.idempotent | Decide whether retries are safe for identical arguments. |
safety.destructive | Hide or require approval for destructive tools. |
safety.humanApprovalRequired | Route to approval or workflow instead of direct call. |
runtime.costTier | Prefer cheaper tools when multiple tools can satisfy the prompt. |
runtime.estimatedLatencyMs | Prefer faster tools for interactive turns. |
lifecycle.status | Prefer active tools and avoid deprecated or retired tools. |
The search result should be small. A practical default is 3 to 12 tools per turn. Larger lists increase token use and can lead to the model choosing an irrelevant tool.
Schema indexing should be bounded. For complex request bodies, index the
top-level property names and descriptions by default, then include nested
properties only when the importer marks them as semantically useful. Deeply
nested OpenAPI schemas can otherwise flood the index with low-value keywords
and increase false-positive tool matches. semanticKeywords should be the
curated override when schema text is noisy.
Ranking And Semantic Weight
Current Gateway Behavior
The gateway MCP router does not calculate a relevance score. Its stateful
tools/list query is a normalized substring predicate, and
routing.semanticWeight is intentionally ignored. For example, these two tools
are equally eligible for a gateway query match even though their weights differ:
- name: get_customer_preferences
description: Get customer preferences
toolMetadata:
routing:
semanticKeywords: [customer preferences]
semanticWeight: 2.0
- name: search_customer_preferences
description: Search customer preferences
toolMetadata:
routing:
semanticKeywords: [customer preferences]
semanticWeight: 0.5
A stateful request with params.query: customer preferences returns both tools,
subject to access-control visibility. It does not guarantee that the 2.0 tool
appears first.
Current Light-Agent Behavior
light-agent consumes the effective catalog and uses semanticWeight as a
multiplier during per-turn tool selection. The effective catalog projects the
nested metadata value as the top-level camel-case field semanticWeight; merely
adding the value to gateway runtime config does not make gateway tools/list
rank its response.
The implemented local ranking calculation is:
weighted_base =
(
0.75 * skill_keyword_score
+ 1.5 * tool_keyword_score
+ routing_score
+ max(skill_priority, 0) / 10
)
* max(semanticWeight, 0.1)
portal_score = max(
first_available(combinedScore, semanticScore, vectorScore),
0.0
)
final_score =
weighted_base
+ portal_score
+ lifecycle_adjustment
+ informational_safety_bonus
The weight defaults to 1.0 and has a lower bound of 0.1. A zero or negative
configured value therefore reduces a local keyword score but cannot erase it.
The multiplier applies only to the locally calculated keyword/routing/priority
portion. A score supplied by portal semantic search is added afterward and is
not multiplied again by light-agent. Portal vector ranking may already have
applied the effective semantic weight when it produced combinedScore; avoiding
a second multiplication preserves that server-owned score.
For example, assume the following component scores:
skill_keyword_score = 1.0
tool_keyword_score = 2.0
routing_score = 2.0
skill_priority = 3
semanticWeight = 1.5
combinedScore from portal = 0.8 (already weighted by portal, when applicable)
lifecycle = active (+0.25)
informational prompt = read-only/idempotent tool (+0.50)
weighted_base = (0.75 + 3.0 + 2.0 + 0.3) * 1.5 = 9.075
final_score = 9.075 + 0.8 + 0.25 + 0.50 = 10.625
Weight changes relative preference; it does not bypass assignment, lifecycle, sensitivity, approval, or other disclosure filters. It also does not make an unrelated tool a local keyword match. When both the weighted base and portal score are zero, the tool is not a scored candidate.
Candidates are ordered by descending final score. Ties prefer lower cost, then
lower estimated latency, skill sequence, and finally tool name. The selected
catalog names are subsequently intersected with live gateway tools/list, so a
high-weight tool that is not currently executable or visible is still removed.
Portal search can use the same metadata for vector or hybrid search:
- Use
semanticDescription,semanticKeywords,description, schema property descriptions, tags, and categories for embeddings. - Use
routing.domain,semanticNamespace,sourceProtocol,sensitivityTier,read_only,destructive, and assignment state as structured filters. - Use
endpointIdas the stable document ID for evaluation and feedback. - Favor
lifecycle.status: activeoverdeprecated, and excluderetiredtools from normal disclosure. - Use
runtime.costTier,runtime.estimatedLatencyMs, and rate-limit metadata as tie-breakers when several tools can satisfy the same intent.
Disclosure Filters
Search ranking should run after coarse assignment and governance filters.
Before a tool schema is sent to the model, the agent or catalog API should remove tools that are not appropriate for the current principal and task:
| Filter | Behavior |
|---|---|
| Agent assignment | Only include tools assigned through the effective agent catalog. |
| Environment | Match hostId, serviceId, and envTag. |
| Runtime availability | Gateway entries intersect live gateway tools/list; runner entries intersect the active lease, approved runtime manifest, and any live local enumeration. Never use one source to validate another placement. |
| Lifecycle | Hide retired tools and prefer active tools over deprecated tools. |
| Sensitivity | Do not disclose tools above the caller or agent sensitivity allowance. |
| Destructive flag | Hide unless an approval path or guarded workflow is configured. |
| Human approval | Route to approval or workflow instead of direct model execution. |
| Read-only preference | Prefer read-only tools unless the user intent requires mutation. |
| Budget and rate limit | Prefer lower-cost tools and avoid tools whose rate budget is exhausted. |
Disclosure is not authorization. A hidden tool should not be shown to the model,
but a visible tool must still be authorized by tools/call.
Tools List
The gateway tools/list endpoint has two jobs:
- Report which configured tools are currently executable through the gateway.
- Optionally filter the list by access-control visibility.
It should not become the primary semantic search API. The catalog or agent cache is a better place for richer semantic ranking because it can include skill assignment, tags, categories, prompt instructions, feedback, and non-runtime governance data.
For gateway-placed catalog entries, the recommended pattern is:
catalog search selects candidate tool names
gateway tools/list confirms executable visible tools
model receives only confirmed tool schemas
This keeps the runtime boundary clean:
- Catalog search can evolve independently.
- Gateway
tools/liststays protocol-compatible and compact. - Gateway
tools/callremains the final enforcement point.
See MCP Tools List Access Control for list visibility filtering.
Catalog Policy And Gateway Visibility
Catalog policy means the portal-side disclosure decision made before a tool is shown to an agent. It includes agent assignment, skill-to-tool links, environment, sensitivity tier, lifecycle status, approval requirements, and any tenant or persona rules owned by the catalog/control plane.
Gateway visibility means the runtime decision made by light-gateway
toolsListAccessControl when tools/list is requested. It checks the current
token, claims, gateway policy, and live runtime configuration.
Both are needed:
visible tools =
assigned gateway-placed catalog tools
intersect live gateway tools/list
intersect gateway toolsListAccessControl result
Catalog policy prevents irrelevant or unassigned tools from reaching the model.
Gateway toolsListAccessControl prevents the agent from seeing tools that are
not visible to the current runtime principal. Neither replaces tools/call
authorization.
Execution Routing
After the model chooses a tool, the agent calls:
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "demo_customer_profile_api_get_customer_preferences",
"arguments": {
"customerId": "CUST-1001",
"channel": "portal"
}
}
}
The gateway then resolves the configured tool by name and executes it
deterministically.
For apiType: mcp:
- Resolve the backend target from
targetHost, or fromserviceId,envTag, andprotocol. - Establish or reuse the backend MCP session.
- Forward a backend
tools/call. - Use
endpointNameas the backend tool name when present. - Return the backend MCP result to the caller after access-control response filtering.
For apiType: openapi:
- Resolve the backend target from
targetHost, or fromserviceId,envTag, andprotocol. - Start from configured
pathandmethod. - Read
toolMetadata.routing.parameters. - Map arguments into path, query, header, cookie, or body.
- Invoke the HTTP endpoint.
- Convert the HTTP response into an MCP result.
- Return the result after access-control response filtering.
Parameter mapping is what lets one model-facing argument object become a correct HTTP request:
toolMetadata:
routing:
parameters:
customerId: path
channel: query
idempotency-key: header
body: body
The model still submits one flat JSON argument object. The gateway owns the split into HTTP destinations. A mapped argument should be consumed exactly once: path, query, header, and cookie arguments should be stripped from fallback body placement, and body arguments should not also become query parameters.
Header and cookie names should come from admin-approved catalog metadata, not from model output. The gateway should validate those names and block hop-by-hop or security-sensitive headers unless an explicit administrative allowlist permits them.
If no mapping is present, the router falls back to method-based behavior:
GETandHEADarguments become query parameters.- JSON-body methods send the argument object as the request body.
That fallback is permitted only when the configured path contains no
placeholder. A path containing OpenAPI {name} syntax must never reach the
method-based fallback.
Path-template tools should not rely on fallback behavior. For a path such as
/customers/{customerId}, the OpenAPI import or catalog authoring process
should generate:
toolMetadata:
routing:
parameters:
customerId: path
Without that mapping, the gateway cannot safely know whether customerId
belongs in the path, query string, header, cookie, or body. Importers can infer
this from OpenAPI parameter locations, but the runtime config should carry the
resolved mapping explicitly.
The current Rust MCP router rejects a missing or incorrect path mapping when the tool is called. The target contract also fails earlier:
- OpenAPI/LightAPI import must emit one explicit
pathmapping for every path placeholder and reject an ambiguous or missing source parameter; - catalog/config publication must reject unsupported placeholder syntax,
placeholders without matching mappings, mappings whose location is not
path, andpathmappings with no matching placeholder; - gateway configuration validation must repeat these checks before publishing a runtime snapshot or accepting traffic;
- runtime validation remains defense in depth and returns a stable routing error without contacting the backend.
Reload keeps the last known-good runtime snapshot when a newly supplied tool is invalid and reports the rejected tool/config digest through diagnostics. It must not silently omit the mapping and enable fallback behavior.
Validation tests cover missing, extra, duplicate, malformed, percent-encoded, and incorrectly located placeholders at import/publication and gateway config load. Every rejected case asserts that no backend request is sent; the valid case asserts one percent-encoded path-segment substitution and no duplicate query/body placement.
The supported path placeholder syntax should be OpenAPI-style {name}. The
gateway should not infer Spring or Express-style :name segments by default;
importers should normalize or reject non-OpenAPI path-template syntax before
publishing gateway config.
Retries, Caching, And Rate Limits
Safety and runtime metadata can help agents and gateways decide how aggressively to retry or cache calls.
Recommended fields:
toolMetadata:
safety:
read_only: true
idempotent: true
destructive: false
humanApprovalRequired: false
runtime:
cacheTtlSeconds: 60
retry:
enabled: true
maxAttempts: 2
retryOn:
- timeout
- 502
- 503
- 504
rateLimit:
key: customer-profile-read
costUnits: 1
costTier: low
estimatedLatencyMs: 100
idempotent is different from read_only. A read-only tool is normally
idempotent, but some write tools can also be idempotent when they use an
idempotency key. Automatic retries should require idempotent: true and an
explicit retry policy. Destructive tools should not be retried automatically
unless the operation is proven idempotent and protected by an idempotency key.
cacheTtlSeconds should only apply to results that are safe to cache. The
current Rust MCP router implements the bounded tools/list visibility cache;
it does not yet implement a gateway tools/call response cache. Before adding
one, preserve this mandatory ordering:
authenticate and run req-acc
-> compute a raw-result cache key
-> load the normalized pre-res-fil MCP result, or call the backend on miss
-> store only the normalized pre-res-fil result when eligible
-> run the current caller's res-fil rules on every hit and miss
-> return the caller-specific filtered result
The shared call cache must never store a post-res-fil result. Reusing a
caller-specific filtered value would couple cached data to an earlier caller’s
claims and policy. Conversely, the pre-filter value can contain more data than
any caller may see, so the cache itself is a sensitive tenant boundary: encrypt
or keep it process-local as policy requires, bound its entries and TTL, exclude
it from logs, and never expose cache inspection to tenant code.
The raw-result key includes tool/config digest, resolved backend and
environment, method, normalized arguments/body, representation-affecting
headers, authenticated tenant/principal by default, and every request dimension
that can change the downstream response. Cross-principal reuse is allowed only
when policy explicitly proves the backend result is principal-invariant; it
does not follow merely from read_only: true. Access control, revocation, and
res-fil are always reevaluated on a cache hit. Do not cache denials, filter
errors, partial/streaming results, secret-bearing results, or unknown outcomes.
There are two distinct cache types:
| Cache | Purpose | Recommended owner |
|---|---|---|
tools/list visibility cache | Reuse the filtered list of visible tool names for the same principal, claims, headers, and query. | Gateway MCP router. |
tools/call raw-result cache | Reuse a normalized pre-res-fil result for an identical safe backend request; run caller-specific filtering on every hit. | Backend service first; gateway only when explicitly configured. |
The tools/list cache is a runtime optimization for discovery. It does not
cache business data and does not change tools/call authorization. It should be
enabled when toolsListAccessControl is enabled and bounded by a maximum entry
count.
Gateway-level tools/call response caching should be opt-in and conservative.
Start with backend-owned caching for expensive read APIs. Add gateway response
caching only for tools with read_only: true, idempotent: true, a positive
cacheTtlSeconds, a pre-filter storage boundary, and a cache key that includes
all tenant/principal, argument, backend, environment, request-header, and tool
configuration dimensions that can change the raw result. If the gateway cannot
prove those conditions or cannot rerun res-fil on a hit, caching stays
disabled for that tool.
Before enabling a gateway call cache, integration tests use callers with
different claims and row/column filters against the same raw backend result.
They prove that req-acc and current res-fil run on every hit, caller outputs
remain distinct, revocation or policy reload takes effect without waiting for
the raw-result TTL, backend-varying identity dimensions prevent unsafe hits,
filter errors are not cached, and pre-filter bytes never appear in logs or
cache diagnostics.
Rate-limit and cost metadata should influence ranking and diagnostics. It can also prevent an agent from repeatedly selecting a tool whose backend quota is already exhausted.
Service Resolution
Routing should avoid model-supplied URLs. The selected tool already carries the deployment routing data.
Resolution order:
- Use
targetHostwhen explicitly configured. - Use direct-registry when a matching static URL is configured.
- Use service discovery through
serviceId,envTag, andprotocol.
The gateway must validate protocol compatibility when direct-registry is used.
For example, a tool configured for protocol: http should not silently route to
an incompatible backend entry.
targetHost is administrative configuration, not model input. Automated imports
must treat targetHost as untrusted until the owning control plane validates
and approves it. Validation should include allowed schemes, allowed hostnames or
service identities, optional CIDR allowlists, DNS and redirect handling, and
environment ownership. This prevents a compromised catalog import from turning
the gateway into an SSRF path to metadata services, loopback addresses, or
internal control-plane endpoints.
DNS and CIDR checks must be enforced on the actual resolved address used by the gateway connector, not only on the URL string. This prevents DNS rebinding from turning an approved-looking hostname into a loopback, link-local, private, or metadata-service address at connection time. Redirect targets should go through the same validation.
Access Control
MCP tool metadata should complement, not replace, access-control policy.
The gateway applies the shared access-control runtime around tools/call:
req-accruns before the downstream tool is invoked or a call-result cache entry is used.res-filruns after the downstream or cached pre-filter result is converted to an MCP result, on every cache hit and miss.- A future gateway call cache stores only the normalized pre-
res-filresult; a shared cache never stores caller-filtered output.
Use metadata as follows:
| Metadata | Access-control use |
|---|---|
endpoint | Stable key for endpoint rules. |
endpointId | Stable audit and governance identifier. |
sensitivityTier | Disclosure and policy input. |
read_only | Safe-tool classification and policy input. |
destructive | Approval or denial input. |
humanApprovalRequired | Workflow or approval routing input. |
sourceProtocol | Policy and diagnostics dimension. |
Do not rely on model instructions for sensitive operations. If
destructive: true or humanApprovalRequired: true, enforcement should be in
policy or workflow, not only in the prompt.
See MCP Tools Access Control for invocation authorization and response filtering.
Metadata Storage
Use one canonical metadata object and derive indexed columns from it.
Recommended storage:
- Store
toolMetadataandinputSchemaas JSON or JSONB in catalog tables. - Store flattened fields such as
routingDomain,semanticNamespace,sourceProtocol,sensitivityTier,semanticWeight,readOnly, anddestructiveas indexed projection columns when search needs them. - Regenerate flattened projections when the canonical JSON changes.
- Store
endpointIdas the stable identity for audit, scoring feedback, and catalog synchronization.
This avoids drift where toolMetadata.routing.domain says one thing and a
flattened routingDomain column says another.
The gateway should not read portal catalog tables directly. Light Portal or the
control plane owns catalog authoring, normalization, approval, and projection.
It publishes a flattened runtime config, such as mcp-router.yml or
config-cache content, to gateway instances. The gateway then executes from that
approved runtime config and live service discovery state.
Catalog import must normalize compatibility fields before publishing gateway
config. If safety.read_only and top-level read_only disagree, or
safety.destructive and top-level destructive disagree, the import should
fail or rewrite the compatibility fields from the canonical safety object.
Agents and gateways should never observe conflicting safety values.
Metadata Contract
The recommended metadata shape is:
toolMetadata:
routing:
domain: Customers
semanticNamespace: API0004
semanticDescription: Get customer preferences
semanticKeywords:
- Customers
- getCustomerPreferences
- preferences
semanticWeight: 1.0
sensitivityTier: internal
sourceProtocol: openapi
parameters:
customerId: path
channel: query
safety:
read_only: true
idempotent: true
destructive: false
humanApprovalRequired: false
runtime:
cacheTtlSeconds: 60
costTier: low
estimatedLatencyMs: 100
lifecycle:
version: 1.0.0
status: active
read_only: true
destructive: false
The duplicated top-level read_only and destructive fields are compatibility
fields. New code should prefer safety.read_only, safety.destructive, and
safety.humanApprovalRequired, then fall back to the top-level fields.
Recommended field semantics:
| Field | Required | Semantics |
|---|---|---|
routing.domain | Recommended | Business capability group. |
routing.semanticNamespace | Recommended | Catalog/API namespace for filtering and grouping. |
routing.semanticDescription | Recommended | Agent-facing capability summary. |
routing.semanticKeywords | Recommended | Search keywords and aliases. |
routing.semanticWeight | Optional | Catalog/light-agent ranking multiplier. Default 1.0, with a 0.1 lower bound in current light-agent selection. It does not affect gateway tools/list. |
routing.sensitivityTier | Recommended | Disclosure and governance tier. |
routing.sourceProtocol | Recommended | Source protocol, such as mcp, openapi, http, or lightapi. |
routing.parameters | Required for non-trivial OpenAPI tools | Argument location mapping. |
safety.read_only | Recommended | True when the tool does not mutate state. |
safety.idempotent | Recommended | True when identical calls can be safely retried. |
safety.destructive | Recommended | True when the tool can delete, reset, revoke, overwrite, or cause irreversible effects. |
safety.humanApprovalRequired | Recommended | True when a workflow or approval step must precede execution. |
runtime.cacheTtlSeconds | Optional | TTL hint for safe raw backend results. It never authorizes caching a post-res-fil caller view; gateway call caching remains disabled until the pre-filter contract is implemented. |
runtime.retry | Optional | Retry policy, only honored for idempotent calls. |
runtime.rateLimit | Optional | Rate-limit grouping and per-call cost units. |
runtime.costTier | Optional | Relative execution cost such as low, medium, or high. |
runtime.estimatedLatencyMs | Optional | Expected latency used for ranking and diagnostics. |
lifecycle.version | Recommended | Tool contract version visible to agents and catalogs. |
lifecycle.status | Recommended | Lifecycle state such as active, deprecated, or retired. |
Use normalized sensitivity reference values such as public, internal,
confidential, and restricted. Older imported values such as Internal-Only
should be normalized during catalog import or edited through the App/GenAI/Tool
dropdowns.
The current Portal reference tables for this metadata are:
| Reference table | Metadata field | Values |
|---|---|---|
sensitivity_tier | toolMetadata.routing.sensitivityTier and sensitivity_tier projection | public, internal, confidential, restricted |
source_protocol | toolMetadata.routing.sourceProtocol and source_protocol projection | openapi, mcp, lightapi, http |
lifecycle_status | toolMetadata.lifecycle.status and lifecycle_status projection | active, deprecated, retired |
parameter_location | values inside toolMetadata.routing.parameters | path, query, header, cookie, body |
cost_tier | toolMetadata.runtime.costTier and cost_tier projection | low, medium, high |
Use openapi when the tool contract is generated from an OpenAPI document. Use
http for manually configured HTTP-family tools without an OpenAPI contract,
including REST-style endpoints or future HTTP transports such as
GraphQL-over-HTTP and gRPC-over-HTTP.
Diagnostics
Operators need to understand why an agent saw or did not see a tool and where a selected call was routed.
Diagnostics should include:
| Event | Useful fields |
|---|---|
| Catalog search | query, selected tool names, scores, score reasons, catalog hash, catalog version. |
| Disclosure filtering | hidden tool names, policy reason, sensitivity tier, destructive flag, approval requirement. |
| Gateway list check | catalog tools missing from gateway, extra gateway tools, gateway list error. |
| Tool call | tool name, endpoint, endpointId, serviceId, envTag, sourceProtocol, policy outcome, correlation ID. |
| Backend routing | target source, selected URL without secrets, discovery node, direct-registry match. |
| Runtime policy | retry attempt, idempotent flag, cache hit or miss, rate-limit decision, cost tier. |
| Response filtering | endpoint, filter rule IDs, filtered result status, policy outcome. |
The agent diagnostics endpoint should compare assigned catalog tools with live
gateway tools/list so operators can see catalog/runtime drift.
The gateway should avoid logging tool arguments in full. When arguments are
logged for debugging, masking should follow the inputSchema and metadata
sensitivity signals.
Distributed tracing should carry selected metadata as span attributes. Useful OpenTelemetry attributes include:
| Attribute | Source |
|---|---|
mcp.tool.name | Tool name. |
mcp.tool.endpoint_id | endpointId. |
mcp.tool.endpoint | endpoint. |
mcp.tool.domain | toolMetadata.routing.domain. |
mcp.tool.namespace | toolMetadata.routing.semanticNamespace. |
mcp.tool.source_protocol | toolMetadata.routing.sourceProtocol. |
mcp.tool.read_only | toolMetadata.safety.read_only. |
mcp.tool.idempotent | toolMetadata.safety.idempotent. |
mcp.tool.cost_tier | toolMetadata.runtime.costTier. |
These attributes let operators group latency, errors, policy denials, and rate limits by domain, namespace, protocol, and tool contract instead of only by URL.
Advanced Metadata Usage
The same metadata contract can support features beyond search and routing.
Dry Run And Mocking
Development and workflow validation can use sandbox metadata:
toolMetadata:
sandbox:
enabled: true
mode: mock
mockResponse:
customerId: CUST-1001
preferences:
channel: portal
Mocking must be opt-in and environment-scoped. Production gateways should not return mock responses unless a deployment explicitly enables sandbox mode for a tool, environment, or test principal.
UI Rendering Hints
Some tool results are easier to inspect as structured UI components:
toolMetadata:
ui:
component: customer-profile-card
resultShape: customerProfile
UI metadata should be treated as a rendering hint, not as trusted executable frontend code. The frontend should map known component names to local UI components and ignore unknown values.
Related Tools
Catalog search can use related-tool hints to pre-warm or prioritize likely next schemas:
toolMetadata:
relatedTools:
- demo_customer_profile_api_get_customer_preferences
- demo_offer_decision_api_search_offers
Related tools should not bypass assignment, visibility, or their
placement-specific live-availability intersection. Gateway-placed entries still
require gateway tools/list; runner entries require the lease/runtime/local
manifest checks. Related links only affect ranking and prefetch.
Sub-Agent Orchestration
In a multi-agent deployment, a tool may require skills owned by a worker agent:
toolMetadata:
orchestration:
requiredSkills:
- data_analysis
- python_execution
preferredAgentId: analytics-worker
The supervisor can use these hints to delegate the user task or to avoid disclosing a tool to an agent that cannot safely execute the surrounding work.
This is orchestration metadata, not a source protocol. Do not use
sourceProtocol: agent. Keep sourceProtocol for concrete protocol or
contract sources such as mcp, openapi, http, and lightapi.
Do not add orchestration reference tables in the first rollout. If sub-agent delegation becomes a product feature, reuse the Light Portal agent and skill model:
- Agent capabilities are the skills assigned to each agent.
requiredSkillsshould be selected from the existing skill catalog.preferredAgentId, if present, should refer to an agent in the managed agent registry.- The catalog or supervisor should validate that the preferred agent has the required skills.
Evaluation Feedback
Metadata should also support closed-loop improvement.
Track these signals by endpointId and tool name:
- Search query text or normalized intent.
- Tool rank and selected rank position.
- Whether the model called the tool.
- Whether the call succeeded.
- Whether the user accepted the result.
- Whether policy denied the call.
- Whether retries, cache hits, rate limits, or cost budgets affected the call.
- Whether schema validation or required arguments failed.
This feedback can tune semanticKeywords, semanticDescription, and
semanticWeight without changing the backend API contract.
Recommended Rollout
Implement metadata usage incrementally:
- Normalize imported
inputSchemaandtoolMetadatato structured JSON. - Validate administrative routing fields such as
targetHostand normalize compatibility safety fields. - Project searchable fields into catalog columns or search documents.
- Add keyword search over
name,endpointName,description,semanticDescription,semanticKeywords, domain, namespace, and schema property names. - Apply disclosure filters for assignment, environment, lifecycle, sensitivity, destructive tools, and approval-required tools.
- Add server-owned tool placement and partition selected candidates into
gateway, runner, workflow, and fixed-service sets. Intersect only gateway
candidates with live gateway
tools/list; require runner candidates to match the lease/runtime/local capability manifests. - Add diagnostics for selected, hidden, missing, conflicting, and placement-incompatible tools.
- Move the existing path-placeholder/mapping checks into importer,
publication, and gateway startup/reload validation while retaining runtime
rejection as defense in depth. In
frameworks/light-pingora/src/mcp.rs, refactor the existingopenapi_path_placeholdersand mapping checks into one shared validator called by bothvalidate_configand request construction so startup and call-time behavior cannot drift. - Add retry, rate-limit, and OpenTelemetry attributes after the core
disclosure path is stable. If gateway
tools/callcaching is later added, implement a bounded pre-res-filcache and rerunreq-acc/res-filfor every caller and hit. In the currenthandle_tool_callpipeline, the cache may replace backend execution after authorization, but it must feed the existingfilter_mcp_responsecall rather than bypass or follow it. - Add semantic or hybrid search after keyword behavior is proven.
- Feed evaluation results back into keywords and semantic weights.
Do not start by changing gateway tools/call. The gateway execution path is
already the right boundary. The first improvement should be better catalog
search and per-turn tool disclosure.
Semantic vector search is needed for large catalogs, but it should be optional. Keep keyword plus structured filtering as the baseline implementation, then add hybrid search as an enhancement for deployments that have enough tools to justify the extra index and operations cost.
Example End-To-End Flow
User prompt:
Show customer CUST-1001 preferences and find available travel offers.
Catalog search:
- Matches
customer,preferences, andCUST-1001against the customer profile tool metadata and schema. - Matches
travelandoffersagainst the offer search tool metadata and schema. - Filters out destructive or approval-required tools.
- Classifies both selected entries as gateway-placed and intersects their
names with gateway
tools/list.
Model tool disclosure:
demo_customer_profile_api_get_customer_preferences
demo_offer_decision_api_search_offers
Execution:
- The model calls
demo_customer_profile_api_get_customer_preferences. - The gateway maps
customerIdto the path andchannelto the query string. - The gateway runs
req-acc. - The gateway invokes the downstream customer profile API.
- The gateway runs
res-filif configured. - The model receives the MCP result.
- The model calls
demo_offer_decision_api_search_offersif more data is needed.
The model never receives backend URLs, discovery nodes, or direct-registry details. It only receives the selected tool schemas.
Resolved Guidance
The recommended default decisions are:
| Topic | Decision |
|---|---|
| Semantic search | Support optional semantic or hybrid vector search. Keyword plus structured filters remain the required baseline. |
| Sensitivity tier | Set a default when API details create endpoints. Store allowed values in the sensitivity_tier reference table and expose normalized dropdowns in the App/GenAI/Tool pages. |
| Destructive tools | Require workflow-backed or approval-backed execution for destructive tools. Do not expose them as direct model-callable tools unless approval is configured. |
| Tool availability | Partition by server-owned execution placement. Gateway entries intersect portal policy and live gateway tools/list; runner entries intersect execution policy, lease allowlist, approved runtime manifest, and live local enumeration. Union only independently authorized, collision-free definitions. |
| Semantic weight | Populate semanticWeight when the endpoint is created, then allow authorized updates from the App/GenAI/Tool page. |
| Caching | Use gateway caching first for tools/list visibility. Keep tools/call caching backend-owned by default. A future gateway call cache stores only normalized pre-res-fil results, reruns req-acc and caller-specific res-fil on every hit, and remains disabled unless the raw-result key and sensitive cache boundary are proven safe. |
| Path parameters | Fail closed at import/publication and gateway startup/reload when a path-template mapping is missing or inconsistent; retain call-time rejection as defense in depth. OpenAPI import generates toolMetadata.routing.parameters, and method fallback applies only to paths without placeholders. |
If a future deployment needs path-template inference, add an explicit opt-in
field such as routing.parameterInference: pathTemplate. The default should
remain explicit mapping because it is safer, easier to audit, and consistent
with OpenAPI parameter locations.
Workflow-Backed MCP Tools
Status: Development implementation; runtime qualification incomplete
The active publication and invocation contract is
Workflow Invoke And Tool Binding Publication.
Published workflow-backed Tools are synchronous only. Later sections in this
document describing asynchronous workflow-backed Tool publication are design
history; asynchronous root Workflow starts use workflow_start.
This repository is still in development. The workflow-backed MCP path has not been exercised against a live workflow deployment, so none of the phase gate scripts or unit-test results constitute production qualification. The current implementation intentionally targets one clean contract; migration adapters, legacy binding formats, and backward-compatibility rollout procedures are out of scope until the runtime behavior is proven.
Implementation And Qualification Status
| Area | Implementation | Qualification |
|---|---|---|
| Phase 0 contracts, canonicalization, and threat-model fixtures | Implemented | Unit/fixture and disposable PostgreSQL contract checks pass; latency evidence not run. |
| Phase 1 synchronous gateway/workflow path | Implemented | Rust tests pass; no deployed end-to-end workflow or concurrency/fairness evidence. |
| Phase 2 asynchronous, effect, cancellation, and compensation path | Implemented | Component tests pass; no live side-effect or recovery exercise. |
| Phase 3 AI-assisted draft authoring | Implemented | Java/UI checks pass; no production model or reviewer workflow qualification. |
| Phase 4 optional skill binding | Implemented | Component and disposable PostgreSQL constraint checks pass; no live agent/catalog exercise. |
| Runtime promotion | Disabled | Remains disabled until the numeric Phase 0/1 qualification evidence is recorded. |
Requirements are indexed by their owning sections: binding and runtime fields under Tool Contract and Invocation Contract; error classes under Failure Mapping; security controls under Authorization, Delegation, and Destination Safety; and verification requirements under each phase’s Exit gates.
This document defines how light-gateway should expose an orchestration as an
ordinary MCP tool while light-workflow owns the durable multi-step execution.
The design lets existing MCP-capable agents consume a higher-level business
capability without adding Light-Portal skill support or reproducing API
sequencing inside the agent.
The core principle is:
The gateway exposes and governs the tool contract; the workflow runtime executes the orchestration.
Problem
The MCP router can currently expose a backend MCP operation or translate an MCP
tools/call into one backend API request. That works well when a backend
endpoint is already meaningful to an agent.
Many enterprise APIs are more granular. A useful business operation may need to:
- load data from several APIs;
- transform and join their responses;
- apply business rules or conditional branches;
- call another API with the derived input;
- normalize transport-specific responses into one stable result;
- retry transient failures or compensate for partial side effects; and
- pause for approval or continue asynchronously.
Making the agent issue each low-level call is inefficient and exposes internal API structure to every agent implementation. A Portal skill can guide an agent through the process, but existing customer agents may understand MCP without understanding Light-Portal skills or skill-to-workflow bindings.
The same high-level capability therefore needs to be available directly from
the gateway’s MCP tools/list and tools/call surface.
Goals
- Expose an orchestration as a normal, schema-bound MCP tool.
- Avoid requiring existing agents to adopt Portal skills.
- Keep one canonical executable workflow definition.
- Keep authorization, input validation, response filtering, and output validation at the gateway boundary.
- Keep sequencing, branching, transformation, retries, durable state, human
tasks, and compensation in
light-workflow. - Support both bounded synchronous tools and explicit asynchronous tools.
- Let users author the workflow manually or ask AI to generate a reviewable
draft in
portal-view. - Publish immutable workflow and schema references to the gateway through the existing configuration control plane.
- Preserve tenant isolation, stable tool identity, audit correlation, and least-privilege delegation for every nested call.
Non-Goals
- Do not embed a general-purpose workflow engine in
light-gateway. - Do not execute user-authored JavaScript, Python, or shell code in the gateway process.
- Do not let model-generated text choose arbitrary backend URLs, credentials, service IDs, or workflow definitions at runtime.
- Do not make every agent turn a workflow.
- Do not require a skill assignment before a workflow-backed MCP tool can be invoked.
- Do not copy the workflow DSL into every gateway instance as the executable source of truth.
- Do not make the gateway read Portal or workflow database tables directly.
Decision
Introduce a workflow-backed MCP tool as a third gateway execution type:
Existing MCP agent
-> light-gateway tools/list
-> light-gateway tools/call
-> gateway authorization and input validation
-> workflow invocation API
-> light-workflow durable execution
-> gateway response filtering and output validation
-> MCP tool result
The tool looks like any other MCP tool to the caller. Its implementation is a versioned workflow reference rather than one HTTP endpoint or downstream MCP operation.
The gateway contains a small workflow dispatch adapter. It does not interpret workflow steps, maintain workflow state, run compensations, or execute transform expressions.
Why Not Orchestrate Inside The Gateway
Gateway-native orchestration may appear to reduce one network hop, but it would create a second orchestration runtime in the data plane. That runtime would need independent solutions for:
- durable state across gateway restarts;
- idempotency and duplicate requests;
- retries, backoff, and per-step deadlines;
- fan-out, joins, and partial failures;
- compensation after side effects;
- human approval and long waits;
- workflow version snapshots;
- cancellation and abandoned callers;
- cycle detection and bounded recursion; and
- per-step audit, metrics, and traces.
Those are workflow-runtime concerns. Keeping them in light-workflow also
prevents a slow or waiting business process from consuming gateway execution
state.
A separate composition service should be considered only if production measurements prove that the durable workflow path cannot meet a required interactive latency target. If introduced, it should execute the same compiled and governed workflow representation instead of creating another authoring language.
Relationship To Skills
A skill and a workflow-backed tool serve different purposes:
| Object | Responsibility |
|---|---|
| Skill | Agent-facing guidance, examples, discovery hints, and optional progressive disclosure. |
| MCP tool | Stable executable input/output contract visible through tools/list. |
| Workflow | Canonical orchestration, transformations, branches, and durable execution. |
The existing Skill Workflow Orchestration
design links a skill to a canonical workflow through skill_workflow_t. This
design adds a second, independent exposure for the same workflow:
workflow definition
|-- optional skill_workflow_t link for skill-aware agents
`-- workflow tool binding for all MCP-capable agents
An agent that supports Portal skills can receive richer instructions and examples. An existing agent can discover and call the composite MCP tool with no skill integration. An agent with a static tool allowlist must add the new tool name or refresh that allowlist, but it does not need a new orchestration framework.
Current Foundation And Gaps
The existing platform already provides most control-plane and runtime pieces:
tool_thas a stable tool reference, model alias, schema digest, dispatch policy reference, andexecution_placement = workflow.wf_definition_tstores the canonical versioned workflow YAML.skill_workflow_tcan optionally connect a skill to a workflow.portal-viewhas a workflow YAML editor, outline/graph support, client and server validation, test input, workflow start, and runtime-state inspection.light-workflowsnapshots the definition digest and resolved execution policy when it consumesWorkflowStartedEvent.light-workflowcan execute HTTP and MCP calls and can useset,switch,assert, and output exports for sequential compositions.light-gatewayalready performs MCP tool authorization, input-schema validation, backend dispatch, response filtering, output-schema validation, resource limits, and audit logging.
The integration is not turnkey yet:
- the gateway runtime accepts only HTTP/OpenAPI and MCP tool execution types;
- workflow start currently enters through
workflow-command, and runtime status is read from workflow query projections; light-workflowdoes not yet expose a stable start/wait/status/result/cancel service boundary for gateway use;- the current runtime expression evaluator supports only a limited path, interpolation, literal, and comparison subset rather than the production CEL expression contract defined below;
- each
light-workflowservice instance currently runs one serial host-task executor over a global cross-tenant claim ordered by priority and age, sleeps for 500 ms after an empty claim, and reclaims stale Boolean locks after five minutes without a fencing token; and - generic fork/join and retry behavior must be completed before advertising broad production orchestration semantics.
These gaps should be closed in the workflow and gateway runtimes rather than worked around with executable logic in gateway configuration.
Component Responsibilities
Light Portal
The Portal control plane owns:
- workflow and composite-tool authoring;
- stable identities and versions;
- schema extraction and validation;
- workflow-to-tool bindings;
- dependency resolution and cycle checks;
- policy and safety metadata;
- test fixtures and promotion evidence;
- review, approval, publication, rollback, and retirement; and
- projection of runtime-ready
mcp-router.toolsconfiguration.
The control plane must normalize flexible UI input into the strict gateway runtime contract before persistence. The gateway must not infer missing workflow identity, safety policy, or schema bindings from model input.
Portal View
portal-view owns the manual and AI-assisted authoring experience. It does not
execute production workflows or publish AI output without validation and
approval.
Light Gateway
The MCP router owns:
tools/listexposure and tools-list access control;- tool name to stable workflow-tool binding resolution;
- caller authentication and composite-tool authorization;
- input-schema validation and argument masking;
- bounded workflow dispatch;
- correlation, delegation, and idempotency context;
- concurrency, payload, and deadline limits;
- mapping workflow terminal state to an MCP result;
- response filtering and output-schema validation; and
- gateway-level audit and diagnostics.
The gateway does not parse or execute workflow tasks.
Light Workflow
The workflow service owns:
- validating the requested workflow reference and expected digest;
- durable instance creation and idempotent start;
- definition and execution-policy snapshots;
- task sequencing, branching, transformations, retries, and joins;
- HTTP, MCP, rule, agent, human, and runner task execution according to policy;
- workflow deadlines, cancellation, and compensation;
- public-result construction and output-schema validation;
- instance, task, event, and audit state; and
- start, wait, status, result, and cancel APIs.
Control-Plane Data Model
Keep the workflow YAML canonical in wf_definition_t.definition. Use
tool_t for the agent-facing tool identity and set:
stable_tool_ref = immutable logical tool identity
execution_placement = workflow
model_alias = gateway-facing MCP tool name
schema_digest = digest of the published input/output contract
dispatch_policy_ref = reference to the approved dispatch policy
Add a dedicated workflow-to-tool binding rather than requiring
skill_workflow_t. A proposed workflow_tool_binding_t contains:
| Field | Purpose |
|---|---|
host_id | Tenant boundary. |
tool_id | Agent-facing tool identity in tool_t. |
wf_def_id | Canonical workflow definition. |
definition_digest | Exact published workflow snapshot expected by the gateway. |
invocation_mode | sync or async. |
sync_wait_ms | Maximum gateway wait for a synchronous result. |
total_deadline_ms | Maximum end-to-end workflow deadline. |
execution_class | Default scheduler class for direct/root invocation; the initial synchronous profile uses interactive, while nested invocation inherits its outer class. |
result_text_mode | Non-executable MCP text rendering mode: compact-json or schema-backed summary. |
idempotency_policy | Required key, derived business key, or read-only handling. |
delegation_policy | Allowed nested tool references, audiences, and maximum depth. |
response_policy_digest | Classification and filtering policy snapshot used for later result reads. |
aggregate_version | Optimistic concurrency and event projection version. |
active | Publication lifecycle state. |
The binding must reference one immutable workflow version and digest. Editing a published workflow creates or promotes a new version; it must not silently change the implementation behind an existing digest.
The optional relationships are:
skill_t
-> skill_tool_t -> workflow-backed tool_t
-> skill_workflow_t -> wf_definition_t
tool_t
-> workflow_tool_binding_t -> wf_definition_t
The two paths may point to the same workflow, but neither path copies the workflow DSL.
Store each resolved nested dependency in a separate
workflow_tool_dependency_t projection keyed by the outer binding and nested
stable tool reference. It records the nested version, contract digest,
compatibility policy, logical authorization tool name, endpoint key, and
authorization-policy reference. A reverse index on the nested stable reference
is required so publication can report every affected composite tool and
require revalidation or reapproval before an incompatible nested contract is
promoted or the nested tool is retired.
Projected Gateway Configuration
Keep apiType as the backend transport dimension (http or mcp). Select a
workflow-backed tool through the existing catalog concept
executionPlacement: workflow; do not add workflow as a third transport.
The gateway runtime therefore dispatches by execution placement first and uses
apiType only for gateway-executed backend calls.
Example projected tool for the later write-capable profile; the Phase 1 variant must be read-only:
- name: recommend_customer_offer
description: Recommend and record the best eligible customer offer.
method: call
executionPlacement: workflow
endpoint: recommend_customer_offer@call
inputSchema:
type: object
additionalProperties: false
required:
- requestId
- customerId
- channel
properties:
requestId:
type: string
description: Stable business request identifier used for idempotency.
customerId:
type: string
channel:
type: string
outputSchema:
type: object
additionalProperties: false
required:
- status
- customerId
properties:
status:
type: string
enum:
- APPROVED
- REJECTED
- NO_CONSENT
- NO_ELIGIBLE_OFFER
customerId:
type: string
selectedOfferId:
type: string
decisionId:
type: string
workflow:
wfDefId: 2695cdee-cb82-4b34-a2d8-f69093c733e3
version: 1.0.0
definitionDigest: sha256:0123456789abcdef
mode: sync
executionClass: interactive
waitTimeoutMs: 20000
totalDeadlineMs: 30000
maximumDefinitionTasks: 8
maximumExecutionAttempts: 8
maximumNestedCalls: 8
maximumParallelism: 1
maximumDelegationDepth: 1
idempotencyInput: requestId
resultReplayMs: 300000
resultTextMode: compact-json
toolMetadata:
routing:
domain: Offers
semanticNamespace: customer-offers
semanticDescription: Recommend an eligible personalized customer offer.
semanticKeywords:
- recommend offer
- personalized offer
- customer eligibility
sensitivityTier: confidential
safety:
read_only: false
idempotent: true
destructive: false
humanApprovalRequired: false
runtime:
costTier: medium
estimatedLatencyMs: 5000
lifecycle:
version: 1.0.0
status: active
Replace the placeholder definitionDigest with the canonical digest of the
exact workflow definition being pinned:
cargo run -p light-workflow --example workflow_definition_digest -- <definition.yaml>
The configuration contains only an approved reference, contract, and bounds.
It does not contain executable scripts or caller-selectable destinations.
The five-minute resultReplayMs in this later write-capable example is an
illustrative business-idempotency policy, not a platform or read-only default.
Read-only tools normally publish a much shorter completed-result freshness
window, including zero when only in-flight deduplication is wanted.
The metadata ownership and compact tools/list rules in
MCP Tool Metadata Usage continue to apply.
The existing metadata contract intentionally uses read_only, idempotent,
and destructive together with humanApprovalRequired; projection must retain
those canonical spellings rather than normalize them ad hoc.
Workflow Start and Invocation API
The current workflow-backed Tool contract is described in Workflow Invoke And Tool Binding Publication. Workflow-backed Tools are synchronous only. Use asynchronous
workflow_startfor editor, scheduler, and other root starts.
Workflow-backed Tool calls enter through the Gateway Tool ACL and its internal
workflow_invoke call. Direct clients cannot list or call workflow_invoke.
The native workflow_start operation remains asynchronous and serves other
root starts; it does not admit a workflow-backed Tool binding.
The Gateway uses these internal operations only after start, to read or wait for the process and to cancel it:
GET /v1/workflow-invocations/{workflowInstanceId}
GET /v1/workflow-invocations/{workflowInstanceId}/result
POST /v1/workflow-invocations/{workflowInstanceId}/wait
DELETE /v1/workflow-invocations/{workflowInstanceId}
The native workflow_start receipt identifies the durably created process
and workflow instance. The lifecycle endpoints must scope reads, waits and
cancellation to the same authenticated user identity.
The invocation and process records are created by the native start handler.
Any later audit or projection event records that accepted start; the legacy
WorkflowStartedEvent consumer ignores start events and cannot create a second
process.
Event consumers still require poison-event isolation for non-start
projections. Deterministic parse, schema, and contract failures are poison and
are quarantined without immediate same-transaction retries. Retryable database
and transport failures roll back the claimed batch and use the consumer’s
outer reconnect backoff; they must never consume poison attempts or block an
aggregate in quarantine. A permanently invalid event is recorded in a
tenant/partition-scoped quarantine with its offset, error, and payload digest
in the same transaction that advances the consumer offset. The
quarantine or deferred-event store also retains the encrypted replayable
payload, or a durable immutable payload reference, plus every source offset and
aggregate version needed to reconstruct order. If replay depends on
outbox_message_t, outbox retention must exceed the maximum quarantine dwell
and unresolved holds must block purging those offsets.
The affected aggregate is marked blocked so later events for that aggregate are deferred or quarantined in order, while unrelated aggregates continue. The consumer raises an operational alert and supports an audited, ordered repair- and-replay operation. A malformed unrelated event must not block workflow start, status, wait, or result retrieval.
Invocation State
Use a small stable state model at the API boundary:
ACCEPTED
RUNNING
WAITING
COMPLETED
FAILED
CANCELLED
The response includes the workflow instance ID, current state, definition digest, timestamps, retryability, and a public result or normalized error when terminal. It never returns the full internal context, credentials, or unfiltered intermediate task outputs.
POST /{id}/wait is a bounded, resumable long poll over durable instance
state, not an in-memory gateway subscription. Any authorized gateway node can
resume it after a restart. Multiple waiters may observe the same instance and
must not change its state. The effective server wait is the minimum of the
published sync_wait_ms, the service-side long-poll cap, and the remaining
workflow deadline. A timeout returns the latest state and instance ID; it does
not imply cancellation or failed acceptance.
The invocation service bounds dedicated PostgreSQL LISTEN connections with
WORKFLOW_WAIT_LISTENER_CONNECTIONS (default 8). Additional waiters use short
durable status polling, so synchronous permit capacity cannot translate into
an unbounded database-connection count.
Public Result
light-workflow must produce the public result only from the canonical
workflow output definition. Executable result expressions do not belong in a
binding row. The public result is validated against the workflow output schema
before the instance is reported as COMPLETED.
The gateway then applies its current response-filter and output-validation pipeline. A workflow must not return raw backend transport envelopes directly to the agent.
MCP Result Envelope
For Phase 1, workflow-backed tools must publish an object-root output schema.
After workflow validation and gateway response filtering, the gateway emits
the filtered object unchanged as structuredContent; this is the authoritative
machine-readable result and the value validated against outputSchema.
Array or scalar public outputs must be modeled explicitly inside an object,
such as { "items": [...] }, rather than receiving an implicit runtime wrapper.
The published contract selects one non-executable text rendering mode:
compact-jsonis the compatibility default and emits one text content block containing a compact serialization of the filteredstructuredContent.summaryrequires a schema-declared, requiredsummarystring in the same public result and emits exactly that field as the text content block.
The summary mode is preferred for large structured results because it avoids
duplicating the entire result in the model-visible text channel. Rendering can
never read hidden workflow context or introduce data absent from the filtered
structuredContent. Text and structured output limits are publication-time
and runtime gates.
A technical failure sets isError: true, omits success
structuredContent, and returns a concise sanitized text block containing the
stable error code and workflow instance ID when allocated. The same code,
instance ID, state, and retryability are carried in bounded gateway _meta for
programmatic clients. Business outcomes remain successful structured results.
Gateway Execution Flow
For tools/call, the gateway performs:
- Resolve the gateway-facing name to one immutable tool and workflow binding.
- Apply tools-call authorization using the composite tool endpoint key.
- Validate arguments against the published input schema.
- Mask or transform arguments according to approved request policy.
- Acquire a synchronous-wait permit from depth zero or the signed delegated depth pool when required.
- Derive the idempotency key and construct correlation, deadline, invocation- budget, effective execution-class, and delegation context.
- Start the workflow using the pinned definition version and digest.
- Wait only when the tool is published as synchronous.
- Map workflow state and public output into an MCP tool result.
- Apply response filtering for the initiating caller.
- Validate successful structured content against
outputSchema. - Emit gateway and workflow correlation/audit attributes.
If the invocation service reports a different workflow digest, tenant, stable tool reference, or schema binding, the gateway fails closed.
Synchronous Tools
Synchronous tools preserve the simplest compatibility contract: an existing
agent issues one tools/call and receives the business result.
The initial synchronous profile should allow only bounded, headless workflows. Recommended starting limits are:
| Limit | Initial value |
|---|---|
| Static definition tasks | 8 |
| Runtime task attempts, including retries | 8 in Phase 1 |
| Nested tool/API calls | 8 |
| Parallel branches | 1 until fork/join is qualified |
| Nested workflow-tool depth | 1 |
| Gateway wait | 20 seconds |
| Total workflow deadline | 30 seconds |
These are starting defaults, not protocol constants. They should be environment-configurable and publication-validated.
Static definition size and runtime execution consumption are separate budgets. Publication counts all reachable tasks, including nested composite dependencies. At invocation time the gateway creates one structured budget covering remaining task attempts, nested calls, delegation depth, parallel branches, request/intermediate/result bytes, wall-clock deadline, and optional cost. The signed delegation token identifies the invocation and carries the immutable budget ceilings, but it is not the mutable counter. The invocation service maintains one durable, atomic budget ledger shared by every task, retry, and parallel branch. Before dispatch, a worker conditionally reserves attempts, calls, bytes, and cost from that ledger in one transaction; dispatch is refused if any remaining counter is insufficient. Actual byte and cost usage reconciles a bounded reservation by idempotent, fenced updates so a crash or duplicate completion cannot release or consume it twice. Fork/join may instead pre-split non-overlapping child reservations, but copied tokens must never create additional budget. Retries consume the same invocation ledger rather than resetting the envelope. Request, intermediate, and result byte ceilings remain distinct: both gateway and invocation service enforce request bytes, the executor consumes intermediate bytes, and public-result construction enforces result bytes.
The first profile should allow deterministic API/MCP/rule calls, set,
switch, and assert. It should reject human ask tasks, unbounded model
calls, runner tasks, schedules, and unbounded loops.
Synchronous eligibility is transitive. Portal validation walks the complete dependency graph and rejects a synchronous tool if any reachable composite is asynchronous, contains a human or unbounded task, exceeds the aggregate invocation budget, or can outlive the outer deadline.
When the gateway wait expires, it returns an MCP error result containing a machine-readable workflow instance ID, state, and retryability. The workflow may continue durably unless its published cancellation policy says otherwise. The gateway must never silently start a second instance when the caller retries and resolves to the same gateway-derived idempotency key.
Each active synchronous wait consumes a gateway workflow-concurrency permit.
When the global, tenant, or tool-specific permit pool is exhausted, the
gateway returns a retryable WORKFLOW_CAPACITY_EXHAUSTED response before
starting an instance; it does not queue the call behind an unbounded wait.
Nested synchronous calls must not reacquire from the same pool held by their outer waits. This design reserves a separate, non-borrowable permit pool for each allowed delegation depth. A root call acquires from depth zero; a valid internal call to another synchronous workflow-backed tool increments the signed call depth and acquires from the corresponding inner pool. Ordinary HTTP or non-composite MCP backend calls do not consume workflow-wait permits. Depth-zero calls cannot consume inner reserves, and ordinary callers cannot claim an inner depth. Global, tenant, and tool limits still apply within each pool.
Every enabled synchronous delegation depth must have explicit non-zero capacity. Portal publication rejects a synchronous dependency graph whose maximum depth has no configured pool, and gateway admission fails the outer call before durable start if a required depth pool is absent, disabled, or unhealthy in the active capacity profile. Separate pools prevent a saturated set of outer waits from holding every permit needed by their own nested calls; they do not pre-reserve one inner permit per outer call. Exhaustion within an inner pool therefore remains a bounded overload failure, not a circular wait.
Interactive Execution Class And Scheduling
Fast durable acceptance is necessary but does not satisfy the synchronous SLO. The result deadline includes durable start, every task queue delay, backend execution, state transitions, final-result construction, gateway filtering, and response rendering.
At outer admission, select one immutable effective execution class for the
complete invocation chain. The initial agent-facing synchronous profile uses
interactive. For a root invocation, the binding supplies the default. For an
internal invocation, the signed delegation context supplies the effective
class and the nested binding’s value is only its direct-root default. All
descendant tasks and nested workflow invocations inherit that effective class.
An agent cannot submit, raise, or spoof it. standard and batch classes use
separate capacity shares. Batch workers may borrow unused interactive
capacity, but interactive work must be able to reclaim its reserved share
without preempting an already running side effect.
Priority alone is insufficient because the current host-task claim is global across tenants. The scheduler must provide bounded per-tenant concurrency and fair selection within each execution class, with aging so a continuously busy tenant or priority tier cannot starve another tenant indefinitely. Horizontal replicas may share the database queue, but their claim protocol and indexes must preserve the same fairness contract rather than reverting to a global oldest-row race.
Task insertion and transition commits must wake eligible executors. A suitable
PostgreSQL implementation establishes LISTEN before catch-up, drains claims
until empty, waits for a non-sensitive NOTIFY, and retains a short fallback
poll for missed notifications and recovery. The current 500 ms sleep occurs
after an empty claim, so it is primarily an idle-to-active dispatch penalty,
not automatically a 500 ms charge for every sequential hop. End-to-end tests
must nevertheless measure every queue interval because other tenants’ tasks
can be selected between transitions.
Executor capacity is a first-class deployment and admission dimension:
- configurable concurrent host-task workers per service instance;
- horizontally scalable service replicas and scheduler partitions;
- reserved interactive workers plus per-tenant and per-tool limits;
- measured queue depth, oldest runnable age, claim latency, service time, and available interactive slots; and
- admission that rejects before acceptance when the remaining deadline cannot be met with the currently advertised interactive capacity.
Gateway wait permits must be coordinated with workflow executor capacity. A deployment must not admit hundreds of synchronous waits merely because gateway connections are available while only one workflow task can execute.
Replace the host task’s Boolean lock and fixed five-minute recovery window with
a renewable lease containing lease_id, monotonically increasing
fencing_token, lease_expires_ts, and worker identity. Task completion and
transition writes succeed only when the lease and fencing token still match.
For interactive work, the initial lease and every renewal are capped by the
remaining workflow deadline. Expired tasks are reclaimed only while useful
work can still finish; otherwise they transition to a stable deadline failure.
Lease duration, renewal interval, and crash-recovery target must be materially
shorter than the synchronous deadline and tested with executor termination.
Fencing prevents a stale worker from committing workflow state, but it cannot
undo or deduplicate an external side effect. Phase 1 remains read-only; later
write-capable tasks also require the protections in
Idempotency And Side Effects and the durable
none/possible/confirmed state defined in
Failure Mapping before lease-based re-execution is allowed.
A reclaimed task in possible or confirmed state cannot automatically repeat
the call unless the downstream idempotency contract proves that replay is safe.
Asynchronous Tools
Use asynchronous publication for workflows that may:
- wait for a human decision;
- call an agent or runner with an uncertain duration;
- perform a long fan-out or batch operation;
- continue for longer than the interactive gateway deadline; or
- require cancellation or compensation after the initiating call returns.
The declared output schema returns a handle:
{
"workflowInstanceId": "019f0000-0000-7000-8000-000000000002",
"status": "ACCEPTED",
"submittedAt": "2026-08-12T20:00:00Z"
}
Expose generic MCP tools for lifecycle operations:
workflow_get_status
workflow_get_result
workflow_cancel
These tools use the same tenant and caller authorization boundary. A caller cannot discover or control another caller’s instance merely by obtaining an instance ID.
Bind instance access to both the authenticated service principal and the
initiating end-user subject/actor claim when one exists. A shared service
principal alone is not sufficient isolation. The instance stores the
publication-time classification and response-filter policy snapshot so
workflow_get_result can reproduce the approved disclosure boundary after the
original MCP session ends.
The snapshot is a maximum disclosure ceiling, not a frozen authorization grant. Every lifecycle call resolves the principal and end-user subject’s current claims, revocation state, and tenant access, then evaluates the current authorization policy. Result rendering applies the more restrictive intersection of that current decision and the stored classification/filter snapshot. A user who has lost access is denied; a later policy change cannot broaden what the accepted instance was allowed to disclose.
Token refresh must not revoke an otherwise unchanged lifecycle identity. The
gateway therefore hashes stable authorization and data-boundary claims while
excluding volatile JWT lifecycle fields such as exp, iat, nbf, jti,
and nonce. A change to roles, scopes, tenant, subject, or any other retained
boundary claim changes the digest and fails closed.
Expose the three lifecycle tools when the authenticated caller’s tenant has at
least one active asynchronous composite tool or the caller can access an
active or retained workflow instance. Retiring the tenant’s last asynchronous
composite must not remove status, result, or cancel from tools/list while an
authorized instance remains discoverable under the retention policy. The
existence check applies the same principal, initiating-subject, tenant, and
current-authorization rules as the lifecycle calls; an inaccessible instance
must not make the tools visible. This avoids expanding every tenant’s
tools/list surface while preserving lifecycle access after retirement.
Do not make one tool unpredictably return either a business result or an async handle unless its published output schema explicitly models both outcomes. Changing a published tool between synchronous business-result and asynchronous handle semantics is a breaking contract change. It requires a new stable tool identity and gateway-facing name rather than an alias rebind.
Underlying API And MCP Calls
A workflow may call registered APIs directly or invoke existing gateway MCP tools.
The preferred default for a composite MCP tool is a workflow call: mcp using
approved, pinned gateway tools because this preserves:
- stable tool identity;
- gateway service discovery and argument mapping;
- fine-grained authorization and response filtering;
- shared audit and diagnostics; and
- consistent MCP/API behavior.
Direct HTTP workflow tasks are appropriate when workflow policy explicitly authorizes a registered service endpoint and service identity. The workflow must not accept a destination URL from an agent or transform expression.
At publication time, resolve each nested tool reference to:
stableToolRef
gateway-facing name
tool version and contract digest
contract compatibility class
logical authorization tool name and endpoint key
authorization-policy reference
lifecycle status
The workflow definition digest is always an exact immutable pin. Nested tool
contracts use versioned dependency resolution: an outer composite continues to
dispatch the approved nested version after a newer alias is published, so an
unrelated inner publication cannot cause an outer runtime outage. An optional
follow-compatible policy may advance only when Portal proves that the new
nested input accepts every previously valid outer request, the new output is
within the contract the outer mapping expects, and authorization or data-
classification policy has not broadened. Because general JSON Schema
compatibility is not decidable for every schema, unsupported or ambiguous
changes are incompatible and require an explicit outer repin, conformance
test, and reapproval.
Portal uses the dependency reverse index to show the inner publisher the affected composites before promotion. Security revocation can still invalidate a pinned dependency immediately; ordinary version promotion cannot.
Phase 1 therefore requires version-aware internal dispatch, not alias lookup.
The control plane projects a private dependency-target registry alongside
mcp-router.tools, keyed by stableToolRef, tool version, and contract digest.
The public alias exposes only the currently promoted version through
tools/list; a workflow delegation token invokes the pinned private target by
stable reference and version. Private targets reuse the same authorization,
argument masking, backend dispatch, response filtering, schema validation, and
audit pipeline, but cannot be invoked by an ordinary external tool name.
Authorization identity belongs to the logical tool; dispatch identity belongs
to the version target. Every private target therefore carries the logical
public tool name and the exact endpoint key used by its alias, such as
accounts@call, and passes those values through the existing authorization and
response-filter pipeline. Its registry key, private target name, version, and
digest never derive a new endpoint key or new rule binding. This is required
for defaultDeny: true, where an unrecognized endpoint or one without request
rules is denied.
An approved version may change backend resolution and contract digest without changing that logical authorization identity. A version that needs a different endpoint key, request rule set, permission boundary, or response-filter policy represents a different capability. Portal classifies it as incompatible and requires an explicit outer repin, conformance tests, and approval rather than silently minting a version-specific authorization identity. Current security revocation continues to take precedence over any pin.
Superseded and retirement-candidate targets remain dispatchable while referenced by any active composite binding, in-flight workflow snapshot, rollback window, or required audit/replay retention. Portal maintains reference counts or equivalent durable reachability evidence. Retirement and garbage collection use that same reachability index. Garbage collection is allowed only after all such references and retention holds are gone, and removal is itself projected and audited. This makes the claimed version pin executable in Phase 1 rather than depending on a future API.
Delegation And Cycle Prevention
The gateway issues a short-lived workflow-task delegation token containing or binding:
- tenant and initiating principal;
- outer stable tool reference;
- allowed nested stable tool references;
- allowed audiences and operations;
- input/data-boundary digest;
- correlation ID;
- immutable effective execution class and current synchronous permit depth;
- structured invocation budget for deadline, task attempts, nested calls, depth, bytes, and cost, plus the identifier and generation of the shared durable budget ledger that owns the mutable counters;
- idempotency context; and
- remaining delegation depth.
Nested calls can only narrow rights and budget reservations. They cannot extend the initiating deadline, add tools, broaden the data boundary, change tenant, or select their own scheduling class. When dispatching another workflow-backed tool, the gateway increments permit depth and copies the effective execution class into the nested token. It accepts these claims only from a gateway-issued delegation token; an external MCP request always enters at depth zero and uses its root binding’s class. Token verification authenticates the immutable ceiling and ledger identity; every mutable consumption decision is an atomic conditional update against that ledger, never a decrement trusted from token contents.
maximumParallelism remains in gateway and Tool-binding wire formats for
backward compatibility, but it is not an enforced invocation-budget dimension.
For both REST and event-driven starts, fork width is governed only by the
light-workflow service’s WORKFLOW_MAXIMUM_PARALLELISM setting.
Portal publication builds a dependency graph for every workflow-backed tool. It rejects:
- a workflow that calls its own composite tool;
- a cycle across two or more workflow-backed tools;
- a call to an unbound or retired tool;
- a nested call whose approved version or contract digest is unresolved or incompatible; and
- a path that exceeds the configured maximum delegation depth.
The runtime also enforces the depth and allowed-tool set so a stale or malicious definition cannot bypass publication checks.
Authorization And Data Protection
Authorization happens at two levels:
- The gateway authorizes the caller to invoke the composite business tool.
- Each workflow task is authorized for its specific underlying API or MCP tool using a narrowed delegation or approved workflow service identity.
The first authorization does not imply unrestricted access to every tool used by the workflow. The binding and workflow policy define the exact internal capability set.
Nested identity is declared per step and defaults to narrowed initiating-user delegation. Workflow service identity is permitted only when the step has explicit publication-time approval evidence because it can authorize an operation the initiating user could not perform directly. Service identity must remain tenant-bound, tool-bound, deadline-bound, and no broader than the published capability set.
The MCP Tools Access Control response-filtering boundary still applies to the final result. Intermediate workflow context and task outputs need their own classification and redaction rules because they may contain more data than the final caller is allowed to receive.
Required protections include:
- registered endpoint and workflow references only;
- no arbitrary URL, credential, or service-id arguments;
- encrypted secret references rather than credentials in workflow YAML;
- bounded request, intermediate context, task output, and final output sizes;
- redaction before logs, events, traces, and AI authoring context;
- tenant-bound workflow instance lookup;
- fail-closed schema and digest mismatches; and
- explicit approval policy for destructive or high-impact tools.
Idempotency And Side Effects
Do not depend on an agent to generate or replay a correct idempotency key. For read-only and ordinary synchronous calls, the gateway derives the key from the tenant, authenticated principal and end-user subject, stable tool reference, workflow definition digest, and canonical effective input. Because input and definition digests participate in this derived key, a different input or version intentionally creates a different invocation; it is not an idempotency conflict.
A client-provided Idempotency-Key, explicit business-key input such as
requestId, or configured business-key expression is accepted only when the
published policy allows it. The gateway scopes and hashes that untrusted key
with tenant, trusted identity, and stable tool fields, and stores the effective
input and definition digests beside it. Reusing an explicit scoped key with
different input or definition produces WORKFLOW_IDEMPOTENCY_CONFLICT.
Side-effecting workflows require a stronger business idempotency contract because two intentionally distinct operations may have identical arguments.
The publication UI must require one of:
- an explicit idempotency input field;
- a deterministic business-key expression;
- an upstream server-enforced idempotency key; or
- a declaration that duplicate effects are impossible or compensated, with approval evidence.
The workflow invocation service stores the accepted key with the stable tool reference, workflow digest, trusted identities, and effective-input digest. The database enforces one current reservation with a unique constraint over tenant, authenticated principal/end-user subject, stable tool reference, and the final scoped key. Definition and input digests are stored values, not part of that uniqueness key, so an explicit-key conflict can be detected. The implementation uses one atomic insert/conflict or compare-and-swap path rather than read-then-write.
Separate two time windows:
- In-flight deduplication lasts until the instance reaches a terminal state or its maximum deadline and uncertain-outcome retry grace have elapsed. A duplicate returns the existing instance.
- Completed-result replay is a separate publication-time freshness policy.
Before
result_replay_until, an identical request returns the completed instance and result. After it expires, an atomic reservation-generation change starts a new instance while retaining immutable history.
Read-only tools should normally use a short completed-result replay window so repeated questions can observe fresh data. Write-capable tools require a window consistent with their downstream side-effect idempotency and retry contract. Retention of invocation and audit history is independent of whether the active reservation can advance to a new generation.
Canonical effective input is the schema-validated workflow input after
approved deterministic defaults and request mappings, before logging
redaction. It uses a versioned JSON Canonicalization Scheme profile based on
RFC 8785: duplicate object keys
are rejected, object properties are recursively sorted by the RFC’s UTF-16
ordering, array order is preserved, and finite numbers use the specified
deterministic representation. Values outside the interoperable IEEE 754 range,
including large integer identifiers, must be schema-declared strings. Absent
properties remain absent and therefore differ from explicit null. Unicode
string code points are preserved exactly; NFC or other Unicode normalization
is not applied. These rules and the profile version are pinned by conformance
fixtures shared by Portal, gateway, and invocation service.
Event-source deduplication, such as a unique source event ID used while
projecting WorkflowStartedEvent, remains a separate safeguard and does not
satisfy caller-invocation idempotency.
For multi-step writes, the workflow definition owns compensation. The gateway does not attempt to reverse completed backend operations.
Failure Mapping
Keep business outcomes separate from technical failures.
Business outcomes such as NO_CONSENT or NO_ELIGIBLE_OFFER are successful,
schema-valid tool results. Technical failures produce isError: true and a
stable machine-readable class such as:
WORKFLOW_INPUT_INVALID
WORKFLOW_START_REJECTED
WORKFLOW_DEFINITION_MISMATCH
WORKFLOW_TIMEOUT
WORKFLOW_CANCELLED
WORKFLOW_TASK_FAILED
WORKFLOW_OUTPUT_INVALID
WORKFLOW_OUTPUT_INVALID_AFTER_EFFECT
WORKFLOW_POLICY_DENIED
WORKFLOW_CAPACITY_EXHAUSTED
WORKFLOW_INVOCATION_UNAVAILABLE
WORKFLOW_IDEMPOTENCY_CONFLICT
The error envelope should include the workflow instance ID when one exists, whether retry is safe, and a correlation ID. It must not expose credentials, raw internal errors, hidden task inputs, or backend responses that have not passed disclosure policy.
The workflow tracks whether externally visible side effects are none,
possible, or confirmed. If public-result construction or output validation
fails after a confirmed effect, return
WORKFLOW_OUTPUT_INVALID_AFTER_EFFECT with retryable: false and the instance
ID. This is operationally distinct from a pre-effect validation failure; an
agent must not repeat the write merely because its result was undeliverable.
Transformation And Aggregation Language
Use the workflow DSL as the only authoring contract. Do not invent a gateway mapping language for composite tools.
Use CEL as the canonical expression language for workflow conditions and data transformations. This keeps one expression contract across Light-Fabric rules, gateway policies, workflow authoring, Portal validation, and AI generation. CEL is not limited to boolean decisions: an expression can also construct lists, maps, and JSON-compatible objects. The rule engine can retain its boolean-only contract while the workflow adapter accepts a typed value.
CEL And jq Comparison
| Concern | CEL | jq |
|---|---|---|
| Primary fit | Policy, conditions, validation, routing, and computed values. | JSON extraction, reshaping, pipelines, and complex aggregation. |
| Validation | Can parse and type-check against declared variables and functions before publication. | Dynamically evaluated; schema and type mistakes normally surface during execution. |
| Result model | Produces one typed value. | A filter can produce zero, one, or many streamed values. |
| Collection support | Provides map, filter, exists, and all; sufficient for common mappings. | Provides concise sort_by, group_by, unique_by, reduce, and recursive traversal. |
| Safety model | Side-effect-free and terminating, but nested collection macros still need cost limits. | Recursion, while, and repeat require time, memory, depth, and output limits. |
| Platform cost | Reuses the existing Light-Fabric language, evaluator experience, security profiles, and Portal contract. | Adds another runtime, validator, editor mode, security profile, compatibility contract, and AI prompt. |
CEL therefore provides the better default for a governed platform. jq is more ergonomic for some advanced JSON transformations, but that advantage does not justify exposing two interchangeable languages throughout every workflow.
References:
CEL Runtime Contract
Provide a shared CEL execution core with location-specific adapters:
- a predicate adapter that requires
boolfor rules,when,switch, retry predicates, and assertions; and - a value adapter that converts one CEL result to JSON for task inputs, exports, derived values, joins, and the final public output.
The existing rule boundary remains unchanged. CEL rule conditions decide whether declarative actions execute; they do not directly mutate a response or become a general workflow runtime. Reuse the compiler, type environment, security validation, value conversion, cost accounting, and diagnostics rather than coupling workflow execution to the rule engine’s boolean API.
The workflow CEL environment exposes only immutable, documented roots:
| Root | Contents |
|---|---|
input | Schema-validated workflow invocation input. |
context | Accumulated workflow state and approved task exports. |
task | Current task input or result where the expression location permits it. |
workflow | Bounded identifiers and execution metadata, never credentials or secret values. |
Each expression location declares its required result category and, when available, its JSON Schema-derived type. Portal publication must parse and type-check the expression against that environment, reject undeclared roots or functions, and persist the normalized expression and digest with the immutable workflow version. Runtime execution uses the same environment declaration and a cached compiled program; it must not reinterpret an expression under a different profile.
The production CEL profile must enforce:
- allowlisted roots, functions, and collection macros;
- no I/O, network access, mutation, service lookup, or dynamic code loading;
- expression-size, evaluation-cost, collection-size, nesting-depth, execution-time, and result-size limits;
- explicit guards for missing or nullable data where required; and
- deterministic failure when evaluation returns the wrong type or cannot be converted to exactly one JSON value.
The CEL value adapter initially needs to support:
- selecting fields;
- reshaping objects and arrays;
- joining previously exported task results;
- computing derived values;
- filtering collections; and
- constructing the public output.
CEL To JSON Contract
The value adapter is new production code, not a thin rename of the existing boolean evaluator. The compile cache, reference inspection, guarded execution, and JSON-to-CEL context conversion are reusable; the reverse conversion and result contract require their own implementation and conformance suite.
The initial cross-runtime CEL-to-JSON mapping is:
| CEL value | JSON representation |
|---|---|
null, bool, string | Corresponding JSON value. Missing remains distinct from explicit null. |
int, uint | JSON number only within the interoperable ranges [-9007199254740991, 9007199254740991] for int and [0, 9007199254740991] for uint; authors must explicitly convert larger identifiers or counters to strings. |
double | Finite JSON number; NaN, positive infinity, and negative infinity are rejected. |
bytes | Standard padded Base64 string and a schema that declares the encoding. |
timestamp | UTC RFC 3339 string normalized with a Z suffix. |
duration | Protobuf JSON duration string in seconds with optional fractional nanoseconds, for example "1.500s". |
list | JSON array after recursively applying this contract. |
map | JSON object only when every key is a unique string; non-string or colliding keys are rejected rather than stringified. |
| opaque, function, optional-without-value, or implementation-specific values | Rejected unless a later versioned profile defines an explicit conversion. |
The adapter must not silently convert an unsupported value to null. The
normalized mapping version is part of the compiled-expression/profile digest
so a library upgrade cannot change persisted results without conformance and
promotion.
Phase 0 must also qualify the concrete evaluator. The currently pinned Rust
cel 0.14 API provides parsing, execution, reference inspection, and a generic
JSON helper, but it does not expose the schema-aware checker or evaluation-cost
budget assumed by this design. Its generic JSON helper also stringifies map
keys and chooses dependency-specific representations for types such as
duration. Before Phase 1, either augment or replace that integration with an
implementation that satisfies the checker, cost, and conversion contracts, or
narrow the published CEL profile and validation claims accordingly. Wall-clock
timeouts alone are not a substitute for deterministic evaluation-cost limits.
The current workflow model advertises jq and JavaScript while the runtime implements only a small jq-like path and comparison evaluator. Before this profile is published, change the workflow default to CEL, execute expressions through the shared CEL core, and make Portal and runtime validation reject jq, JavaScript, and any other unimplemented language.
Optional Advanced jq Transform
Do not enable jq in the initial production profile. If representative customer
workflows later demonstrate a material need for operations such as grouping,
sorting, reducing, or recursive JSON traversal that would otherwise require
non-portable CEL extensions, jq may be introduced as an explicit advanced
transform task. It must not become a per-expression alternative for
conditions, policies, retries, or ordinary mappings.
An optional jq task requires a separate versioned compatibility and security profile. It must accept one JSON input and return exactly one JSON output; zero-result and multi-result filters fail unless the task contract explicitly collects them into one array. The allowed subset must exclude unbounded recursion and repetition, imports and modules, environment or input access, and debug or stderr output. The runtime must enforce fuel or cost, time, memory, depth, input, and output limits.
JavaScript should not be enabled merely because it appears in the workflow model. When declarative CEL transformations are insufficient and the optional jq profile is not appropriate, an approved isolated runner task may be used. Its input, output, image or template digest, resource limit, and execution policy must be pinned. It cannot execute in the gateway process.
Parallel aggregation requires explicit fork/join semantics with bounded
parallelism and a deterministic merge rule. Step retries require explicit
attempt count, retryable error classes, backoff, jitter, and idempotency
requirements. These semantics belong in light-workflow, not in the MCP tool
configuration.
Portal Authoring Experience
Add a Composite MCP Tool workspace to portal-view. Reuse the existing
workflow editor, validation, graph, and test-run surfaces.
Contract
The user defines:
- MCP name, description, semantic metadata, and examples;
- input and output JSON Schemas;
- synchronous or asynchronous mode;
- MCP result text mode and completed-result freshness window;
- latency, deadline, fan-out, and step limits;
- read-only, idempotent, destructive, and approval metadata; and
- target gateway instances or environments.
Flow
The user can edit YAML directly or assemble a graph from registered API endpoints, MCP tools, rules, and supported workflow tasks. Selecting a source operation inserts its stable reference and current schema digest rather than a free-form URL.
Mappings
Each task exposes editors for:
- workflow input to call arguments;
- task output to workflow context;
- branch expressions;
- join and aggregation expressions; and
- final public-output mapping.
All expressions use CEL in the initial production profile. The editor shows the allowed roots and functions for that location, provides schema-aware completion and syntax/type diagnostics, states the expected result type, and previews the input and output shape at each step.
Generate With AI
AI generation is a draft-authoring feature, not a production execution path. The generation request contains:
- the user’s business objective;
- only the APIs and MCP tools selected or authorized for the author;
- their schemas, descriptions, examples, and safety metadata;
- the supported workflow DSL and CEL profile;
- organization policy and runtime bounds; and
- optional sample input and expected output.
The model must not receive credentials or unrestricted catalog access. It must not invent endpoints, tools, schema fields, or runtime features.
Generation produces:
- a workflow draft;
- a proposed input and output contract;
- dependency and mapping explanations;
- positive, edge, and failure fixtures; and
- assumptions and unresolved questions.
The draft remains unpublished until it passes deterministic validation and a human approves the diff.
Validate And Test
The Portal runs, in order:
- YAML and workflow-schema validation.
- Runtime-supported task and expression validation.
- Input/output JSON Schema validation.
- Stable tool, endpoint, and workflow reference resolution.
- Schema-digest and lifecycle checks.
- Dependency-cycle and delegation-depth checks.
- Safety, approval, idempotency, result-rendering, and data-boundary policy checks.
- Mock fixture tests.
- Optional live sandbox tests with failure injection.
- Gateway
tools/listand invocation qualification against a non-production instance.
AI-generated and manually authored definitions use the same validation and publication pipeline.
Publication, Versioning, And Rollback
Publication creates an immutable bundle containing:
stable tool reference
gateway-facing alias
input/output schemas and schema digest
workflow definition ID, version, and digest
nested dependency snapshot
dispatch and delegation policy references
response-classification and filtering policy digest
execution class, runtime bounds, result text mode, and replay windows
test evidence
approver and publication metadata
The Portal projects the runtime subset into mcp-router.tools. The gateway
continues to use last-known-good configuration when a new snapshot is invalid.
Promotion atomically moves the tool alias to the new approved binding. New calls use the new binding; in-flight workflows continue using their stored definition and policy snapshots.
Before moving an inner tool alias, Portal queries the dependency reverse index, classifies the contract change, and shows the affected composite tools. An incompatible change cannot strand existing outer bindings at runtime: either the old nested version remains dispatchable, or the outer tools are explicitly repinned, retested, and reapproved in the same promotion plan.
Retiring an inner tool uses the same reverse-index gate. Portal blocks retirement while an active outer binding references the tool unless one atomic plan repins or retires every affected outer binding. Retirement prevents direct new starts and new dependency publication, but it does not invalidate the pinned dependency snapshot of a workflow accepted before that plan; its private target remains dispatchable until the in-flight and retention references are released. An emergency security revocation is a separate fail-closed operation and may deliberately break pinned calls with an explicit impact report and audit record.
Promotion must not change a stable tool between synchronous business-result and asynchronous handle contracts. That change requires a new stable tool reference and gateway-facing name so cached schemas and static allowlists do not observe a semantic type change behind an alias.
Rollback republishes the previous approved binding. It does not mutate or delete historical definitions or running instances.
Retirement removes the tool from new tools/list responses and rejects direct
new starts while preserving the pinned execution dependencies, status, result,
cancellation, and audit access needed by already accepted instances according
to retention policy.
Observability And Audit
Use one correlation ID across:
outer MCP request
gateway workflow dispatch
workflow instance
workflow tasks
nested gateway/API/MCP calls
final MCP result
Recommended gateway span and audit attributes include:
mcp.tool.name
mcp.tool.stable_ref
mcp.tool.endpoint_id
mcp.tool.execution_placement
workflow.definition_id
workflow.definition_digest
workflow.instance_id
workflow.invocation_mode
workflow.execution_class
workflow.permit_depth
workflow.state
workflow.task_count
workflow.nested_call_count
workflow.delegation_depth
workflow.wait_ms
workflow.total_ms
High-cardinality identifiers such as workflow.instance_id, correlation ID,
and raw digest values belong in spans and audit events, not metric labels.
Metrics use bounded dimensions such as tenant tier, stable tool reference where
cardinality policy permits it, workflow version, state, and normalized error
class.
Metrics should cover:
- starts, completions, failures, cancellations, and timeouts;
- durable-acceptance, runnable-to-claim, task service, transition, synchronous wait, result-rendering, and total workflow latency;
- interactive queue depth, oldest runnable age, executor saturation, lease expiry/reclaim, and deadline-aware admission rejection;
- active and waiting instances;
- duplicate/idempotent start hits;
- definition, schema, and policy mismatch rejections;
- nested-call denials and cycle/depth rejections;
- output-validation failures; and
- capacity rejection by tenant, tool, workflow version, and bounded permit depth.
Do not attach raw inputs, intermediate context, or final results to metrics. Trace and audit payload capture follows classification and redaction policy.
Capacity And Availability
The gateway must bound workflow dispatch independently from HTTP and MCP backend dispatch. Recommended controls include:
- global and per-tenant concurrent workflow starts;
- per-tool and per-delegation-depth concurrent synchronous waits;
- executor-advertised interactive slots, queue age, and throughput estimates in admission decisions;
- workflow-invocation connection and response timeouts;
- circuit health for the invocation service;
- request and public-result size limits;
- maximum pending asynchronous instances where policy requires it; and
- overload responses that distinguish safe retry from an accepted workflow.
A gateway timeout must not be reported as “not started” after the workflow service has durably accepted the instance. The invocation service returns the instance ID as part of durable acceptance, and retries use idempotency lookup to resolve uncertain outcomes.
Invocation-service health does not remove an already published tool from
tools/list; discovery is a stable contract and may be cached by agents. Calls
fail with retryable WORKFLOW_INVOCATION_UNAVAILABLE before acceptance while
the circuit is open. After durable acceptance, errors return the instance ID
and current state instead of an ambiguous unavailable response.
The gateway remains stateless with respect to workflow progress. Gateway restart or reload does not lose the workflow instance.
Implementation Phases
Phase 1 is primarily a light-workflow runtime qualification project with a
gateway feature attached. Direct invocation, interactive scheduling, fair
claiming, notification wake-up, fenced leases, deadline-aware admission, and
the CEL evaluator decision are release prerequisites rather than follow-up
gateway optimizations. Delivery ownership, staffing, and milestones must
reflect that dependency order.
Phase 0: Contract And Threat Model
Owners: light-fabric, light-workflow, portal-db, and light-portal.
The versioned implementation artifacts live under
contracts/workflow-invocation/v1. The shared Rust types and strict
canonicalizer live in workflow-invocation-contract; the direct transactional
acceptance boundary is light-workflow::invocation; and the matching Portal
schema patch is patch_20260812_01_workflow_mcp_phase0.sql. Run
scripts/run-workflow-mcp-phase0-gates.sh with a disposable PostgreSQL URL to
verify both repositories. The qualification manifest deliberately keeps
runtime promotion disabled until the exact Phase 1 scheduler/executor topology
and replacement or augmented CEL evaluator produce passing evidence.
- Define the invocation API, state model, public result, and error envelope.
- Define
workflow_tool_binding_t, the nested-dependency reverse index, the invocation/idempotency store, the durable atomic invocation-budget ledger, and their command/query events. - Define exact workflow pins, supported nested-schema compatibility rules, dependency impact reporting, private version-target retention/garbage collection, and policy digest rules.
- Define delegation claims, depth-partitioned synchronous permit pools, inherited execution-class semantics, idempotency semantics, cancellation behavior, and synchronous eligibility rules.
- Define the MCP result envelope, canonical-input profile, completed-result freshness window, and explicit-key conflict behavior with shared fixtures.
- Spike and load-test the complete interactive result path, including direct transactional acceptance, fair task scheduling, wake-up, executor capacity, and crash recovery. Durable acceptance must create the invocation, process, initial task, snapshots, idempotency reservation, and audit outbox event without waiting for the shared event consumer. Measure acceptance separately as a necessary sub-gate, but gate the architecture on end-to-end result latency.
- Qualify the CEL implementation for schema-aware checking, deterministic cost enforcement, and the pinned CEL-to-JSON conversion table. Decide whether to augment, upgrade, or replace the current crate before freezing fixtures.
- Publish OpenAPI/JSON Schema fixtures and positive/negative conformance tests.
Exit gates:
- Portal, gateway, and invocation-service fixtures give the same canonical
digest for reordered objects, numeric spellings, absent versus
null, preserved Unicode, and rejected duplicate keys; - a committed start returns its preallocated instance ID and an immediate
status read returns
ACCEPTEDor a later state, never projection-lag404; - numeric acceptance p95/p99 sub-gates and end-to-end result p95/p99 gates are selected before Phase 1. The result gate uses controlled one-task and maximum-task workflows under concurrent interactive load, cross-tenant batch backlog, unrelated outbox traffic, and poison-event injection, with every queue and execution stage measured separately;
- the selected result-latency gate is met without tenant starvation, and an executor killed after claim is recovered within the remaining interactive deadline while its stale fencing token cannot commit;
- CEL conformance fixtures pin large integers, finite/non-finite doubles, timestamps, durations, bytes, null versus missing, map keys, opaque values, result types, checker behavior, and cost exhaustion;
- stale workflow definitions fail closed, while nested compatible changes and incompatible repin requirements follow the documented rules;
- an inner contract change produces a complete reverse-dependency impact report before promotion, an existing outer binding continues to call its pinned private version, and garbage collection refuses a referenced target;
- cross-tenant start/status/result/cancel tests fail closed;
- concurrent derived-key duplicates resolve to one workflow instance, changed
derived input starts a new instance, reuse of one permitted explicit key with
changed input returns
WORKFLOW_IDEMPOTENCY_CONFLICT, and expiry of the completed-result replay window atomically starts a new generation; - parallel consumers of copied delegation tokens share one ledger: with only
Nattempts, calls, bytes, or cost units remaining, at mostNreservations commit, retries do not reset counters, and duplicate or stale fenced reconciliation cannot double-release a reservation; - compact-JSON and summary-mode fixtures pin
content,structuredContent, filtering, schema validation, size limits, and technical-error rendering.
Phase 1: Read-Only Synchronous MVP
Owners: light-gateway, light-workflow, light-portal, and portal-view.
- Add the workflow execution-placement dispatch branch and configuration parser
to the gateway while keeping
apiTypelimited to backend transports. - Add depth-partitioned synchronous permit pools whose inner reserves are reachable only through signed workflow delegation.
- Add the direct durable invocation façade and resumable bounded long-poll operation; do not drive gateway starts through the shared Portal event log.
- Add bounded poison-event quarantine and audited replay to workflow event consumers.
- Add the immutable interactive execution class, fair per-tenant claiming, wake-on-insert, configurable concurrent executors, deadline-aware admission, and renewable fenced task leases.
- Support headless, bounded sequential compositions.
- Replace the placeholder jq-like evaluator with the shared CEL predicate and value adapters, change the workflow default to CEL, and validate expressions before publication.
- Add manual Composite MCP Tool authoring, validation, test, and publication.
- Project the private version-target registry and dispatch nested workflow calls by pinned stable reference, version, and contract digest rather than alias; preserve the logical authorization identity and retirement reachability.
- Use narrowed initiating-user delegation by default for nested MCP calls; require approval evidence for each service-identity step.
- Restrict the first profile to read-only operations.
Exit gates:
- an unchanged generic MCP client discovers and calls a composite tool;
- the final result passes gateway response filtering and output validation,
appears unchanged in
structuredContent, and uses the published text mode; - gateway restart does not lose an accepted workflow;
- any gateway node can resume a wait, concurrent waiters observe one durable instance, and capacity exhaustion rejects before starting or queueing;
- with
Npermits configured at root and first-nested depth,Nconcurrent root workflows that each call one controlled nested composite complete without root saturation consuming the nested reserve; direct requests cannot claim or spoof an inner-depth permit; - workflow/gateway traces share one correlation ID;
- the Phase 0 end-to-end p95/p99 result gates continue to pass at the declared executor concurrency and cross-tenant backlog, with bounded fairness and no starvation;
- an idle executor is woken without waiting for the fallback poll, and an executor crash reclaims an interactive task before its deadline while a stale completion is rejected by fencing;
- Portal validation and workflow execution agree on CEL conformance fixtures, and jq or JavaScript definitions fail publication and runtime loading;
- recursive workflow-tool dependencies and transitively async or unbounded dependencies are rejected for synchronous publication;
- a nested workflow inherits the outer effective execution class even when its binding has a different direct-root default;
- with
defaultDeny: true, promoting an inner alias does not change either the version or logical authorization identity used by an existing outer binding, and its pinned call still passes the alias endpoint’s rules without a new private-target rule; - retirement is rejected while active outer references exist unless one plan repins or retires them, and referenced or in-flight private targets cannot be retired from dispatch or garbage-collected;
- the configured static-task, runtime-attempt, nested-call, payload, cost, wait, and total-deadline budgets are enforced by the shared invocation ledger rather than mutable token claims;
- a malformed event is quarantined after bounded attempts and cannot block start, status, wait, result, or unrelated aggregates in its tenant partition; every deferred offset remains replayable, and retention or purge refuses to remove its payload dependency.
Phase 2: Production Orchestration
Owners: light-workflow and the workflow client in light-gateway.
- Complete deterministic, schema-aware CEL transformations and aggregation conformance with cost and result-size enforcement.
- Add bounded fork/join aggregation and generic task retries.
- Add explicit task and workflow deadlines.
- Add asynchronous start/status/result/cancel tools.
- Add side-effect idempotency, approval, and compensation policies.
- Distinguish post-effect output failures as non-retryable and preserve their side-effect state in status and audit responses.
- Add dependency-drift checks and version promotion/rollback.
- Use representative customer workflows to decide whether the optional restricted jq transform task is justified; keep jq rejected otherwise.
Exit gates:
- parallel partial failure is deterministic and auditable;
- fork/join siblings and concurrent retries carrying copies of the same signed
parent token cannot exceed aggregate attempt, call, byte, or cost ceilings;
tests race
N+1reservations against a remaining budget ofNand prove that exactly one is rejected without overspend; - retry tests never duplicate protected side effects;
- long-running and human-task workflows return and enforce authorized handles;
- revoking the initiating subject after acceptance denies status/result access, and result filtering never discloses more than the intersection of current authorization and the stored publication-time ceiling;
- cancellation reaches a terminal state or reports a stable non-cancellable reason; and
- rollback changes new starts without changing in-flight snapshots.
Phase 3: AI-Assisted Authoring
Owners: portal-view, Portal GenAI services, and workflow validation.
- Generate drafts from selected registered operations and schemas.
- Generate contract, mapping, edge, and failure fixtures.
- Show assumptions, dependency graph, policy findings, and a human-readable diff.
- Record generator model, prompt/template version, source schema digests, and reviewer approval as provenance.
- Prohibit direct AI-to-production publication.
Implementation contract:
workflow-queryexposesgenerateWfDefinitionDraftas a draft-only query. It never creates, updates, or publishes a workflow.- The query sends only explicitly selected, authorization-filtered tool metadata to an OpenAI-compatible authoring model. It strips secret-bearing fields, caps the operation count and context size, treats descriptions and schemas as untrusted data, and refuses credentials in the intent or existing definition.
- Configure the authoring service with
WORKFLOW_AUTHORING_LLM_URLandWORKFLOW_AUTHORING_LLM_MODEL.WORKFLOW_AUTHORING_LLM_BEARER_TOKENis optional, andWORKFLOW_AUTHORING_LLM_TIMEOUT_SECONDSis bounded to 1-60 seconds. These values remain server-side and are never returned to Portal. - The model response is accepted only as strict JSON containing definition,
assumptions, policy findings, and contract, mapping, edge, and failure
fixtures. A deterministic
workflow-mcp-phase3validator then rejects unavailable tools, non-MCP generated calls, unsupported tasks, nested or unbounded forks, jq, JavaScript, and non-CEL expression profiles. - Portal shows the proposal, bounded human-readable diff, assumptions,
dependency graph, fixture categories, policy findings, and generator
provenance. Applying it requires a signed-in reviewer checkbox and records
model, prompt-template, source-schema, request, definition-digest, and
reviewer evidence under
document.metadata.aiAuthoring. - Strict server validation binds
reviewerUserIdto the authenticated subject; it does not trust the reviewer identity supplied by the browser. - The first save of an AI-authored draft is private. Server validation fails closed when the authorization-filtered catalog or validator is unavailable, and recomputes the semantic definition digest after removing only provenance metadata so post-approval edits require another review. Production exposure remains a separate promotion action using the normal gates.
- Workflow create and update command handlers independently reject attempts to
set
catalogVisible: trueon an AI-authored definition, so direct RPC calls cannot bypass the private-first rule.
Exit gates:
- generated definitions cannot reference unavailable tools or unsupported DSL features;
- secrets and unauthorized catalog entries never enter generation context;
- deterministic validators reject unsafe or inconsistent drafts; and
- manual and AI-authored workflows pass the same promotion gates.
Run scripts/run-workflow-mcp-phase3-gates.sh from light-fabric to execute
the existing Phase 2 runtime gates plus the authoring-service tests, Portal
review tests, lint checks for the Phase 3 UI, and a release-mode Portal build.
This is an implementation check, not production qualification.
Phase 4: Optional Skill Integration
Owners: Portal skill registry and agent catalog.
skill_workflow_tmay carryworkflow_binding_idand the binding-derivedworkflow_tool_id. Both are nullable so ordinary skill/workflow links remain valid, but they must either both be absent or both be present.- A composite foreign key binds the skill link to the exact
(host, binding, workflow, tool)tuple. A second foreign key requires that exact tool to be present inskill_tool_t, so progressive disclosure cannot advertise a capability the skill was not granted. - Portal lists only active workflow-backed bindings whose workflow matches the
selected definition and whose current tool schema digest matches the pinned
binding digest. The command service derives
workflow_tool_idfrom the trusted binding; browsers cannot supply it. - Query and effective-agent-catalog responses include the tool name, bound workflow version, definition digest, and schema digest. The Skill Workspace validates these invariants and renders the resolved contract for reviewers.
- Direct MCP discovery remains independent:
workflow_tool_binding_thas no foreign key to a skill, and a published binding with no skill link continues to be projected and invoked normally.
Exit gates:
- a linked skill, tool, binding, and workflow resolve to one exact pinned contract;
- a mismatched workflow or a tool absent from
skill_tool_tis rejected by database constraints and command validation; - a workflow-backed MCP tool without a skill link remains valid and directly discoverable;
- create/update and validation RPC schemas expose the optional binding without accepting a caller-selected tool id; and
- Portal exposes the optional selector and the resolved version/digests.
Run scripts/run-workflow-mcp-phase4-gates.sh from light-fabric to execute
the prior runtime/authoring gates, the Portal database schema/constraint gate
when a disposable PostgreSQL URL is supplied, Portal persistence and GenAI
command/query tests, and the Portal UI lint/build checks. The development
contract does not include legacy-data migration or backward-compatibility
validation.
Acceptance Criteria
The design is complete when:
- an existing MCP client can invoke a meaningful multi-API capability through
one ordinary
tools/call; - no agent-side skill or orchestration implementation is required;
- the gateway contains no general workflow interpreter or executable user scripts;
- one canonical workflow definition backs both skill-aware and direct MCP exposure;
- every published tool is bound to stable tool, schema, workflow, dependency, and policy digests;
- synchronous result latency and tenant fairness are qualified against declared executor capacity, backlog, and crash-recovery gates;
- nested synchronous calls cannot be starved by root wait permits and inherit the outer execution class through signed delegation;
- synchronous and asynchronous behaviors are explicit in the tool contract;
- retries, cancellation, idempotency, and compensation have one runtime owner;
- outer and nested authorization remain tenant- and caller-bound;
- private version dispatch preserves the logical tool’s authorization identity, and promotion or retirement cannot strand active outer dependencies;
- shared-principal asynchronous access is also bound to the initiating end-user subject, current authorization, and stored response-filter ceiling;
- AI-generated workflows are drafts until deterministic checks and human approval complete; and
- publication, promotion, retirement, and rollback preserve in-flight workflow snapshots and audit history.
Settled And Remaining Decisions
This design settles the architecture choices that gate implementation:
- Phase 1 is read-only.
- Gateway starts use direct transactional durable acceptance rather than the shared event consumer.
- Nested calls select identity per step, defaulting to narrowed initiating-user delegation; workflow service identity requires explicit approval evidence.
- The workflow
outputblock is the only executable public-result mapping. - Workflow execution is selected by
executionPlacement, not a new transport value inapiType.
The remaining product and operational choices require measured customer or environment data rather than another runtime architecture:
- What p95/p99 acceptance, execution, and wait SLOs should each deployment use, and what maximum synchronous wait follows from those measurements?
- How many customer agents dynamically refresh
tools/list, and how many use static allowlists that require an explicit rollout procedure? - Which asynchronous result fields remain queryable after the initiating MCP session ends, and what retention and legal-hold policies apply?
Related Documentation
- MCP Router
- MCP Tool Metadata Usage
- MCP Tools Access Control
- MCP Tools List Access Control
- Skill Workflow Orchestration
- Workflow Client Architecture
- Start Workflow
Service Identity, mTLS, And Deferred Revocation
Status: design decision; revocation implementation deferred until production need.
This document defines the intended role of light-identity-issuer for controlled,
long-running service workloads. It also records when Light Fabric should combine
mutual TLS (mTLS), an application token, and a user access token.
The Light CLI is deliberately outside this workload-certificate model. It is a public client that may be installed on intermittently connected machines, so it uses ordinary server-authenticated TLS and the signed-in user’s access token when calling the Gateway. It does not receive or present a client certificate.
Decisions
- Portal and Config Server are the future authority for workload-certificate revocation policy.
- Revocation distribution is not implemented yet. The current in-memory revocation list is not a production revocation mechanism.
- An expired workload certificate cannot renew. Controlled services must renew
before expiry, using the issuer-provided
renewAtdeadline. - Official environments should use mTLS on sensitive service-to-service hops where both endpoints are controlled by Light Fabric.
- A user-delegated service-to-service call should normally carry all three
independently verified proofs:
- mTLS identifies the connecting workload instance and protects the hop;
- the app token identifies and authorizes the immediate calling application;
- the user access token identifies the user whose authority is being delegated.
- A service-only operation carries mTLS and an app token, but must not invent a user identity when no user is involved.
- Browser and public CLI ingress uses server TLS and a user token. It does not require mTLS or a distributable application secret.
Why The Three Proofs Are Not Redundant
Each proof answers a different question:
| Proof | Question answered | Typical enforcement |
|---|---|---|
| mTLS certificate | Which deployed workload instance opened this connection? | CA chain, validity, environment, role, service ID, and optionally install ID |
| App token | Which immediate application is calling, and what application operations may it perform? | Signature, issuer, audience, expiry, token_use=app, sid, scopes, and route policy |
| User access token | Which user authorized the operation, and what may that user access? | Signature, issuer, audience, expiry, host/tenant, subject, roles, and resource ACL |
For a delegated request, authorization is the intersection of these identities. A valid user is not permission for an arbitrary service to act as that user. A valid application is not permission to act for an arbitrary user. A trusted certificate is not permission to call every route.
The receiver must bind the transport identity to the app identity. Accepting any certificate signed by the platform CA is insufficient. At minimum:
- the certificate SAN role must match the route’s expected role;
- the certificate SAN service ID must equal the app token
sid; - the certificate environment must match the receiving environment;
- the certificate must be valid and chained to the environment’s approved CA;
- the app and user tokens must independently satisfy their route contracts.
Where stronger instance binding is required, the app token can additionally carry a certificate confirmation claim or install ID. That is a later hardening step; service-ID binding is the required baseline.
Recommended Connection Profiles
| Connection | Transport and credentials | Reason |
|---|---|---|
| Browser to Gateway | Server TLS + user token | Browsers are public clients and cannot protect a platform client key |
| Light CLI to Gateway | Server TLS + user token | The CLI is public and intermittently connected; no client certificate |
| Gateway to internal API/MCP service for a user | mTLS + Gateway app token + user token | Authenticate the hop, immediate caller, and delegated user |
| Workflow or Agent to Gateway for a user | mTLS + workload app token + user token; action reference where required | Prevent a copied bearer token from changing workload origin |
| Service maintenance or control operation with no user | mTLS + app token | Do not manufacture a user leg |
| Health check carrying no authority | Network-restricted endpoint; server TLS where it crosses a host boundary | Keep privileged credentials off non-privileged probes |
| External third-party provider | Provider-approved TLS and authentication profile | Platform mTLS is not assumed to be supported externally |
This is a policy recommendation, not a requirement that every loopback call or every sidecar connection terminate its own mTLS session. A service mesh or local authenticated proxy may terminate mTLS, but the downstream service may trust the forwarded peer identity only over an authenticated internal channel that strips client-supplied identity headers.
Security Benefit Of mTLS
mTLS materially improves the service-to-service threat model when it is bound to the app identity.
What it adds
- Bearer-token replay resistance. A stolen app token alone is insufficient from a machine that does not hold an accepted workload private key.
- Workload provenance. The receiver authenticates the process or instance at the other end of the TLS connection, not merely a claim in an HTTP header.
- Defense in depth. A mistake in token routing, logging, or storage does not automatically become service impersonation.
- Per-install attribution. Unique leaves can identify which instance made a connection, improving incident response and audit evidence.
- Narrower network trust. Network reachability alone does not make a caller a trusted internal service.
- Independent credential rotation. App-token signing keys and workload CAs can rotate on different schedules and respond to different compromises.
What it does not add
- It does not protect a fully compromised workload host that can use the private key and tokens while they are available.
- It does not correct excessive app scopes, user permissions, or route-policy mistakes.
- It does not make a forwarded user bearer resource-specific. Every receiver must still validate the user token and its own ACL.
- It does not help when the receiver accepts a certificate for service A with an app token for service B. Explicit peer/app binding remains mandatory.
- It does not preserve end-to-end identity through an unauthenticated TLS terminator or proxy.
- It does not eliminate revocation and renewal operations. Poor certificate lifecycle management can create availability failures.
Costs And Risks Of mTLS
- A CA and issuing service become security-critical infrastructure.
- Every workload needs secure private-key generation, storage, rotation, and destruction.
- Renewal must complete before expiry; clock skew, issuer outages, and failed reloads must be observable and rehearsed.
- Load balancers, sidecars, and service meshes must preserve authenticated peer identity without trusting spoofable inbound headers.
- Separate environments need separate trust boundaries. Prefer a distinct CA per environment; if a CA is shared, the receiver must enforce the certificate’s environment attribute explicitly.
- Debugging is more involved because failures can occur during the TLS handshake before application logs or HTTP error responses exist.
- CA compromise has a large blast radius and requires CA rotation plus trust migration.
- Certificate issuance and renewal add operational load, although connection pooling keeps per-request TLS cost modest.
These costs are justified for controlled services that already have managed deployment, secret storage, monitoring, and continuous runtime. They are not justified for public clients such as Light CLI.
Security Without mTLS
A no-mTLS service profile uses server-authenticated TLS plus app and user tokens. It is simpler and can still be secure when all of the following hold:
- app tokens are short-lived, audience-restricted, and narrowly scoped;
- every receiver validates the immediate app identity and user identity;
- tokens are never placed in URLs or logs;
- credential-bearing redirects are disabled;
- network policy restricts service ingress;
- token rotation and revocation are reliable;
- sensitive services do not accept caller identity from unauthenticated headers.
The principal residual risk is bearer replay: anyone who steals a valid app token can use it from another reachable machine until it expires or is revoked. Network policy reduces that exposure but does not cryptographically bind the token to a workload. Proof-of-possession tokens can reduce the gap, but they introduce a key and lifecycle problem similar to mTLS.
For low-risk internal APIs, a short-lived app token over server TLS may be an acceptable operational trade-off. For Gateway, Workflow, Agent, Knowledge, and other services that forward user authority or perform privileged operations, mTLS plus app-token binding is the recommended official-environment profile.
Deferred Revocation Authority
Revocation policy will be owned by Portal and distributed by Config Server when
production requires it. Portal provides the authorization, approval, audit, and
event history. Config Server distributes the resulting immutable policy snapshot.
Neither service receives or uses the CA private key; certificate signing remains
the sole responsibility of light-identity-issuer.
The intended future flow is:
flowchart LR
O[Authorized Portal operator] --> E[Append-only revocation event]
E --> P[Revocation projection]
P --> C[Published Config Server snapshot]
C --> I[Issuer instances]
C --> G[Gateway instances, when immediate denial is required]
I --> D[Durable last-known-good cache]
Policy identity
A revocation entry should be scoped by at least:
- environment;
- issuing CA identity or generation;
- certificate install ID;
- revocation timestamp;
- non-secret reason/category;
- actor and policy revision in the Portal audit record.
The effective key is (environment, CA generation, install ID), not just a
service ID. Revoking one failed instance must not revoke every healthy instance
of the same service.
Renewal lineage and supersession
Revocation policy and renewal lineage solve different problems. Even before an install is
revoked, a copied still-valid certificate and key must not be able to fork an unlimited renewal
chain after the legitimate workload has rotated. A production issuer therefore also needs a
durable, atomic current-certificate record keyed by (environment, CA generation, install ID).
Successful renewal must compare-and-swap the presented certificate serial or fingerprint to the new certificate. Once that succeeds, every later attempt using the predecessor is refused even if the predecessor has not expired. The state must be shared consistently by every issuer replica and survive restart; an in-memory nonce cache is not sufficient. A race is intentionally first-writer wins and must emit enough audit evidence for an operator to recover the displaced legitimate workload by revoking the install and enrolling a new install ID.
This lineage store is deferred with the production revocation work. Until both exist, the issuer is appropriate for development and qualification, not as the final production credential authority.
Monotonic behavior
Revocation is not an ordinary replaceable configuration value. Once accepted, an older snapshot must not silently remove it. The implementation must:
- reject malformed, empty-by-accident, or revision-regressing updates;
- atomically apply validated updates;
- preserve accepted revocations in a durable last-known-good cache;
- load that cache before serving issuance or renewal after restart;
- avoid casual un-revocation; recovery should normally enroll a new install ID.
If deliberate un-revocation is ever supported, it requires a separate audited operation rather than snapshot rollback.
Availability behavior
If Config Server is temporarily unavailable, the issuer continues enforcing its durable last-known-good revocations. It must never replace them with an empty list. Production policy should define a maximum snapshot age; after that limit, the issuer should fail closed for renewal because it cannot prove it has current revocation information.
Issuer-only enforcement prevents a revoked installation from renewing. Its already-issued certificate remains usable until expiry. If production requires immediate denial, Gateway and other mTLS receivers must consume the same policy and reject the install ID during connection admission.
Deferred Implementation Boundary
No revocation distribution, Portal command, projection, Config Server property, Gateway enforcement, or local cache is authorized by this document. Those pieces are intentionally deferred.
Before implementation, define and approve:
- the Portal command/event schema and required operator roles;
- the immutable projection and Config Server property contract;
- revision and rollback rules;
- issuer cache format, locking, atomic replacement, and maximum staleness;
- whether issuer-only expiry-bounded enforcement is sufficient or Gateway must enforce install revocation immediately;
- CA and environment separation;
- metrics, alerts, audit evidence, and disaster-recovery exercises.
Qualification Requirements For A Future Rollout
- A revoked install cannot renew on any issuer replica.
- Restarting an issuer cannot forget a revocation.
- Config Server outage retains the last-known-good list.
- Empty, malformed, stale, and revision-regressing snapshots fail closed.
- Snapshot rollback cannot resurrect a revoked install.
- An expired certificate cannot renew.
- A certificate for one service or environment cannot be combined with another service’s app token or admitted in another environment.
- CA rotation supports an overlap period without weakening service/app binding.
- When immediate Gateway enforcement is enabled, an already-issued revoked leaf is denied on a new connection.
Recommendation
Use mTLS together with app tokens for privileged communication between managed, long-running Light Fabric services. Add the user token only when a service is acting on behalf of a user. Keep public clients on ordinary server TLS and user authentication.
This layered model is more secure than bearer tokens alone because compromise of one credential class is insufficient for service impersonation. The benefit is real only when certificate identity is bound to the app token, environments are isolated, certificates renew before expiry, and the operational lifecycle is treated as production infrastructure rather than static files copied at deploy time.
Light-Workflow
light-workflow is the workflow execution service for Agentic Workflow
documents.
It loads workflow definitions, executes workflow tasks, integrates with
light-rule for rule-backed checks, and exposes workflow execution APIs.
Key Dependencies
workflow-corelight-ruleaxumsqlxreqwest
Role
light-workflow is the runtime service that turns workflow specifications into
long-running execution state. It is used by agentic flows, human-in-the-loop
orchestration, and integration-test style automation.
Bounded Workflow Expressions — Evaluator Decision Pending
Status: proposed design for review, with acceptance criteria revised by the
owner on 2026-10-02 (see below). Extended CEL was selected at E02. The E03
production contract for profile cel-workflow-v2 (revision 3) is in
implementation/light-workflow/2026-10-02-ExpressionEvaluatorE03ProductionContract.md. It is summarized in 2.1, 2.3, 2.5, 2.7, 5.1 and 6; that contract
governs the details. No evaluator is selected and no runtime
support is implemented. Every numerical budget in this document is a proposal
pending review, not an established guarantee.
Tracking: light-fabric #431.
Source baseline: 900ddd1a4aa94540de272ecc03a47d3357bd3db5, plus the locally
uncommitted test-only G03 probes in apps/light-workflow/src/executor.rs.
Owner decision 2026-10-02: revised operating model
After reviewing the E01 CEL investigation, the owner revised the acceptance criteria for both candidates. This is an explicit owner decision made after results were seen, not an executor adjustment. It is recorded here and in the E01 plan, §13. The original E00 criteria are kept below and labelled historical. Where historical text conflicts with this section, this section governs.
- Operating model: workflow definitions are written by trusted users and reviewed before publishing. Arbitrary untrusted users cannot publish executable expressions. If that changes, strong isolation must be revisited before such publishing is allowed.
- Strong isolation is future hardening, not a prerequisite. Provable per-expression work and allocation bounds, and a separate evaluator process (3.3), are not required for expression support. Neither CEL nor jq is designed as a sandbox, and neither is required to become one. The product must not advertise hard per-expression isolation.
- Still required: capability isolation (2.9), correctness, the size limits in 4.3, predictable errors, and reviewed publication.
- Pathological expressions can still exhaust a worker’s CPU or memory. That is an operational limitation, contained by worker and container limits. An OOM can interrupt other work in the same process.
- Evidence already collected is preserved. The E01 CEL result stays “Fail against the original strict criteria”. Reassessment under these criteria is recorded separately.
Revised gates, identical for both candidates (full definitions in 4.4):
| Gate | Revised status |
|---|---|
| F1–F9, projection, input limits, capability isolation, result contract, existing CEL unchanged, effort and modification ceilings | Required, unchanged |
| Allocation peaks on realistic fixtures, performance | Reported, not pass/fail |
| Amplification outcomes, work/allocation enforcement arguments | Not executed in the current scope (scope reduction below); earlier results kept as history |
| Determinism | Required for outcomes the evaluator returns on ordinary fixtures and limit-boundary cases |
| Crash resistance | Unqualified. Documented limitation and future hardening item (scope reduction below). The earlier draft’s crash-free gate is withdrawn |
| JSON number boundary | New, required (2.3) |
| Error-category mapping | New, required (2.7) |
Scope reduction, 2026-10-02 (owner-approved)
Later the same day, the owner removed deliberate crash and resource-exhaustion testing from the current spike, equally for CEL and jq, under the trusted-author, reviewed-publication model.
- Not executed: deliberate crash, OOM, timeout and resource-amplification executions, including repeated guard stress probes. The A1–A10 cases (4.2) and the guard-based A-case runs are out of the current scope. A4’s compile-time limit checks are kept as ordinary validation tests (see below).
- Kept: ordinary correctness tests, realistic G03 fixtures (F1–F9, projection), result-contract cases, bounded validation tests for the configured limits (input-limit boundaries; source, AST and nesting limits), static source review, and benchmarks on realistic fixtures.
- Unqualified, not passed: crash resistance and resource isolation. No claim of crash resistance or strong resource isolation is made. Parser safeguards are assessed by source review and ordinary validation tests only.
- Known observation: the E01 CEL run had a stack overflow in a debug test build on a 2 MiB thread while compiling a large expression. The input was not kept. This is recorded as a limitation, not investigated further in the current scope.
- Historical results, including the earlier amplification and guard evidence, are preserved unchanged.
- Crash resistance and resource isolation are future hardening (5.1).
Purpose
Workflow definitions need generic JSON transformations over ordinary tool results: string splitting and extraction, field selection, array projection, JSON serialization and UTF-8 size accounting. These operations are reusable across API integrations and belong in the definition, not in provider-specific engine code.
The immediate consumer is the GitHub issue-to-design workflow (G03). Gateway continues to own authentication, ACL enforcement, upstream credentials and routing. The engine does not acquire a GitHub provider, capture records, receipts or integration-specific authorization.
Two candidates are compared on equal terms:
- Extended CEL: standard CEL extensions and bounded comprehensions added to the existing workflow value profile.
- In-process jq: a Rust jq implementation, for example
jaq, selected byevaluate.language: jq.
Neither outcome is predetermined. This document authorizes no spike execution, runtime change, publication or deployment. Approval of this revision is a documentation decision only.
1. Executable expression-field inventory
This inventory records how expressions execute at the baseline. Both candidates must account for every row. Line numbers drift, so rows cite functions.
1.1 Admission
validate_runtime_definition (runtime_definition.rs) accepts only an explicit
evaluate.language: cel. It rejects other languages and an omitted evaluate
block. It does not inspect evaluate.mode.
workflow-core defaults a missing language to jq during deserialization,
so evaluate: {} currently reaches validation as jq and is rejected. Any change
must preserve that rejection by checking the original definition, not the
defaulted model.
Admission performs no expression compilation for most fields. Invalid expressions are discovered at run time, where most of them fall back silently (see 1.3).
1.2 Expression sites
| DSL position | Executor entry | Context (. / CEL variables) | Result use |
|---|---|---|---|
set map values and set expression | resolve_json_value | run context | task output |
call: http endpoint URI | resolve_template_to_string | run context | string |
call: http endpoint {name} placeholders | rewritten to ${{ name }}, then template | run context | string |
call: http body | resolve_json_value | run context | JSON |
call: http query, headers | resolve_http_string_map | run context | strings |
idempotencyKey (HTTP, MCP, A2A) | resolve_template_to_string | run context | string |
call: jsonrpc/openrpc URI, params, headers | template / resolve_json_value | run context | JSON / strings |
call: mcp params, resource URI | resolve_json_value / template | run context | JSON / string |
call: a2a parameters | resolve_json_value | run context | JSON |
call: agent input, mockOutput | resolve_json_value | run context | JSON |
Agent instructions, prompt | resolve_template_to_string | run context | string |
ask assignment category, reason, assignee, role | resolve_template_to_string | run context | string |
switch case when (or the case name when when is absent) | evaluate_condition | run context | boolean |
assert value, equals, contains | resolve_json_value | run context | JSON |
assert.json comparison expression | evaluate_condition | the selected value, not the run context | boolean |
export map values | apply_exports | see 1.4 | context entries |
Workflow output.as string | direct evaluate_cel_value | final run context | public output object |
Workflow output.as object | resolve_json_value | final run context | public output object |
call output selectors (output: result | response | raw) and assert.json
paths are fixed selectors handled by lookup_json_path, not expressions.
The run context starts as the invocation input (process_info_t.context_data
is initialized from input_data). Task output enters the context only through
export.
1.3 Template and fallback layer
TEMPLATE_REGEX recognizes ${{ ... }} and ${ ... }. A whole-string
template keeps the JSON type of its result; embedded templates are stringified
into the surrounding string. The regex does not understand quotes or nested
braces.
evaluate_expression_to_value is not a pure CEL evaluator. In order, it:
- tries the CEL value profile when the expression does not start with
.; - evaluates as a predicate if a comparison operator appears in the text;
- treats
.a.bas a jq-style path lookup; - parses
true,falseandnull; - parses a quoted string;
- parses a number as f64;
- returns the expression text itself as a string.
A failed template is replaced by its original text. evaluate_condition
tries a CEL predicate, then splits on the first comparison operator and
compares with f64 arithmetic, and finally applies truthiness. A failed operand
becomes null and a failed expression becomes false.
Consequences:
- Existing “CEL” definitions may already contain jq-shaped
.pathexpressions. - A typo in a template produces literal text or
false, not a failure. - The f64 parse and the f64 comparison contradict the CEL value profile, which rejects doubles.
1.4 Export semantics
get_export_map reads export.as (or export) from the raw YAML into a
HashMap. apply_exports then, per entry:
.output→ the whole task output;.output.<path>→ a path lookup into the task output;- anything else →
evaluate_expression_to_valueagainst the context being built, which already contains previously applied exports.
Because HashMap iteration order is unspecified, an export that reads another
export key from the same map is nondeterministic. Failed lookups are skipped
silently. Exports cannot evaluate a general expression over the task output.
1.5 Parsed but not executed
light-workflow never reads task-level if, input, or output, and the
workflow-level input is not executed as a transformation. Admission does not
reject them, so a definition that uses them is accepted and the field is
ignored. This is an existing gap. New-profile admission must reject these
unsupported fields; implementing them is not required by this work. Existing
legacy definitions and runs retain their behavior. Broader legacy validation
changes require separate review.
run.script is accepted by admission and policy mapping. That does not
establish a qualified script evaluator, and neither candidate uses it.
1.6 Existing CEL value profile
evaluate_cel_value (crates/light-rule/src/engine.rs) compiles with the
standard profile, rejects comprehensions, caps the AST node count and converts
results through the workflow JSON profile. That profile rejects doubles, unsafe
integers, non-string map keys and opaque values on output. Inputs are added as
CEL variables, one per top-level context key that is a valid identifier. The
installed context lacks string splitting, substring extraction and JSON
serialization. Evaluation is wrapped in catch_unwind; there is no runtime
work or memory metering.
2. Shared contract, independent of evaluator
These rules apply to whichever candidate is selected.
2.1 Language selection
- One language per definition through
evaluate.language. No per-task selector, mixing or automatic detection. - An omitted
evaluateblock,evaluate: {}, an unknown language and any nonemptyevaluate.modeare rejected through every admission and publication path for the new profile. Existing legacy definitions and runs retain their current validation and execution behavior; this does not retroactively reject an existing CEL definition usingevaluate.mode. - If extended CEL is selected, it remains
language: cel, with a new evaluator profile (2.6). If jq is selected,language: jqis accepted only after implementation. - A CEL definition opts into the new profile only through an explicit selector
in the definition. Because legacy and new CEL definitions share
language: cel, the profile is never inferred from publication date, DSL version or any other implicit signal, and republishing an existing definition without the selector keeps the legacy profile. E03 selector:document.metadata.lightExpressionProfile: cel-workflow-v2, together withevaluate.language: celand noevaluate.mode. An absent selector means the implicit legacy profilecel-workflow-v1. An unknown or disabled identifier is rejected withEVALUATOR_PROFILE_UNSUPPORTED, never downgraded to legacy. The key is reserved: E04’s migration preflight aborts if a stored definition or snapshot already contains it.
2.2 Results
- A value position produces exactly one JSON value. JSON
nullis a value. For jq, zero results and two or more results fail; the evaluator consumes enough of the stream to detect a second result or a subsequent error. - A predicate produces exactly one boolean. There is no truthiness conversion.
- An embedded template must produce a string; other values are serialized explicitly in the expression.
- The new profile has no literal-text fallback, no
falsefallback and no silent truncation. Existing CEL definitions keep the 1.3 behavior unless an explicit migration is approved.
2.3 Numbers
Numbers are handled per stage, not by validating the whole context:
-
Input preservation: values the expression does not touch pass through unchanged, including decimals and large integers. An unrelated decimal in the context must not fail an expression.
-
Input conversion (revised 2026-10-02): an integer that fits
i64becomes the language’s signed integer (CELInt). An integer in(i64::MAX, u64::MAX]is preserved exactly (CELUIntis acceptable), or represented so that any operation reading it returns an error. Rejecting the whole input is not allowed (see input preservation above). A non-integer passes through as the language’s floating type. Conversion never truncates, wraps, rounds or reinterprets a number. -
Arithmetic (revised 2026-10-02): inside the language, arithmetic follows that language’s documented semantics. CEL integer division truncates toward zero and integer overflow is an error; jq division yields a float. These results are recorded, not treated as failures. Extension functions the profile adds use checked arithmetic and no truncating casts. Only silent corruption at the JSON boundary fails.
-
Output (production,
cel-workflow-v2, E03):- integers are emitted exactly over the full
i64/u64range; - finite doubles are emitted in shortest round-trip form;
- NaN, ±Infinity, non-string map keys and non-JSON values fail with the JSON-profile category.
Values are preserved as parsed, not as original text. Consumers that use binary64 numbers lose precision above 2^53; a public output schema protects them only if it constrains the range or declares such fields as strings. v2 adds checked
double(),int()anduint()conversions; implicit int/double mixing stays an error. See the E03 contract §4. - integers are emitted exactly over the full
-
E01 qualification output boundary (historical, not the product policy): “Integers in
[-9007199254740991, 9007199254740991]are emitted exactly. Integers outside that range, floats and non-string map keys fail with the invalid-JSON-profile category.” The spike evidence was collected under that boundary and is unchanged. -
Number precision before conversion: in E01, the shared harness parses JSON with
serde_jsonwithoutarbitrary_precision, so integers beyond theu64/i64range, and decimals, arrive asf64for both candidates. This is recorded as a harness property, not a candidate failure. Production integration must decide its own parse precision. -
Historical E00 arithmetic rule (superseded 2026-10-02): “operations on numbers outside the safe integer range
[-9007199254740991, 9007199254740991], or producing a fraction, fail rather than rounding. Whether decimal arithmetic is supported at all is a spike output.”
2.4 Scope and size
- Only ordinary workflow JSON crosses the evaluator boundary: no token caches, bearer tokens, LONG registrations, credentials, authority objects or service handles.
- The run context is converted into evaluator form once per task step and shared by that step’s expressions, not once per expression.
- The proposed spike input ceiling is 2 MiB of compact UTF-8 JSON, depth 64
and 100,000 JSON nodes, across context and task-output inputs together.
Measure a fixed harness envelope
{"context": ..., "taskOutput": ...}; count its keys and structure too. This is a test interface, not new DSL syntax.MAX_HTTP_RESPONSE_BYTESremains 1 MiB. The byte budget accommodates a 1 MiB response alongside a bounded context, but the node and depth limits apply independently and may still reject it: a dense 1 MiB response, such as an array of small objects or numbers, can exceed 100,000 nodes. The envelope does not promise to accommodate arbitrary accumulated context and does not raise any production limit. - Projection before accumulation is required: a new-profile export must be able to transform the separate task output while reading the pre-export context, then atomically merge only selected results into context. Raw task output must not be automatically copied into context by this operation. This does not change any separate task-output persistence contract.
- The spike tests this operation through a harness adapter with separate
context and task-output inputs. E03 chooses the actual language bindings and
export syntax after evaluator selection. Neither task-level
output.assupport nor a new DSL surface is a prerequisite for the spike. - Runtime integration must check the combined input before evaluator conversion, without first creating an unbounded copy. Existing production contexts above the new profile’s limit fail explicitly only when opting into that profile; legacy CEL behavior is unchanged.
2.5 Variable scopes and wrapper syntax
Resolved at E03 for extended CEL (E03 contract §2–§3):
- Bindings:
context,workflow.input,output(exports only) andvalue(assert.jsoncomparisons only). Context keys are not bound as top- level identifiers, so nothing can shadow a binding. - Delimiter: a single
${ … }, recognized by a scanner that understands CEL strings.$${writes a literal${, and${{is rejected. A string that is exactly one expression keeps its JSON type; expressions embedded in text must produce strings. Predicates and export strings must be expressions. - Exports: every entry reads one pre-export snapshot. They are merged together only if all succeed.
switch: every non-default case needswhen, anddefaultcomes last.- Endpoints:
{name}placeholders are allowed only in the URI path. The value comes fromcontext;.,..and empty values are rejected; the value is encoded once as a path segment and never rescanned. Endpoint URIs may not contain expression spans: this is a documented v2 restriction. Query expressions stay available throughwith.query, and legacy behavior is unchanged.
The original open questions, kept for reference:
- The task output in
exportmust be reachable without a reserved.outputkey that can shadow a context key. For jq, the modeled Serverless Workflow DSL passes runtime arguments as variables ($context,$input,$output,$task,$workflow), which avoids shadowing. The exact variable set must be checked against the DSL versionworkflow-coremodels and against the positions the executor actually runs (section 1), not assumed. ${{ ... }}is ambiguous with a jq object constructor (${{a: .x}}). One rule must be chosen, and the template scanner must understand quoted strings, escapes and nested braces. For CEL, map literals have the same issue.- Export entries must read one pre-export snapshot and be applied together after all succeed, removing the ordering dependency in 1.4.
- The
{name}endpoint placeholder rewrite generates expressions. Under jq,nameis not a valid path; the rewrite must emit syntax that is valid for the selected language.
2.6 Evaluator profile enforcement
Every definition admitted under the new contract records an evaluator profile
identifier (for example cel-workflow-v2 or jq-workflow-v1). Existing
definitions are treated as an implicit legacy CEL profile. Runs already keep
process_info_t.definition_snapshot; where the profile identifier lives is
decided at E04. Claim-time filtering likely needs the identifier in a
queryable column rather than inside the snapshot, which may require a
migration; E04 assesses this (section 5.1).
Enforcement policy:
- A worker keeps an implementation for every profile it advertises. Profiles are immutable: a behavior change is a new profile identifier.
- A worker that does not support a run’s profile does not claim that run’s tasks. Before admitting any new-profile definition, every task-claiming worker must understand profile filtering, or deployment must stop the old workers first. Today’s workers cannot be assumed to implement this rule.
- Admission rejects a profile not supported by the deployment’s explicitly
configured capability policy. Temporary absence of a compatible worker
does not prove a profile is unsupported and does not fail a persisted run.
Compatible tasks remain pending under existing operational timeout rules.
An explicit operator decision to retire a profile may fail affected runs
with non-retryable
EVALUATOR_PROFILE_UNSUPPORTED; runtime integration must identify the authority applying that decision. Never execute a different profile as a fallback. - Removing a profile requires that no active run uses it, or an explicit operator decision to fail those runs.
2.7 Errors and effects
-
Stable error categories: invalid expression, unsupported capability, invalid JSON profile, wrong result count or type, and resource-limit exhaustion.
-
Category mapping (2026-10-02). E01 keeps the spike’s frozen category set and tells cases apart by category plus phase:
Condition Category Phase Syntax or parse error; malformed macro invalid expression compile Function or feature outside the admitted subset unsupported capability compile Source, AST or nesting limit resource limit compile Runtime language error: type mismatch or no overload, missing field or key, index out of range, division or modulo by zero, integer overflow, invalid extension-function argument, jq error/1invalid expression evaluation Number/JSON boundary violation (2.3) invalid JSON profile conversion or evaluation Zero or multiple results; non-boolean predicate wrong result count or type evaluation Output bytes, nodes or depth resource limit evaluation When more than one condition applies, the evaluator returns the first one it meets in a deterministic traversal order (for example, ascending map-key order), never one that depends on hash-iteration order. Production integration (E04) adds a separate evaluation error category for runtime language errors, rather than reusing invalid expression.
-
Production categories (E03):
EXPRESSION_INVALID,EXPRESSION_UNSUPPORTED,EXPRESSION_LIMIT,EXPRESSION_EVALUATION,EXPRESSION_JSON_PROFILE,EXPRESSION_RESULT_TYPEandEVALUATOR_PROFILE_UNSUPPORTED, with the phase rules in the E03 contract §7.4. The table above remains the E01 spike mapping. -
Diagnostics identify definition, task, field and source offset. Admission errors may include an excerpt of the definition-authored expression. Runtime errors must not include context values, returned data or interpolated library errors.
-
Compile errors fail admission. Runtime expression errors fail the task through the existing failure machinery. They are not transient upstream errors and do not create their own retry loop.
-
All request arguments are resolved before dispatch; failure means zero dispatches. A failure after a successful external call does not roll that call back, and existing retry and idempotency contracts still apply.
2.8 Scheduling and leases
Evaluation is CPU-bound and synchronous. It must not run on Tokio worker
threads; it runs on spawn_blocking or a dedicated bounded pool. The executor
runs up to DEFAULT_HOST_EXECUTOR_CONCURRENCY (8) tasks with 30-second task
leases, so slot acquisition, cancellation and lease renewal must be designed
together:
- The total time a task can spend waiting for and holding evaluation slots must fit inside a renewable lease, or the lease must be renewed while waiting.
- A queue timeout alone does not prevent duplicate execution. A task whose lease is lost must not commit an evaluation result.
- Cancellation must stop evaluation and release its resources before releasing its slot. A watchdog is defense in depth, never the primary bound.
- Revised operating model (2026-10-02): an in-process evaluation cannot be
preempted unless the selected library offers metering. Cancellation and lease
loss are therefore honored at evaluation boundaries: a result produced after
either is never committed. A runaway evaluation is contained by worker and
container limits, not by a per-expression bound. Crash resistance is
unqualified (scope reduction above). As defense in depth, not
qualification, integration should compile and evaluate on threads whose
stack is larger than the 2 MiB default of Rust
std::threadand Tokio’s blocking pool, and record the chosen size.
2.9 Isolation and determinism
No environment, filesystem, network, module import, external input stream, debug output, clock or random access. The same expression, input and profile yields the same result or the same deterministic limit failure. Library defaults are audited, not only the functions the application calls.
Capability isolation remains a required gate. Resource isolation, meaning hard per-expression work and memory bounds, is future hardening under the 2026-10-02 operating model. Determinism applies to outcomes the evaluator returns. A process killed by an OS guard is recorded separately.
3. Candidates
3.1 Extended CEL
Scope under evaluation:
- Standard string extensions:
split,substring, and related functions from the cel-gostringsextension, implemented with the same signatures. map,filter,allandexistsmacros with a fixed nesting limit.- One JSON serialization function with pinned key ordering and escaping.
- UTF-8 byte sizing (already possible through
size(bytes(s))).
Questions the spike must answer:
- Can the
celcrate’s public API add these functions and macros without a fork? - A nesting limit and AST cap do not bound work by themselves. Work depends on list sizes, intermediate collections and string growth. Can a cost estimate computed from the AST and actual input sizes reject an expression before execution? If not, can execution be metered? (Since 2026-10-02 the answer is reported, not required.)
- Does widening the value profile change behavior for existing CEL definitions? Any change requires a new profile identifier (2.6).
3.2 In-process jq
Scope under evaluation: a pinned Rust jq library with a published function/operator inventory, unqualified features rejected at compilation.
Questions the spike must answer:
- Can the library meter work and live allocation during evaluation, including built-ins, through its public API? An async timeout cannot stop a CPU-bound evaluator, and a post-evaluation length check cannot prevent allocation exhaustion. (Since 2026-10-02 the answer is reported, not required.)
- Restricting the subset does not suffice by itself. Without recursion, ranges or loops, comma duplication, variable cross products and string repetition still grow exponentially or polynomially within the AST cap.
- Can number pass-through (2.3) be preserved?
3.3 Isolated evaluator process (future hardening, separately approved)
Under the 2026-10-02 operating model this is future hardening, tracked separately. Failing the resource gates no longer leads to it. It becomes necessary if untrusted users are ever allowed to publish expressions. Historical E00 wording: “If neither in-process candidate passes, a dedicated evaluator child process may provide stronger isolation.” This document does not approve it. A proposal would need its own contract covering:
- Memory:
RLIMIT_ASlimits virtual address space, which allocator reservations can trip; a cgroup memory limit is the likely mechanism. - Time:
RLIMIT_CPUis not a per-request wall-clock deadline; the parent must enforce the deadline and kill the child. - Bounded IPC framing, process pooling and reuse rules, cancellation and restart behavior, and filesystem/network confinement.
- Its own qualification, separate from the spike in section 4.
No jq CLI, shell or external executable is permitted in the two in-process
spike candidates. A dedicated evaluator worker executable belongs only to the
separately approved fallback above; it is not a general script runner.
4. Spike criteria, fixed before execution
These criteria are set before the spike starts and are not adjusted after results are known. Budgets are proposals pending review; review may change them only before the spike begins.
Exception, 2026-10-02: the owner explicitly revised the pass criteria after the CEL investigation (see the owner decision at the top of this document). The fixtures, amplification cases and limits in 4.1–4.3 are unchanged as definitions. The work and allocation ceilings are not enforced gates. After the scope reduction, the amplification cases (4.2) are not executed in the current spike. 4.4 gives the revised gates and keeps the original ones as historical.
4.1 Fixtures
The G03 fixtures are JSON files at three sizes: typical, page-maximum (30 comments of realistic length) and limit-sized (an input at the proposed input limit).
| ID | Transformation |
|---|---|
| F1 | Split a canonical GitHub issue URL into owner, repository and issue number |
| F2 | Project an issue to {title, body} with a nullable body |
| F3 | Project a comment page to [{author, body}] with nullable user |
| F4 | Concatenate comment pages into one list |
| F5 | Serialize the comment list to a JSON string |
| F6 | Measure the UTF-8 byte size of the serialized aggregate |
| F7 | Predicate: issue response contains pull_request |
| F8 | Predicate: a page has exactly 30 entries (continue pagination) |
| F9 | Pass through an unrelated decimal value in context while evaluating F2 |
Each candidate writes F1–F9 in its own syntax. A fixture passes when the result matches the expected JSON exactly.
Shared setup must freeze the exact JSON bytes, expected outputs, sizes and
SHA256 hashes in a manifest before either candidate starts. Typical and
page-maximum fixtures use identical inputs for both candidates. Limit-sized
fixtures use bounded unrelated padding where necessary so valid F1–F9 results
still fit the output ceiling; an oversized serialization result is an A5
rejection case, not a successful F5 fixture. Include a projection fixture with
a 1 MiB task response and a 48 KiB context, preserving an existing context key
named output, and prove only projected data enters the next context.
Limit-boundary fixtures test each input limit separately. Each set has one input exactly at the limit, which must be accepted, and one just over it, which must be rejected with the input-limit category:
| ID | Binding limit | Construction |
|---|---|---|
| L1 | Bytes (2 MiB) | Few large strings: under 100,000 nodes and depth 64 |
| L2 | Nodes (100,000) | Dense small values: well under 2 MiB and depth 64 |
| L3 | Depth (64) | Narrow nesting: well under 2 MiB and 100,000 nodes |
The manifest records, for every limit-sized fixture, its byte size, node count and depth, and which limit binds first. E00 approves these construction rules; shared setup materializes the files before candidate measurements. Later changes invalidate affected comparisons and require review before rerunning either candidate.
4.2 Amplification cases
Not executed in the current scope (scope reduction, 2026-10-02). The definitions are kept for future hardening, and earlier results are preserved.
Each case is written in the most damaging form each candidate’s admitted subset allows. Cases that cannot be expressed in a candidate are recorded as “not expressible”, with the compile-time rejection shown.
| ID | Case |
|---|---|
| A1 | Duplication chain: repeatedly double a list ([.[],.[]] in jq; list concatenation in CEL) |
| A2 | Cross product: nested iteration over the same large list |
| A3 | String growth: repeated concatenation or repetition of a long string |
| A4 | Deep nesting: deeply nested input, and construction of deeply nested output |
| A5 | Large serialization: serialize a limit-sized input, and serialize repeatedly |
| A6 | Unbounded constructs: recursion, ranges, loops and generators |
| A7 | Sorting, grouping and comparison over a limit-sized array |
| A8 | Number edge cases: unsafe integers, fractional division, overflow |
| A9 | Large result streams (jq) or large list construction (CEL) |
| A10 | Regex or pattern functions, if included in the admitted subset |
4.3 Proposed limits (pending review)
| Resource | Proposed ceiling per evaluation |
|---|---|
| Expression source | 16 KiB UTF-8 |
| Parsed expression | 2,048 AST nodes; nesting depth 64 |
| Input | 2 MiB combined envelope (2.4); depth 64; 100,000 JSON nodes |
| Work | 100,000 metered units, accepted only under the 4.3.2 conditions. Reported, not a gate, since 2026-10-02 |
| Live allocation | 16 MiB charged evaluator memory, under the 4.3.1 conditions. Reported, not a gate, since 2026-10-02 |
| Output | 1 MiB compact JSON; depth 64; 100,000 JSON nodes |
| Result count | Exactly one value |
| Compilation cache | 128 entries; 16 MiB retained per process |
| Concurrency | Bounded per worker; value decided with 2.8 |
These are evaluator ceilings, not increases to task, HTTP, agent or context limits. The lowest applicable limit wins, and definitions cannot raise them.
4.3.1 Allocation ceiling conditions
Since 2026-10-02 these conditions govern how allocation is measured and reported. A missing enforcement mechanism or bound is no longer a failure.
16 MiB is an experimental ceiling. It is evaluated only after both of these are fixed in the spike plan, before any measurement:
- Accounting: peak simultaneously live incremental bytes attributable to compilation, input conversion, evaluation and result serialization, including converted inputs, intermediate values, retained compiled expressions and output buffers. Freed bytes cease to count; this is not cumulative allocation traffic. Report three peaks separately: input conversion, compilation and evaluation, with cold and warm evaluation peaks reported separately. Warm measurements include the retained compiled expression and converted input even though allocation happened before timing. Each evaluation is charged only for its own compiled expression; the shared compilation cache budget (4.3) is measured and enforced separately. Report any excluded fixed library/harness overhead separately and bound it before measurements.
- Input size: the proposed 2 MiB envelope, depth and node bounds from 2.4.
If either is still open when measurement would start, the allocation criterion is reported as inconclusive for both candidates.
A counting allocator supplies measurement evidence, not production enforcement. Passing also requires a reviewable enforcement mechanism or conservative bound covering all admitted operations, including compilation and conversion. Passing the fixed tests alone is insufficient proof for arbitrary admitted expressions.
4.3.2 Work ceiling conditions
Since 2026-10-02 a metered definition or cost bound is reported as an extra capability when a candidate provides one. Its absence is not a failure.
A unit count is not comparable across engines. The 100,000-unit ceiling is accepted for a candidate only if that candidate provides one of:
- Metered accounting: a definition of what one unit charges, covering AST steps, built-in loops, sorting and comparison, string expansion and serialization, with evidence that every A1–A10 case is charged; or
- A conservative cost bound: a pre-execution bound computed from the AST and actual input sizes, shown to be an upper bound for every A1–A10 case.
Engines are compared on measured outcomes at the limit, not on unit counts: for each amplification case, the wall time and peak charged allocation at the point of rejection or completion.
4.4 Pass criteria
Revised pass criteria (2026-10-02, governing)
Both candidates are judged by the same gates:
- Compatibility: F1–F9 pass; the projection fixture holds; existing CEL qualification tests pass with the existing profile unchanged.
- Input limits: inputs exactly at each limit pass validation, and inputs just over are rejected with the input-limit category.
- Capability isolation: as in 2.9.
- Result contract: as in 2.2. Zero results, multiple results, or a value followed by an error never return a value. Predicates are strictly boolean. For jq, the stream is not consumed past the second result.
- Determinism of returned outcomes: for every limit-boundary case and every ordinary fixture that is repeated, the returned outcome and category agree across runs. Errors are chosen in a deterministic traversal order (2.7).
- Configured-limit validation: source, AST and nesting limits reject with the correct category at one unit over the limit and accept at the limit, shown by ordinary tests.
- JSON number boundary: as in 2.3.
- Error-category mapping: as in 2.7.
- Effort and modification ceilings: as in 4.5.
Reported, not pass/fail: allocation peaks per phase on realistic fixtures, and performance.
Unqualified in the current scope (scope reduction at the top): crash resistance, resource isolation, amplification behavior and work/allocation enforcement. Reports list these as excluded qualification, never as passed. If an ordinary test crashes, the crash is reported as a finding.
Each candidate’s report states its outcome under the revised criteria and, from the same evidence, under the historical criteria below.
Historical E00 pass criteria (superseded 2026-10-02)
Safety, enforceable limits:
- Every amplification case either completes within all limits or is rejected deterministically, at compile time or at run time, before any limit is exceeded. “Usually terminates” is a failure.
- Peak live allocation is measured with a counting allocator in the test harness and must stay within the allocation ceiling plus a documented, bounded fixed overhead.
- Repeated runs of the same case produce the same outcome and the same failure category.
Compatibility:
- F1–F9 pass.
- Existing CEL qualification tests still pass with the existing profile unchanged.
Performance, measured and reported, not an enforced limit:
- p50 and p99 latency per fixture at each size, over at least 1,000
evaluations on a recorded machine, reported in three separate measurements:
- Cold compilation: parsing, validation and compilation of an expression not in the cache.
- Context conversion: converting the run context into evaluator form, once per task step (2.4).
- Warm evaluation: evaluating an already-compiled expression against already-converted input.
- Provisional target, pending review: warm-evaluation p99 under 5 ms per F1–F9 fixture at page-maximum size. The target informs review; meeting it does not approve a candidate, and missing it is a finding, not a safety failure.
4.5 Effort and modification ceilings
Time:
- Time is counted in active engineering hours, not calendar days. One working day is 6 active hours.
- Each candidate gets the same budget: proposed 18 active hours (3 days).
- The whole spike, including shared setup and the report, is time-boxed to a proposed 42 active hours (7 days).
- Shared setup is done once, before either candidate’s clock starts, and is charged only to the overall time-box: fixtures, amplification inputs, the counting allocator, the benchmark harness and the report template.
- Candidate-specific setup is charged to that candidate: adding and building the dependency, license review, and learning its API.
Code:
- Each candidate may use only the library’s public API. A vendored fork, a
patched dependency or a
[patch]override exceeds the ceiling. - Implementation code is limited to a proposed 800 non-test lines per candidate, counted as non-blank, non-comment lines. The count includes helper crates, build scripts, macros and generated implementation code, whether committed or produced by a build script. It excludes tests, fixtures and the shared harness.
- Exceeding a ceiling ends that candidate’s investigation with the result “exceeds modification ceiling”.
4.6 Outcomes
Each candidate ends as pass, fail or inconclusive, with the following precedence:
- Fail: a required pass/fail criterion has failed with evidence. A demonstrated failure remains a failure even if the time-box expires before other criteria are answered. List those criteria as unanswered; do not infer their results.
- Pass: every required pass/fail criterion has passed with evidence.
- Inconclusive: no required criterion has demonstrably failed, but required evidence remains missing when the time-box expires.
The measured latency target is advisory and does not determine pass or fail. This precedence matches the E01 implementation plan’s section 8; it does not change any fixture, resource limit or candidate budget.
Rows are read top to bottom; the first matching row applies.
| Result | Next step |
|---|---|
| One passes; the other is inconclusive | Report the qualified pass and what remains unanswered for the other. The owner may select the passing candidate or approve extending the other investigation. Inconclusive is not treated as failure, and the report claims no comparative superiority |
| Both inconclusive | Report what remains unanswered; no selection recommendation; extending the spike requires approval |
| One inconclusive; the other fails | Report evidence for both; no selection recommendation; extending the inconclusive investigation or proposing the fallback (3.3) requires approval |
| Only CEL passes; jq fails | Propose extended CEL to unblock G03; jq remains a separate track |
| Only jq passes; CEL fails | Propose jq integration under the shared contract |
| Both pass | Owner decides on effort, compatibility and measured performance; the report does not choose |
| Both fail | Report evidence; the owner decides the next step. Since 2026-10-02 this does not imply the isolated process (3.3), which is future hardening. Historical wording: “the isolated process (3.3) may be proposed for separate approval”. |
Since 2026-10-02, this table is applied to the outcomes under the revised criteria (4.4). Outcomes under the historical criteria are reported as history and do not choose the row.
The report records for each candidate: dependency name, version and license; admitted function inventory; enforced limits and how each is enforced, including the work accounting or cost bound (4.3.2); results for F1–F9 and A1–A10 with wall time and peak charged allocation; the three latency measurements (4.4); active hours used; and implementation line count.
5. Checkpoints
| Checkpoint | Deliverable and review gate |
|---|---|
| E00 | Approve this revision, including section 1 and the section 4 criteria |
| E01 | Run the spike under section 4; deliver the report; no runtime integration |
| E02 | Owner selects a candidate (or the fallback proposal) from the report |
| E03 | Resolve 2.5 for the selected candidate; add the complete YAML example (section 6) |
| E04 | Integrate admission, dispatch, profile enforcement, errors, scheduling and publication/digest compatibility; qualify CEL regressions |
| E05 | Review executor and hostile-input evidence, upgrade behavior and rollout instructions |
| G03 resume | After separate approval, implement and qualify the G03 definition |
Stop for owner review at each checkpoint. Keep the test-only G03 probes and unrelated workspace changes intact.
5.1 Open items
| Item | Resolved at | Requirement |
|---|---|---|
| CEL profile selector | E03 (resolved) | document.metadata.lightExpressionProfile: cel-workflow-v2 (2.1) |
| Profile storage | E04 | E03 contract §7.3 proposes: • process_info_t.expression_profile with a snapshot-consistency CHECK• a database-held v2 admission switch • claim function v2, a legacy-only predicate on claim v1, and profile checks on every execution path The schema alone doesn’t protect paths in already-running old binaries. Rollout order: install the schema → upgrade or stop every admitting, evaluating or completing binary → confirm the capability inventory → enable admission, with reserved-key preflights before and after and no automatic rewriting. Rollback checks are split: the database checks the switch and active v2 runs and keeps the legacy-only protection; the deployment verifies every admission writer honors the switch. SQL alone doesn’t prove application admission is disabled |
| Compiled-expression accounting | E04 | Charge each evaluation for its own compiled expression; enforce the shared compilation cache budget separately (4.3.1) |
| Memory reporting | E04 | Report input-conversion peak memory separately from compilation and evaluation peaks, as in the spike (4.3.1) |
| Runtime error category | E04 | Implement EXPRESSION_EVALUATION and the other E03 categories (2.7) |
| Strong isolation | Future hardening | Per-expression work/memory bounds or an isolated process (3.3). Required before untrusted users may publish expressions |
| Crash resistance | Future hardening | Qualify compiler/evaluator stack safety (recursion bounded before each recursive phase, a defined thread stack) and amplification behavior. Unqualified in E01; see the scope reduction at the top |
6. Complete workflow example
E03 provides one complete, neutral cel-workflow-v2 example (E03 contract §6).
It is labelled as proposed and not executable until E04. It shows workflow
input and schema, set, HTTP arguments with path placeholders, switch,
atomic exports and workflow output.as, with explicit page-request,
aggregate-count and serialized-size bounds enforced in the definition.
It is deliberately not a G03 definition. G03 still has to supply GitHub’s
canonical owner and repository validation (rejecting . and ..) and the
registered lightapi://<capabilityRef> calling pattern with its Portal-issued
metadata.workflowTool pin, after E04.
7. Integration requirements for the selected candidate
- One internal dispatch boundary for all section 1 positions, with distinct
value, string and predicate result contracts. Direct calls, such as the
workflow
output.aspath, must not bypass it. - Definition validation, Portal publication, native start and Invoke admission agree on the supported language and profile. All known expressions are compiled at admission. Runtime checks remain for input-dependent types and resource use.
- The definition digest includes the language selector and expression source through its normal canonical representation. Earlier definitions are not rewritten and their identities are not recomputed.
- No changes to Gateway ACLs, ToolBinding rules, creator identity, task leases beyond 2.8, cancellation, deadlines or LONG token semantics. Local transformations do not trigger token exchange.
Operational requirements
- Metrics for evaluations, latency and limit failures by category.
- An author validation path: compile and evaluate an expression against a supplied sample input without starting a run.
- A versioned process for adding built-ins: each addition is qualified against the amplification cases and produces a new profile identifier (2.6).
- Operator documentation states the operational limitation: under the 2026-10-02 operating model, a pathological expression in a reviewed definition can exhaust a worker’s CPU or memory. Containment is worker and container limits plus definition review. Hard per-expression isolation is not advertised.
- Crash resistance is unqualified. Operator documentation says so, and evaluation threads use a recorded stack size larger than the 2 MiB default as defense in depth.
8. Required qualification after integration
- Language admission: legacy CEL unchanged; the new profile accepted; missing block or language, unknown language and nonempty mode rejected on every new-profile admission and publication path without retroactive legacy validation changes.
- Syntax and scope: every section 1 position, the wrapper rule, quotes and braces, literal strings and export semantics. Unsupported fields from 1.5 fail rather than being ignored.
- Results: null, result count, strict predicates, serialization, Unicode byte counts and number handling per 2.3.
- Hostile input: deferred to the crash-resistance and strong-isolation future hardening items (5.1). Integration qualifies the configured-limit validation and deterministic returned outcomes only. Historical E00 wording: “asserting bounded termination and memory, not only an error string”.
- Isolation: attempted environment, file, network, module, input-stream and clock access; sentinel secrets absent from evaluator input and diagnostics.
- Executor: Set, HTTP arguments, switch, assertions, exports, agent arguments and public output; invalid arguments cause zero dispatches; restart and resume; lease loss during evaluation; profile enforcement across a rolling upgrade.
- Publication and digest: language or expression changes change the definition identity; native and Tool-bound execution select the same evaluator.
Component tests and simulated fixtures are component evidence. Live Portal publication, deployed execution and a real coding-agent handoff are reported separately. No GitHub credential is needed for evaluator qualification.
9. G03 follow-through
The expression capability removes a transformation gap; it does not complete
the integration. The G03 definition still needs canonical URL checks,
registered Gateway HTTP calls, body-based pagination at per_page=30, an extra
empty request for exact page multiples, explicit aggregate and page bounds,
fixed-delay task retries and rejection of issue responses containing
pull_request.
Issue and comment text remains untrusted content in ordinary context. Attachment links are metadata unless a later definition explicitly retrieves them. The existing coding-agent entry requires valid workspace and task setup and its existing authority; the expression capability does not create those inputs. The real design handoff is qualified directly, rather than treating a prepared JSON payload as successful execution.
Related documentation
- Workflow Invoke and Tool Binding Publication
- Native Agent Call
- Personal Development Workflow Orchestration
Start Workflow (Archived Instructions)
Use the native workflow_start MCP tool through light-gateway for
asynchronous root Workflow starts. It rejects stableToolRef: a
workflow-backed Tool must enter through its Gateway Tool and the
Gateway-internal workflow_invoke operation. workflow_start accepts the
optional expectedDefinitionDigest to ensure that it starts the definition
revision acknowledged by Workflow. The Workflow Editor uses the Portal
StartWorkflow command, which saves and acknowledges the current definition
before calling workflow_start. See
Workflow Invoke And Tool Binding Publication.
The remaining examples in this archived guide describe the former Portal event path. They are retained for historical context and must not be used to start a current Workflow.
Runtime Path
The local start flow is:
- Create or update a workflow definition in
light-portal. - Start the workflow with the
workflowservicestartWorkflowcommand. workflow-commandwrites a workflow started event into the event store and outbox tables.light-workflowpolls the same database, loads the definition bywfDefId, creates the process and task records, and executes the workflow.
For this reason, the DATABASE_URL used by light-workflow must point to the
same database used by the local portal stack.
Prerequisites
Start the local portal stack first. For the Rust local stack, use the normal
portal-config-local deployment command from the portal-config-loc checkout:
./scripts/deploy-local.sh pg rust
Make sure the workflow command and query services are available in that stack.
The workflow definition pages in portal-view depend on those services.
Then build light-workflow:
cd /home/steve/workspace/light-fabric/apps/light-workflow
cargo build -p light-workflow --locked
Start light-workflow Locally
Create light-workflow.env in
/home/steve/workspace/light-fabric/apps/light-workflow:
DATABASE_URL=postgres://postgres:secret@localhost:5432/configserver
LIGHT_PORTAL_AUTHORIZATION="Bearer <workflow-service-token>"
SERVER_ENVIRONMENT=dev
LIGHT_WORKFLOW_CONFIG_MODE=local
RUST_LOG=light_workflow=debug,info
WORKFLOW_LOG_ANSI=false
Start the service with the debug binary:
./run.sh --debug-binary
The script loads light-workflow.env automatically. If you do not use the env
file, export the values before running the script:
export DATABASE_URL=postgres://postgres:secret@localhost:5432/configserver
export LIGHT_PORTAL_AUTHORIZATION="Bearer <workflow-service-token>"
export SERVER_ENVIRONMENT=dev
export LIGHT_WORKFLOW_CONFIG_MODE=local
export RUST_LOG=light_workflow=debug,info
export WORKFLOW_LOG_ANSI=false
./run.sh --debug-binary
Do not set the variables on separate shell lines without export. That creates
shell variables only for the current shell and run.sh will not receive them.
Recommended UI Test
The easiest local test is to create the definition in the portal UI and start it from the workflow editor test action.
-
Open
light-portal. -
Go to the workflow definition page.
-
Create a workflow definition.
-
Paste one of the example workflow YAML files from:
/home/steve/workspace/light-fabric/apps/light-workflow/examples -
Save the definition.
-
Open the definition in the workflow editor.
-
Use the editor test run action with a JSON input object.
For the basic example, use
apps/light-workflow/examples/simple-set-assert.yaml and this input:
{
"applicantId": "APP-001"
}
The editor test action is preferred for local testing because it parses the
input text as JSON and sends input as an object.
The table run button opens the generic startWorkflow form. If using that path,
make sure the request sends input as a JSON object, not as a string. If the
input is submitted as a string, the workflow command may accept the request but
the runtime context will not have the expected object fields.
Start with Postman or curl
You can also start the workflow directly through the portal command endpoint.
Send the request to the same light-gateway or light-portal host used by the UI.
Do not send this request to light-workflow; light-workflow is the executor,
not the command API.
The command envelope is:
{
"host": "lightapi.net",
"service": "workflow",
"action": "startWorkflow",
"version": "0.1.0",
"data": {
"hostId": "<host-id>",
"wfDefId": "<workflow-definition-id>",
"input": {
"applicantId": "APP-001"
}
}
}
Example curl shape:
curl -k -X POST "https://localhost:8443/portal/command" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <access-token>" \
-d '{
"host": "lightapi.net",
"service": "workflow",
"action": "startWorkflow",
"version": "0.1.0",
"data": {
"hostId": "<host-id>",
"wfDefId": "<workflow-definition-id>",
"input": {
"applicantId": "APP-001"
}
}
}'
If your local UI uses a session cookie instead of a bearer token, use Postman with the same authenticated session or copy the current local authorization header from the browser request.
Creating the Definition by API
For most local tests, create the definition in the UI. It is easier because the YAML can be pasted directly.
If you create the definition through the command API, send a workflow definition
command first and use the returned definition id as wfDefId in the
startWorkflow command.
The command shape is:
{
"host": "lightapi.net",
"service": "workflow",
"action": "createWfDefinition",
"version": "0.1.0",
"data": {
"hostId": "<host-id>",
"namespace": "light-portal",
"name": "simple-set-assert",
"version": "1.0.0",
"definition": "<workflow-yaml-as-json-string>"
}
}
When calling this from Postman, remember that the YAML definition is a JSON string field. Newlines must be escaped correctly by the JSON editor or sent by a tool that can build the JSON body safely.
Example Workflows
The current examples are in
/home/steve/workspace/light-fabric/apps/light-workflow/examples:
| File | Purpose | Input |
|---|---|---|
simple-set-assert.yaml | Basic local smoke test with no external dependency. | { "applicantId": "APP-001" } |
http-risk-decision.yaml | Calls a risk evaluation HTTP endpoint and branches on the result. | { "applicantId": "APP-001", "loanAmount": 25000, "creditScore": 720 } |
human-approval.yaml | Creates a human approval style workflow and waits for a later decision. | { "requestId": "REQ-001", "summary": "Approve test request" } |
insurance-claim-rest-v1.yaml | Complete product demo with direct HTTP API orchestration, native agent tasks, and human tasks. | See examples/README.md. |
insurance-claim-mcp-v1.yaml | Complete product demo with gateway MCP tool orchestration, native agent tasks, and human tasks. | See examples/README.md. |
insurance-claim-headless-v1.yaml | Headless insurance-claim regression workflow with deterministic agent outputs and no human-task pauses. | See examples/README.md. |
Start with simple-set-assert.yaml. It is the best smoke test because it does
not require another service.
For a complete multi-agent product demo, use the insurance claim suite in
apps/light-workflow/examples. The product walkthrough is
Insurance Claim Agentic Workflow, and
the operational runbook is in apps/light-workflow/examples/README.md.
For http-risk-decision.yaml, start a local mock service for the URL used by
the definition. When light-workflow runs natively with run.sh,
127.0.0.1 means the host machine. When light-workflow runs in Docker,
127.0.0.1 means the container itself, so change the workflow endpoint to a
Compose service name or host.docker.internal.
For human-approval.yaml, the first run should create a waiting task. Completing
that flow requires the worklist or task-completion API path.
Verify Execution
Watch the light-workflow log after sending startWorkflow. A successful run
should show that the start event was received, the first task was initialized,
and the executor picked up task work.
Useful database checks:
select wf_def_id, namespace, name, version
from wf_definition_t
order by update_ts desc
limit 5;
select process_id, wf_instance_id, status_code, context_data
from process_info_t
order by started_ts desc
limit 5;
select wf_task_id, task_type, status_code, task_output
from task_info_t
order by started_ts desc
limit 10;
select c_offset, event_type, aggregate_id, payload
from outbox_message_t
order by c_offset desc
limit 10;
If outbox_message_t has the workflow started event but no process or task
records appear, check that light-workflow is running against the same
DATABASE_URL as the portal stack.
Troubleshooting
DATABASE_URL is required: PutDATABASE_URLinlight-workflow.env, export it before runningrun.sh, or put the assignment on the same command line as./run.sh.function make_interval(mins => bigint) does not exist: Rebuild and restartlight-workflow. The runtime query must cast the retry value tointbefore passing it tomake_interval.- Workflow definition list is empty in the UI: Confirm the workflow query service is running and the local stack is using the jar or binary that contains the workflow definition owner-scope fix. Some local stacks run copied service artifacts, so rebuilding a source checkout is not enough unless the deployed artifact is refreshed.
- No tasks are created after starting the workflow: Confirm the
startWorkflowcommand wrote a workflow started event to the outbox table, and confirmlight-workflowpoints to that same database. - The workflow input is missing fields: Confirm
inputwas submitted as a JSON object. A string that contains JSON text is not the same as a JSON object in the workflow context.
Workflow Invoke And Tool Binding Publication
Status: Component implementation for issue #415; live qualification remains unverified. Revised after the 2026-09-26 plan review (R1–R13).
This document defines two related changes to workflow-backed MCP Tools:
- A new native
workflow_invokeMCP operation onlight-workflow. The Gateway calls it for every workflow-backed Tool. It admits the call against a Workflow-owned, owner-approved binding revision, waits for completion within the binding’s deadline, and returns the output or a clear timeout or failure result. - Publication of workflow definitions, grants and Tool bindings to
light-workflowthrough Workflow MCP operations, replacing the direct database sync.
The Gateway is the only caller of workflow_invoke. An Agent that needs a
workflow calls the workflow-backed MCP Tool on the Gateway like any other
Tool. A skill can guide the Agent to that Tool through skill_tool_t, but it
never calls the workflow directly.
It supersedes the Workflow Start and Invocation API section of
Workflow-Backed MCP Tools for
workflow-backed Tools. workflow_start remains the entry point for editors,
schedulers and other asynchronous callers, as described in
Start Workflow.
The Portal StartWorkflow command initializes and acknowledges both the saved
definition and its current grant set before it calls workflow_start. It
rechecks the acknowledged definition revision and digest after grant delivery
and passes that digest as expectedDefinitionDigest. Workflow still fences the
definition digest and enforces the live grants at admission and execution.
The request body’s idempotencyKey takes precedence over the
Idempotency-Key header. If neither is supplied, each Start request generates
a fresh key and represents new work, even with identical workflow input. A
caller retrying one intended start must reuse an explicit key.
The Workflow Editor attempts synchronization when a user saves or publishes. Portal keeps the desired definition and grant revisions alongside Workflow’s acknowledged revisions. If delivery fails after a local save, the editor shows both revision pairs and the last error; Sync now retries with the current user’s bearer. If the status query also fails, the editor retains an unconfirmed-sync warning and keeps Sync now available until a later status read confirms both revisions. There is no periodic synchronization job. Start may also sync pending revisions during the authenticated user request, and runs only after the required revision is acknowledged.
Operator setup
The local stack uses the main all-in-lt/docker-compose.yml. Provision the
Workflow run credential keyring through its idempotent runtime secrets init
service; WORKFLOW_LONG_KEYRING_FILE points to
/run/secrets/run-credential-keyring.json. Workflow refuses invoke admission
when a usable sealing key is unavailable. Configure
the Gateway application identity accepted by Workflow. All Portal publication,
definition and grant synchronization calls use the acting user’s bearer in Authorization;
Gateway supplies its own application bearer in X-Scope-Token to Workflow.
At the Gateway entry, Authorization must contain a valid user bearer. If a
caller supplies X-Scope-Token, Gateway independently verifies it as an app
bearer; a missing scope header is allowed, and an invalid one is refused.
The scope token never substitutes for user authentication. The same user role
and MCP Tool access rules apply in either case. This path requires no mTLS.
JWT-only Gateway-to-Workflow operation is supported; mTLS is optional future
setup.
In Portal Rule Admin (/app/rule/admin), assign the established user role rule
to workflow_definition_save and workflow_definition_grants_sync in MCP
Gateway Setup (/app/mcp/setup) → Access Control. Allow admin, host-admin,
and workflow-admin for these two endpoints. An app bearer in Authorization
must be denied even if it carries a user-like claim.
On the other Workflow endpoint cards, add the established role-based rule and
assign roles: definition publish/retire to admin, host-admin,
workflow-admin, genai-admin; binding publish/retire to admin,
host-admin, genai-admin; binding get/list/decide/revoke to admin,
host-admin, workflow-admin, genai-admin. In Role Permission
(/app/access/rolePermission), assign the Portal commands
decideWorkflowToolBinding, revokeWorkflowToolBinding, and
refreshWorkflowToolBindingStatus to those four roles. Assign
retireWorkflowToolBinding to admin, host-admin, genai-admin.
Keep retryWorkflowOperation under its existing command ACL and original-user
check. Check each endpoint ID before saving.
Definition and grant synchronization travels from Portal through Gateway MCP
to Workflow’s authenticated publication operations. The Portal needs no direct
connection to the workflow-ops database. The Gateway keeps
workflow_invoke internal: it is absent from client tools/list and direct
client tools/call is refused. Gateway still applies the published Tool ACL
before its internal call.
Historical problem
The issue #415 checkpoint unified every root start on the native
workflow_start MCP tool. That is right for editor and scheduler starts, but
workflow-backed Tools lost the binding contract on the way:
- Permanent idempotent replay. The Gateway passes the derived idempotency
key to
workflow_start, which hard-codesresult_replay_untilto the maximum timestamp. A binding withresultReplayMs: 0(the normalizer default) now replays the first result forever, so a read-only Tool called twice with the same input by the same user never sees fresh data. - Binding runtime fields are dropped.
workflow_starthard-codes the async mode, the interactive class, a 30-day deadline,permit_depth: 0and no parent action. The binding’stotalDeadlineMs,executionClass,runtimeBoundsanddelegationPolicynever reach admission, and a nested Tool call from a running workflow becomes an unlinked root. - Workflow-backed admission is skipped.
workflow_startadmits with thePortalExecutionprofile. TheWorkflowBackedchecks (verify_binding, orchestration and pinned dependency validation, approval evidence, deadline-aware admission) do not run. - Sync waiting lives in the Gateway. The Gateway loops on
workflow_wait_result, which returns null for failed or cancelled runs, so a terminal failure is indistinguishable from a slow run until the Gateway deadline. Lifecycle tool errors are then collapsed toWORKFLOW_INVOCATION_UNAVAILABLE. - Workflow data reaches Workflow by database access. The development
workflow-projection-syncCompose service runspublish-workflow-projections.shevery 30 seconds. It reads Portal tables overpostgres_fdwand writes Workflow’sworkflow_opsschema directly: saved definitions, bindings, dependencies, endpoint targets and grants. In production the Workflow runtime may belong to another team, and the Portal cannot access its operations database. The sync is also asynchronous to Gateway publication, so the Gateway can serve a binding that Workflow has not seen yet, or the reverse. - Binding rows are mutable and read live. An admitted run reads the
active binding, its dependencies and its endpoint targets again at each task
(
executor.rsgrant and endpoint reads,bound_mcp.rsdependency join). A republish or revocation during a run changes what the run may reach.
Goals
- Workflow-backed Tools are admitted by Workflow from a Workflow-owned, immutable binding revision, not from Gateway-supplied runtime fields.
- An admitted run keeps the revision it was admitted with until it ends.
- One MCP call returns the output, a terminal failure, or a timeout that identifies the running instance.
- Portal writes definitions, grants and bindings only through Workflow MCP operations, and receives receipts it can pin in the Gateway configuration.
- The Gateway snapshot serves a Tool only when Workflow has an active revision with the same digests, so the two sides cannot skew silently.
Non-Goals
- Changing
workflow_startsemantics for editor, scheduler or API callers, other than the saved-definition digest pin (see Saved definitions). - Changing the workflow definition language or task execution.
- Agents calling
workflow_invokedirectly. They use the workflow-backed Tool, so there is one access-control and admission path. - Nested workflow-backed Tool calls. They are refused until the mTLS action dispatch is restored (see Parent action).
- Running publication itself as a workflow. Publication records a pending revision and the owner decides separately (see Workflow owner approval).
- Backward compatibility with the projection sync. It is removed, not adapted.
Tool Binding Model
A Tool binding connects one MCP Tool (tool_t) to one published workflow
definition version. The Portal creates it when a Tool is saved with
executionPlacement=workflow and a workflowVersionRef.
WorkflowToolBindingNormalizer fills defaults and rejects unsafe
combinations. The Portal stores the author’s binding in its
workflow_tool_binding_t; Workflow stores each published version of it as an
immutable revision (see Binding revisions).
| Field | Meaning | Normalizer default |
|---|---|---|
bindingId, toolId, toolName | Portal binding identity and the stable Tool reference | from the Tool |
wfDefId, workflowVersion | Pinned definition version | from workflowVersionRef |
definitionDigest, schemaDigest | Hashes of the definition and its input schema | computed |
invocationMode | Always sync for Tools (see Sync only) | sync |
syncWaitMs | How long the caller waits for output | 20000, max 20000 for sync |
totalDeadlineMs | Hard deadline for the run | 30000, max 30000 for sync |
executionClass | Scheduling class | interactive for sync |
resultTextMode | How output is rendered to MCP text | compact-json |
cancellationPolicy (new) | What happens to a run past its deadline | before-effects-only |
idempotencyPolicy | Key kind and replay windows | kind derived, resultReplayMs 0 for read-only Tools |
delegationPolicy | Nested call rules | maximumDelegationDepth 1 |
runtimeBounds | Attempts, nested calls, parallelism, bytes, cost units | 8, 8, 1, 1 MiB / 4 MiB / 1 MiB, 1000 |
admissionLimits (new) | Concurrent runs and start rate, total and per user | see Concurrency and rate limits |
callerPolicy (new) | Roles a user needs to reach the workflow through this Tool | empty: the Tool ACL decides |
toolAnnotations (new) | The Tool’s readOnly and destructive flags | from the Tool |
policyDigest, responsePolicyDigest | Hashes of the admission and response profiles | computed |
Binding cancellationPolicy accepts exactly before-effects-only, cooperative,
and disabled in publication JSON and the Portal and Workflow binding tables.
The default is before-effects-only. Admission explicitly maps these strings
to CancellationPolicy::BeforeEffectsOnly, Cooperative, and Disabled.
The existing invocation enum serialization and invocation-row storage remain
BEFORE_EFFECTS_ONLY, COOPERATIVE, and DISABLED; the new binding contract
does not accept those uppercase spellings.
The database enforces total_deadline_ms >= sync_wait_ms and that sync mode
uses the interactive class.
Related records travel with the binding:
- Dependencies (
workflow_tool_dependency_t): nested Tools that the workflow may call, with contract digest, authorization key and dispatch target. - Endpoint targets (
workflow_endpoint_target_t): the resolved HTTP endpoints that tasks may call, with allowed methods and authorization policy digest. - Task approval evidence (
workflow_tool_approval_evidence_t): one row per write task, generated by Workflow when the owner approves the revision (see Task evidence).
Grants (workflow_tool_grant_t) are not part of the binding. They belong to
the workflow definition: the Tool owner grants a workflow access to a Tool
(see Grants).
The binding has two consumers:
- The Gateway needs just enough to list and route the Tool: name, input
schema, the pins and digests, the mode, and the wait time for its HTTP
timeout. It gets these from
mcp-router.yml. - Workflow needs the full revision to admit a call. It gets it from publication (below) and treats it as authoritative.
Sync only
Every workflow-backed Tool is synchronous. The caller waits for the output
within one MCP call, and the whole run fits inside totalDeadlineMs (at most
30 seconds). The normalizer rejects invocationMode: async, and the async
Tool behaviour in
Workflow-Backed MCP Tools
is out of scope. A long-running or human-in-the-loop process is started with
workflow_start instead.
Read-only and write Tools
Workflows that update data or drive a business process can be Tools. Today
the normalizer and Workflow admission (rule_api.rs, the sync nested-target
check) reject a sync binding unless every reachable Tool is readOnly, not
destructive and not humanApprovalRequired, so write workflows could only
be async Tools.
The read-only rule exists because a sync caller can give up before the run finishes. An MCP client or Agent that times out usually retries. With a read, a duplicate run costs capacity. With a write, the client cannot tell whether the first attempt changed anything, and the retry may do it twice.
With workflow_invoke the retry re-attaches to the in-flight run or replays
its result, which removes that risk as long as the replay window covers the
client’s retry horizon. The Tool’s own annotations decide what the workflow
may do:
| Tool | Required |
|---|---|
Read-only (readOnly: true) | the workflow reaches only reads (matrix below) |
Writes (readOnly: false) | idempotencyPolicy.resultReplayMs ≥ 600000 for every key kind |
Destructive (destructive: true) | as for writes; clients see the flag and ask for confirmation |
Contains a human task or reaches a humanApprovalRequired Tool | rejected; use workflow_start |
Sync effect matrix. Workflow checks each task of the pinned definition against the outer Tool’s annotations at binding publish and again at admission:
| Task | Read-only Tool | Write Tool | Destructive Tool |
|---|---|---|---|
HTTP GET/HEAD | allowed | allowed | allowed |
| HTTP write method | rejected | allowed, with task evidence | allowed, with task evidence |
Nested MCP Tool, readOnly: true | allowed | allowed | allowed |
Nested MCP Tool, readOnly: false | rejected | allowed, with task evidence | allowed, with task evidence |
Nested MCP Tool, destructive: true | rejected | rejected | allowed, with task evidence |
Nested Tool with humanApprovalRequired, or a human task | rejected | rejected | rejected |
The static fit check at publish reuses the existing retry and budget envelope calculation, so a write Tool’s worst-case attempts still fit the deadline.
Concurrency and rate limits
runtimeBounds limits one run. admissionLimits limits how many runs the
Tool can start, so the workflow owner can approve a known load:
{
"maximumConcurrentRuns": 20,
"maximumConcurrentRunsPerUser": 2,
"startsPerMinute": 120,
"startsPerMinutePerUser": 10
}
The normalizer fills these defaults when the Tool author leaves them out,
the same way it fills runtimeBounds. The author can change them, and the
workflow owner approves the final values.
The four numbers must agree with each other. For a sync Tool, the number of runs in flight is roughly the start rate times the average run time. At 120 starts per minute (2 per second) and a typical 5 second run, about 10 runs are in flight, well under 20. The concurrent cap only bites when runs slow down toward the 30 second deadline: 2 per second × 30 seconds would be 60, and the cap holds it at 20. That is the case it is for, since a slow downstream API should not have more and more runs piling onto it. The per-user values (2 concurrent, 10 per minute) stop one user or a looping Agent from taking the whole allowance.
Workflow enforces the limits inside the admission transaction, under a
per-Tool advisory lock so concurrent admissions cannot both pass. It counts
non-terminal runs and starts in a sliding one-minute window for the Tool
across all of its revisions, so a republish does not reset the counters. A
re-attach or replay does not count as a new start. When a limit is hit,
Workflow returns WORKFLOW_CAPACITY_EXHAUSTED with retryAfterMs. The
Gateway permit pool still protects the Gateway itself, but Workflow’s limits
are the ones the owner approves.
Publication Through Workflow MCP
Operations
Workflow exposes these operations on its native MCP endpoint, next to
workflow_start. They are declared in
contracts/workflow-admin/workflow-tools-list-full.json with schemas and examples like
the other workflow-admin tools.
| Tool | Purpose |
|---|---|
workflow_definition_save | Upsert the saved (editable) head of a definition |
workflow_definition_publish | Publish one immutable definition version |
workflow_definition_retire | Stop new admissions to a definition version |
workflow_definition_grants_sync | Replace the Tool grants of one definition |
workflow_binding_publish | Publish a binding revision with its dependencies and endpoint targets |
workflow_binding_retire | Retire a Tool’s binding, called by the Tool owner |
workflow_binding_get | Read one revision with its status and decision history; headOnly:true with toolId and wfDefId reads its publication head |
workflow_binding_list | List revisions for definitions the caller owns, filterable by status |
workflow_binding_decide | Approve or reject a pending revision, called by the definition owner |
workflow_binding_revoke | Withdraw approval of an active revision, called by the definition owner |
Every call carries the acting user’s token in authorization. The Gateway adds its own token in
x-scope-token, and Workflow checks it against
workflow.invocation.allowedCallerServiceIds, as for workflow_start.
The Portal reaches these tools through the same Gateway /mcp route that
StartWorkflow uses, and the Gateway connects to Workflow. The Portal never
connects to Workflow or to the Workflow database directly. Today the Gateway
uses one configured Workflow endpoint. When each host runs its own
light-workflow instance, the Gateway will locate it through controller-rs
service discovery; the publication contract does not change.
Request host and identity
The ten definition-publication and binding-management tools carry a required
hostId as a target-host assertion. It does not supply identity or grant
cross-host access. Workflow authenticates first and rejects a mismatch with
WORKFLOW_POLICY_DENIED before any scoped read, mutation or operation-receipt
lookup/replay. Database scope comes from the verified user host. The request
body, user bearer, and Gateway scope bearer must identify the same host.
Gateway caller-service validation remains required for every call.
workflow_invoke has no hostId argument; its host is derived from trusted
invocation context. The manifest retains identitySource: trustedInvocationContext.
The input-schema validator permits only the exact root hostId paths on the
ten tools, without relaxing its identity/fencing guard for existing tools,
alternate spellings or nested fields.
The owner objects on definition save/publish describe user-authenticated
resource metadata, not caller identity. Only save changes current ownership;
authorization uses the verified caller and stored owner. The role field on
binding list is only the owner/requester relationship filter relative to
that caller. These are exact tool/path validator exceptions with constrained
schemas. A sync payload’s actor is audit metadata and grants no authority.
Gateway assertion
The user token identifies who clicked. Gateway checks that user’s endpoint
permissions before forwarding, and Workflow independently checks the user host
and Gateway service identity. A payload actor cannot replace either token.
All definition, grant, and binding operations carry the acting user’s bearer
in Authorization. Gateway verifies that user and the endpoint ACL. On the
Gateway-to-Workflow request, Gateway supplies its own app bearer in
X-Scope-Token; Workflow verifies its service identity, host, and environment
independently from the user bearer and checks that the user host matches the
request host. hybrid-command does not mint an application token for Workflow.
The payload actor remains audit metadata, never a credential.
- A missing or invalid user or Gateway token returns
isErrorwithWORKFLOW_POLICY_DENIED.
workflow_binding_get, workflow_binding_list, workflow_binding_decide
and workflow_binding_revoke use the same user and Gateway tokens. Workflow checks the
caller against the definition owner it stored, so they do not depend on the
Portal’s word.
Ownership. The definition owner (owner_user_id, owner_position_id) is
set by workflow_definition_save, so an ownership change in the Portal
reaches Workflow through the publisher-authenticated save. Active approvals
stay active after a transfer. Pending revisions move to the new owner.
Carry-over requires that the owner recorded on the basis revision equals the
current owner (see Carry-over). A position holder is checked from the
position claims of the user token, so position changes take effect when the
user’s token is renewed.
Saved definitions
workflow_start runs the saved, editable head of a definition
(wf_definition_t), and the FDW sync kept that head in step with the Portal.
Without the sync, Workflow needs its own write path for it.
workflow_definition_save input:
{
"hostId": "01964b05-552a-7c4b-9184-6857e7f3dc5f",
"wfDefId": "0198a3c2-...",
"sourceRevision": 7,
"actor": "steve",
"namespace": "claims",
"name": "summarize-claim",
"version": "1.2.0",
"definition": "document:\n dsl: 1.0.0\n ...",
"lifecycleStatus": "PUBLISHED",
"catalogVisible": false,
"owner": { "userId": "...", "positionId": "..." },
"active": true
}
definition is the YAML text as stored in the Portal. catalogVisible is a
required boolean independent of lifecycle status; a published definition may
remain private. Portal sends its stored value from the same consistent snapshot
as the other save fields and normalizes legacy database NULL to false. Visibility
is outside the definition-text digest. The saved head accepts at most 126
characters for actor, namespace and name, and 20 for version, matching
the existing table. Published immutable versions retain a 64 character version
limit. sourceRevision is
the Portal definition’s aggregate version. Workflow keeps the last applied
revision on the head row:
- a lower revision returns
result: "stale"and changes nothing, so a delayed older save can never restore older text, owner or deletion state; - the same revision with the same content and visibility returns
unchanged; with different content or visibility,WORKFLOW_IDEMPOTENCY_CONFLICT; - a higher revision upserts the head and returns
saved.
Every receipt, stale included, returns appliedRevision (the revision
Workflow now holds) and the definitionDigest of the stored head at that
revision, so the pair always describes Workflow’s state. active: false
records a delete. The Portal delivers saves through the durable
synchronization described below.
Save-then-start. A start must run exactly what the user saved:
- Portal
StartWorkflowchecks that Workflow has acknowledged the current Portal revision of the definition. If not, it delivers the save at once and fails the start if that delivery fails. - It also initializes and acknowledges the current full grant set, including legacy grants with no sync row. If grant delivery fails, it does not start. After grant delivery it rechecks the acknowledged definition revision/digest pair. This is a readiness check; grants remain live at runtime.
- It then calls
workflow_startwith the new optionalexpectedDefinitionDigest, set to the digest acknowledged together with that revision. Workflow compares it to the head it loads and returnsWORKFLOW_DEFINITION_MISMATCHon a difference.
The workflow_start MCP catalog input schema must declare this optional
digest with ^sha256:[0-9a-f]{64}$. Gateway validates arguments against its
published tool schema before calling Workflow. The catalog must first be
reimported into Portal’s MCP API endpoint (api_endpoint_t.tool_schema), then
the Gateway Tool publication must be previewed and published and its config
snapshot activated. Rebuilding and restarting Workflow or republishing Gateway
Tools alone leaves an older Portal endpoint schema in place.
Every Portal start path goes through StartWorkflow, including the workflow
editor’s Start button. A body idempotency key takes precedence over the header.
With neither, Start means new work and allocates a fresh key; callers retrying
one intended start must reuse an explicit key.
workflow_definition_publish never changes the head.
Digest parity. The Rust digest is sha256: plus
execution_runner_protocol::canonical_sha256 of the YAML parsed to JSON. The
Java WorkflowDefinitionDigest uses SnakeYAML with YAML 1.1 rules. The two
disagree on some inputs: yes/no/on/off, unquoted timestamps, 1 vs
1.0, merge keys and explicit nulls. Workflow’s digest is authoritative, and
Portal compares its own only as a precheck. A shared fixture set covers these
cases, and both sides must agree on it before the Portal precheck is trusted.
Definition publish
Input:
{
"hostId": "01964b05-552a-7c4b-9184-6857e7f3dc5f",
"wfDefId": "0198a3c2-...",
"namespace": "claims",
"name": "summarize-claim",
"version": "1.2.0",
"definition": "document:\n dsl: 1.0.0\n ...",
"expectedDefinitionDigest": "sha256:...",
"bindingApproval": "carryOver",
"owner": { "userId": "...", "positionId": "..." },
"operationId": "0198a3c3-..."
}
definition is YAML text, as in save.
Workflow must:
- Parse the definition and recompute its definition and input schema
digests. It rejects the call if the result differs from
expectedDefinitionDigest. - Run the generic definition validation (runtime definition and CEL checks). Binding-specific limits such as effect mode and budget are checked at binding publish, not here.
- Store the version in a version table keyed by
(hostId, wfDefId, version). Versions are immutable.
Receipt:
{
"result": "published",
"status": "active",
"wfDefId": "0198a3c2-...",
"version": "1.2.0",
"definitionDigest": "sha256:...",
"schemaDigest": "sha256:...",
"bindingApproval": "carryOver"
}
Version states. A version is active or retired; retired is
terminal.
- Republishing the same digest returns
result: "unchanged"with the current status, includingretired. - A different digest for an existing version is rejected with
WORKFLOW_DEFINITION_MISMATCH. workflow_definition_retireis refused while an active binding revision pins the version. Pending revisions on it becomewithdrawn, each with a decision record.- Admission against a retired version returns
WORKFLOW_DEFINITION_RETIRED. Running runs finish.
Existing binding rows are not backfilled with version rows. They are admissible again only after the Portal republishes them.
Grants
A grant lets a workflow call a Tool through LightAPI. The Tool owner approves
it through RequestWorkflowToolAccess and DecideWorkflowToolAccess, so it
belongs to the definition, not to a binding.
workflow_definition_grants_sync input:
{
"hostId": "...",
"wfDefId": "...",
"sourceRevision": 12,
"actor": "steve",
"grants": [
{
"grantId": "...",
"toolId": "...",
"toolVersion": "1.0.0",
"lightapiDigest": "sha256:...",
"allowedEnvironments": ["dev"]
}
],
}
It replaces the full grant set of that definition in one transaction.
sourceRevision works as in save: a lower revision is stale and changes
nothing, so a delayed older set can never restore a removed grant. The
Portal delivers it after DecideWorkflowToolAccess, after an access
revocation, and in Publish Selected before binding publish.
Grants are read live, not pinned to a run. A revocation is enforced once Workflow acknowledges the grant set that removes it; until then the Portal shows “Sync pending” on the access list. After that, every later dispatch that needs the grant is refused, including the next call of a run already in flight. An operation already dispatched is not undone.
Synchronization
Definition saves and grant sets are delivered by authenticated user actions:
- The Portal keeps a sync row per definition and kind (
definition,grants) with a desired revision and the user of the last change, written in the same transaction as the event that changed the data, and the revision and digest Workflow last acknowledged. - Delivery reads the revision, the user and the data in one consistent
database snapshot, so content is never paired with another revision’s
number. When several grant changes are delivered as one set, the user of
the last change is sent as
actor. - Save and Publish in the Workflow Editor attempt delivery with the current
user bearer. A committed local revision remains pending if delivery fails.
getWfDefinitionByIdreturns desired and acknowledged definition and grant revisions, and the editor shows the difference with a Sync now action.syncWfDefinitionretries both kinds under the current user bearer. There is no scheduled synchronization worker or stored user credential. - The Portal stores the acknowledged revision and digest together and only moves them forward, so a late receipt for an older revision is ignored. Errors are recorded and shown in the UI.
- Definitions and grants that existed before this change get their sync row on first use: the definition is delivered before its first start, and the full grant set before its first binding publication.
- A fresh or rebuilt sync row may start at or below the revision Workflow
already holds. Workflow reports its stored revision in a
stalereceipt or in the equal-revision conflict. For grants, which the Portal owns, the Portal moves its revision above Workflow’s and sends its current set again. For a definition, whose revision is the Portal aggregate version, the sync stops with an error for investigation until a later edit moves the revision past Workflow’s.
Lost receipts
Final review amendment (2026-09-27): remediation is specified in
implementation/light-workflow/workflow-invoke-execution/final-remediation-handoff.md.
The ledger classifies operation outcome separately from retryability. Proven
rejection of the exact stored operation may close it as failed. Authentication,
transport, Gateway, unreadable-response and unproven JSON-RPC errors preserve
pending state. Initially, definitive rejection is limited to validated binding
publish/retire VERSION_CONFLICT responses from paths that check the original
receipt before rejecting its stale expected version. Retry never edits the
stored request. Other errors do not prove the original outcome.
F2 request-validation evidence (2026-09-27): For binding publish and
retire, a negative expectedAggregateVersion is a request-only validation
failure. The authenticated publisher and host checks precede parsing. A
parseable request with hostId and operationId reaches the serialized
operation_begin insert/row lock and exact tool/request-digest comparison
before this check. An existing success receipt wins. The rejection is stored
under that same operation ID while holding the operation row lock, so an
overlapping attempt and later replay return the same rejection even if
validation rules change. The error remains WORKFLOW_INPUT_INVALID,
afterEffect=false, and includes additive details.requestValidation:
{version:1, discriminator:"negativeExpectedAggregateVersion", operationId:"<UUID>", toolName:"workflow_binding_publish|workflow_binding_retire"}.
Portal accepts this evidence only from a Workflow Tool error for the matching
stored operation/tool and a numeric negative expected version in the stored
request. Missing, malformed, mismatched, or unmarked errors remain unknown.
Additive binding request-validation evidence (version 2)
After authenticated host verification and serialized operation lookup, binding
publish may store a request-only rejection for either bindingFields or
bindingReach. Evidence is
{version:2, discriminator, operationId, toolName:"workflow_binding_publish", requestSection}.
bindingFields requires requestSection:"binding"; bindingReach requires
requestSection:"reach" and covers the typed dependencies and
endpointTargets arrays together. The section identifies the immutable request
member inspected by Workflow. The rejection stores the
original WORKFLOW_INPUT_INVALID code, exact message, and complete evidence;
replay returns those stored values. Portal accepts it only for a Workflow Tool
error with afterEffect=false, a matching operation and tool, a nonnegative
expected version, and the named section present with its contract type in its
stored exact request. Portal does not reinterpret a bare business error code as
proof. Both versions require exactly their documented evidence members; extra
members are malformed and remain unknown. The earlier version 1 negative-version
marker remains supported.
bindingFields covers every explicit field, bound, policy, digest-format,
role, annotation, and schema-type rejection in validate_binding.
bindingReach covers reach size, dependency field/digest/policy/uniqueness,
and endpoint field/digest/uniqueness/URI/method/document rejections in
normalize_payload. Digest format checks use immutable supplied strings.
digests serialization and canonical-hash computation errors do not establish
invalid input and carry no evidence. Definition, grants, owner, current head,
static fit, runtime configuration, database, authentication, transport, and
other state-dependent failures remain unknown unless separately proven.
Deserialization errors cannot safely identify and fence an operation, so they
remain unknown. Other WORKFLOW_INPUT_INVALID paths remain unknown, including
state-dependent missing binding and limits. The existing validated
VERSION_CONFLICT rule remains independent.
A validated definition publication receipt with result=unchanged and
status=retired is a completed D21 outcome. The definition prerequisite for
binding publication additionally requires status=active; a retired receipt
is neither cached as active nor followed by a binding send. Original-user
Retry replays the retired receipt, and publishing another version is deliberate
new work. Gateway-removal retirement results retain Portal-owned identity,
status, code and message; unconfirmed results expose only validated recovery
metadata. A forbidden Retry means only that the caller is not the original
requester. It does not establish completion, failure or expiry.
UI recovery follow-up: A recovered definition receipt with either
result=published or result=unchanged and status=active enables an explicit
binding-publication continuation; it never sends the binding automatically.
When the authenticated Portal ledger returns a stored failed operation, Portal
adds metadata.operationState="failed" to the command error while preserving
the stored business error code. This Portal-owned marker, rather than the code
or retryable flag, lets the UI stop Retry and offer deliberate new work.
Unproven remote errors do not receive the marker.
Ledger database errors and unknown operation-prefixed Retry errors remain
recoverable with the original operation ID; neither establishes failure or
expiry. The UI handles explicit expiry, forbidden Retry, and not-found
responses separately.
Concurrent finalization rereads the committed outcome when its guarded update loses. Remote reconciliation happens outside the short receipt/event/projection transaction. Publication event-version allocation is serialized using the existing global aggregate identity. Wrong-user Retry is forbidden regardless of state and must not expose receipts or masquerade as pending.
Every other mutating call the Portal makes to Workflow (definition publish
and retire, binding publish, retire, decide and revoke) goes through a
Portal operation ledger. The Portal stores the operationId, the exact
request and the requesting user before calling. After a timeout the
operation shows “Unconfirmed — Retry”. Only the same user can retry. Retry
is its own Portal command that names the original operationId; it carries
that user’s current token and resends the same request with the same
operationId, and it never creates a new operation. Workflow returns the
stored receipt, so the retry resolves the original operation instead of
making a second mutation. Running the original command again while the
operation is unconfirmed sends nothing and points to Retry. Another user cannot take over that
operation; they see who has it unconfirmed. This covers one operation
stream (one operation on one subject). Other authorized operations on the
same Tool or definition are not blocked; Workflow’s own concurrency checks
govern them.
An unconfirmed operation expires 29 days after it was created. Workflow drops its record 30 days after its own creation. The deadline never moves. Every retry checks it first, and an expired operation is never resent; it shows “Expired; remote outcome unconfirmed”, because Workflow may or may not have applied it. The UI offers Refresh status and then a separately labelled action to publish (or decide, revoke, retire) again. That action is new work and creates a new operation.
The one-day margin is an operational condition, not a guarantee: a resend is safe when the clock skew between Portal and Workflow plus the time from the Portal’s expiry check to Workflow processing the request is under 24 hours. Workflow does not enforce the Portal deadline itself.
These operations need a user token, and the ledger does not store tokens, so nothing retries them in the background. A receipt updates the Portal projection only when its aggregate version is newer than what the Portal already holds.
Binding revisions
Every Workflow workflow_tool_binding_t row is an immutable revision.
binding_id is the revision id, generated by Workflow. The Portal’s binding
id is stored as source_binding_id.
A revision has a revision_status:
| Status | Meaning |
|---|---|
pendingApproval | Waiting for the definition owner |
approved | The one active revision for the Tool (active = true) |
rejected | Refused by the owner, with a comment |
superseded | Replaced by a newer approved or pending revision |
withdrawn | Pending on a definition version that was retired |
retired | Retired by the Tool owner |
revoked | Approval withdrawn by the owner |
legacy | Written by the projection sync before this change; never admitted |
active is true only for approved, and at most one revision per Tool is
active and at most one is pending. Dependencies, endpoint targets and task
evidence are written once per revision and never updated.
An admitted run records its binding_id. Every later read in the run uses
that revision without checking active: the grant join and endpoint lookup
in executor.rs, the dependency join in bound_mcp.rs, and
validate_pinned_dependencies. Revoking or superseding a revision therefore
stops new admissions but not runs in flight; they finish. The idempotency
replay trigger also uses the pinned revision’s policy.
Each Tool has a publication head row that holds its aggregate version, the active revision id and the pending revision id. Every publication operation locks it (see Concurrency).
Binding publish
The input carries the normalized binding and the records Workflow needs. The
Portal resolves endpoint targets from tool_t, api_* and the permission
tables before the call, which the projection script used to do over the FDW.
Workflow receives only resolved values.
{
"hostId": "01964b05-552a-7c4b-9184-6857e7f3dc5f",
"binding": {
"sourceBindingId": "...",
"toolId": "...",
"toolName": "claims.summarize",
"wfDefId": "...",
"workflowVersion": "1.2.0",
"definitionDigest": "sha256:...",
"schemaDigest": "sha256:...",
"invocationMode": "sync",
"syncWaitMs": 20000,
"totalDeadlineMs": 30000,
"executionClass": "interactive",
"resultTextMode": "compact-json",
"cancellationPolicy": "before-effects-only",
"idempotencyPolicy": { "kind": "derived", "resultReplayMs": 0 },
"delegationPolicy": { "maximumDelegationDepth": 1 },
"runtimeBounds": { "maximumTaskAttempts": 8, "...": "..." },
"admissionLimits": { "maximumConcurrentRuns": 20, "...": "..." },
"callerPolicy": {},
"toolAnnotations": { "readOnly": true, "destructive": false },
"policyDigest": "sha256:...",
"responsePolicyDigest": "sha256:..."
},
"dependencies": [],
"endpointTargets": [],
"expectedAggregateVersion": 3,
"operationId": "0198a3c4-..."
}
There is no approvalEvidence in the input. Workflow generates task evidence
itself when the revision is approved.
Workflow must:
- Require that the pinned definition version is published, active, and that its digests match.
- Validate the binding fields and constraints, including the sync effect matrix for the Tool’s annotations.
- Run
validate_orchestration_definitionandvalidate_pinned_dependenciesagainst the payload, and the static fit check (see Carry-over). - Check each endpoint target against Workflow’s own destination policy. The Workflow owner can refuse a target that the Portal has resolved.
- Compute
bindingDigestandapprovalDigest(see Digests). - Decide the status (see Status rules), insert the revision and its related rows, update the Tool head, and store the receipt, all in one transaction.
Receipt:
{
"result": "published",
"status": "active",
"toolId": "...",
"bindingId": "<revision id>",
"sourceBindingId": "...",
"workflowVersion": "1.2.0",
"definitionDigest": "sha256:...",
"bindingDigest": "sha256:...",
"approvalDigest": "sha256:...",
"aggregateVersion": 4,
"decisionId": "..."
}
status is active or pendingApproval. decisionId is present when the
publish itself recorded a decision (self-approval or carry-over).
Digests
bindingDigest identifies a revision. It is the canonical SHA-256 of:
- every binding field except
sourceBindingId,operationIdandexpectedAggregateVersion; - the dependencies, sorted by
(authorizationToolName, nestedToolId, nestedToolVersion); - the endpoint targets, sorted by
endpointRef, with methods uppercased, deduplicated and sorted, and environments sorted.
Null or absent fields are omitted. A duplicate set key is rejected, not merged.
approvalDigest covers the fields that change load or behaviour: mode,
syncWaitMs, totalDeadlineMs, execution class, cancellation, idempotency
and replay, delegation, runtime bounds, admission limits, caller policy and
toolAnnotations. It excludes reach (dependencies and endpoint targets),
which carry-over compares as sets, and the definition version. Changing the
Tool name, description or input examples changes neither what the owner
approved nor the approval digest.
The previous payloadDigest is removed.
Publish ordering
The Portal Publish Selected action works per Tool:
workflow_definition_saveandworkflow_definition_publishfor each pinned definition version, andworkflow_definition_grants_syncfor each definition involved.workflow_binding_publishfor each workflow-backed Tool in the selection.- Build
mcp-router.yml:- a Tool whose receipt is
activeis included with the receipt’sbindingDigestanddefinitionDigest; - a Tool whose publish is pending or failed keeps its previous Gateway entry, if it has one, or stays out;
- the rest of the selection is promoted normally.
- a Tool whose receipt is
The result dialog lists every Tool with its outcome, so one pending or failed Tool does not block the others, and the Gateway never routes to a revision that Workflow has not made active.
Retirement runs in the reverse order. The Portal removes the Tool from the
Gateway snapshot first, then calls workflow_binding_retire. Runs in flight
keep their pinned revision and finish. New workflow_invoke calls for the
retired Tool are rejected.
If Gateway removal was staged but retirement remains unconfirmed until its
Portal operation expires, the requester may start explicit new work with the
Portal RetireWorkflowToolBinding {hostId, instanceId, toolId, expectedAggregateVersion} command. Portal checks that the Tool is absent
from that instance’s current staged Gateway publication and refuses retirement
if it was re-added. The caller supplies the Workflow Tool-head version read
after Refresh status; Portal rejects a changed head and never substitutes a
newer version. A live pending retirement returns
WORKFLOW_OPERATION_PENDING. Once the old operation expires, this command
allocates a new operation ID and uses the existing D21 ledger, publisher-token
path, and retirement completion event. RetryWorkflowOperation still targets
only the old operation and sends nothing after expiry. Gateway removal is not
staged again by this command. Snapshot activation and live routing remain
separate operator steps.
Preview. The Portal stores published_request_digest on its binding row:
the canonical digest of the binding publish input it sent, without
operationId and expectedAggregateVersion. Preview and the server-side
PublishGatewayTools include a Tool only when the payload rebuilt from the
current Portal rows has that digest and the stored status is active. An edit
since the last publish therefore shows as “needs publish”, not as served.
Concurrency
Each Tool has a head row in workflow_tool_publication_t with
aggregate_version, active_binding_id and pending_binding_id. Every
publication operation:
- inserts the head with
ON CONFLICT DO NOTHING, then selects itFOR UPDATE, which also serializes the first publication of a Tool; - locks in a fixed order: the definition version row first (
FOR SHAREfor publish and decide,FOR UPDATEfor retire), then the Tool head.
expectedAggregateVersion applies to binding publish and binding retire. A
mismatch is a conflict, with the same semantics as other aggregate writes.
Decide and revoke pin expectedBindingDigest instead, so the owner acts on
exactly the revision they reviewed.
Operation idempotency. Every write carries an operationId. Workflow
stores (hostId, operationId) with the tool name, a request digest and the
receipt. A repeat with the same request returns the stored receipt. A repeat
with a different request returns WORKFLOW_IDEMPOTENCY_CONFLICT. Records are
swept after 30 days.
The binding publish receipt includes optional carryOverDeniedReason when
carry-over is denied and the new revision is pendingApproval. Active publish
receipts, including self-approval and successful carry-over, omit the field.
The operation record stores the complete response, so replay returns the
original reason unchanged.
Migration 0019 adds immutable input_schema and output_schema JSON columns
to each binding revision, preserving the optional Tool schemas covered by
bindingDigest. It also stores carry_over_denied_reason on the revision.
A new operation returning an unchanged pending revision includes that stored
reason, and workflow_binding_get reads it after operation receipts expire.
Binding get and list use read-only REPEATABLE READ transactions so their
head pointers, revision statuses, decisions, items and counts come from one
committed snapshot.
After a timeout or a lost receipt, the Portal refreshes from
workflow_binding_get and reconciles status, digests and aggregate version.
Workflow owner approval
Binding a Tool to a workflow lets a new audience start that workflow, with a deadline and budget the Tool author chooses. The workflow owner is accountable for that load and for what the workflow does, so the owner must approve a binding created by someone else. This is the reverse of the existing Tool access approval:
| Direction | Requester | Approver | Mechanism |
|---|---|---|---|
| A workflow calls a Tool | Workflow author | Tool owner | RequestWorkflowToolAccess → approval workflow → DecideWorkflowToolAccess → workflow_definition_grants_sync |
| A Tool invokes a workflow | Tool author | Workflow owner | workflow_binding_publish → pending revision in Workflow → workflow_binding_decide |
The pending revision and the decision are held by Workflow, not the Portal. Workflow belongs to the owner’s team, and the owner’s decision reaches it with the owner’s own user token, so Workflow does not have to trust the Portal’s word that approval happened.
Status rules
workflow_binding_publish:
- unchanged when the new
bindingDigestequals the active or pending revision’s digest; the existing receipt is returned. - active (selfApprove) when the publishing user is the definition owner.
The previous active and pending revisions become
superseded. - active (carryOver) when the carry-over rules below hold.
- pendingApproval otherwise. A previous pending revision becomes
superseded. The active revision, if any, keeps serving.
workflow_binding_decide with approve revalidates the revision inside the
transaction (definition version active, static fit, effect matrix) and then
activates it; the previous active revision becomes superseded. reject
requires a comment.
workflow_binding_revoke sets the active revision to revoked. New
workflow_invoke calls for the Tool return WORKFLOW_POLICY_DENIED with
“binding revoked by workflow owner” until the Portal removes the Tool from
the Gateway. Runs in flight finish.
workflow_binding_retire sets the active revision to retired and a pending
one to withdrawn.
Carry-over
The definition version is not part of the approval digest. When the Tool author re-pins to a newer version of the same definition, the approval of the current active revision carries over if all of these hold:
- same host, Tool and
wfDefId; - the owner recorded on the active revision equals the current definition owner;
- equal
approvalDigest; - the new reach is a subset of the active revision’s reach, where a
dependency is keyed by
(nestedToolId, nestedToolVersion, contractDigest, authorizationToolName)and a target by(endpointRef, endpointUri, sorted allowedMethods); - the new version’s write tasks are a subset of the active revision’s, keyed by qualified task name, kind, and method or Tool;
- the new version passes the static fit check; and
- the new version was published with
bindingApproval: carryOver.
Otherwise the revision is pendingApproval.
A heavier new version cannot overload the workflow under a carried-over
approval, because the approved limits still apply to it. To stop that from
surfacing as runtime failures, workflow_binding_publish checks the
version’s static requirements against the binding: fork width against
maximumParallelism, declared nested calls against maximumNestedCalls, and
the retry and budget envelope against maximumTaskAttempts and the deadline.
A version that cannot fit is rejected with a message that names the limit.
The author raises that limit, which changes the approval digest and sends the
revision to the owner.
The owner knows best when a version is heavier, so
workflow_definition_publish takes bindingApproval of carryOver
(default) or reapprove. With reapprove, every revision that pins that
version goes to the owner.
workflow_binding_decide input:
{
"hostId": "...",
"bindingId": "<revision id>",
"expectedBindingDigest": "sha256:...",
"decision": "approve",
"comment": "Approved for claims read-only lookups",
"operationId": "..."
}
Decision history
Every status change appends a row to workflow_tool_binding_decision_t:
decision id, Tool, revision, action (approve, reject, revoke, retire,
supersede, withdraw, carryOver, selfApprove), actor, comment,
approval digest, time and operation id. The table is append-only: the runtime
role has INSERT and SELECT only. workflow_binding_get returns it as the
revision history.
Task evidence
validate_approval_evidence requires, for each write task (a non-GET/HEAD
HTTP call, or an MCP call to a Tool that is not readOnly), the task’s
approvalEvidenceDigest metadata and an active evidence row for
(host, binding_id, task_name, digest). The projection sync never copied
these rows, so write tasks could not be admitted.
When a revision becomes active by self-approval, approval or carry-over,
Workflow writes one evidence row per write task of the pinned version, with
evidence_digest set to the task’s approvalEvidenceDigest and
approved_by set to the definition owner. A write task without
approvalEvidenceDigest metadata fails binding publish. The evidence table
keeps its current shape.
UI flow
All screens are in portal-view. Portal commands call the Workflow tools
through the Gateway with the acting user’s token, and store the returned
status on the Portal workflow_tool_binding_t row so lists do not need a
Workflow call. The review form reads the live revision from Workflow.
Tool author, GenAPI Admin → Tool
- The author creates or edits a Tool with
executionPlacement=workflowand picks a published workflow version. The form shows the definition owner. When the author is not the owner it shows “This binding needs approval from <owner> before it can be served.” - The author runs Publish Selected. The result dialog lists each Tool: “Published”, “Waiting for approval from <owner>”, or the Workflow error. The rest of the batch is published normally.
- The Tool list shows a binding status column: Active, Pending approval, Rejected, Revoked, Needs publish. Rejected and Revoked show the owner’s comment. The author can edit and publish again, which creates a new pending revision.
- After approval the status becomes “Approved, not yet on Gateway” until the next Publish Selected includes it. Gateway publication stays an explicit action, as it is today.
Workflow owner, Workflow → Definitions
- The definition list shows a “Tool bindings” column with a badge for the pending count, for example “2 pending”. The owner can filter the list to definitions with pending revisions.
- The row action “Tool Bindings” opens a list of revisions for all versions of that definition: Tool name, Tool owner, definition version, status, requested or decided time and who decided.
- Selecting a pending revision opens the review form, loaded with
workflow_binding_get:- Requester: Tool name, description, Tool owner, host.
- Target: definition name, version and digest.
- Execution: mode, sync wait, total deadline, execution class, cancellation policy, read-only and destructive flags.
- Limits: idempotency and replay window, delegation depth, task attempts, nested calls, parallelism, request, intermediate and result bytes, cost units, admission limits, caller policy.
- Reach: dependencies, endpoint targets and write tasks.
- Change: when an active revision exists for the Tool, the changed fields are highlighted against it.
- A comment box and Approve / Reject buttons. Reject requires a comment.
- An active revision has a Revoke action with a required reason.
Worklist. A pending revision also appears in the owner’s Worklist, like Tool access requests do today. Opening it goes to the same review form.
What is removed
- The
workflow-projection-syncservice in the Compose files. crates/workflow-store/deployment/publish-workflow-projections.sh, and thepostgres_fdwserver and user mapping it needs.- Any documentation step that tells operators to run the sync.
Workflow schema changes
The writer changes from the FDW script to the publication handlers, and the
tables change with it (migration 0018):
workflow_tool_binding_tgainssource_binding_id,revision_status,binding_digest,approval_digest, the owner snapshot (owner_user_id,owner_position_id),requested_by,requested_ts,approved_by,approved_ts,approval_basis_id,cancellation_policy,admission_limits,caller_policyandtool_annotations. A check enforcesNOT active OR revision_status = 'approved'. Partial unique indexes allow one active and one pending revision per Tool.- The
UNIQUE (host_id, tool_id, workflow_version)constraint onworkflow_tool_binding_tis dropped, so a Tool can have several revisions on the same definition version (a legacy row plus its republish, or a changed binding). Foreign keys use(host_id, binding_id). workflow_endpoint_target_tis already keyed by(host_id, binding_id, endpoint_ref)since migration 0005, so each revision has its own targets. Migration 0018 preserves this key; fresh and upgrade tests verify it.wf_definition_tgainssource_revision, and the newworkflow_definition_grant_sync_tstores the last applied grant revision and digest per definition (see Saved definitions and Grants).workflow_claim_idempotency_v2is added (see Idempotency).- New tables: the definition version table,
workflow_tool_publication_t,workflow_publication_operation_tandworkflow_tool_binding_decision_t. - Every existing binding row becomes
legacywithactive = false. A Tool serves again only after it is republished. workflow_action_authority_tgainscredential_kindandworkflow_run_credential_tis added (see Run credential selection).- The idempotency replay trigger reads the pinned revision’s policy.
Workflow Invoke
Callers and routing
| Caller | Operation | Admission profile |
|---|---|---|
| Workflow Editor, scheduler, API client | workflow_start | PortalExecution |
| Gateway, executing a workflow-backed Tool for any MCP client, including Agents | workflow_invoke | WorkflowBacked |
Identity and user token
workflow_invoke receives the same two credentials as workflow_start: the
original user token in authorization, and the Gateway’s token in
x-scope-token. Workflow authenticates them with the existing authenticate
path in rule_api.rs: both JWTs must verify, and the scope token’s service
id must be in workflow.invocation.allowedCallerServiceIds.
Unlike workflow_start, there is no light-oauth registration and no token
exchange. The run uses the original user token for its outbound calls.
Token lifetime is a configuration matter. The Gateway renews user tokens
before they expire (renew_before_seconds in spa_auth), and JWT
verification allows clock_skew_in_seconds after exp (security.yml).
Operators set these so that a token accepted at admission lasts a sync run.
Workflow does not add a check of its own. If a token stops verifying during
a run, the next outbound call fails with an authentication error.
Run credential selection
This rule applies to every run, including async runs started with
workflow_start. For each outbound HTTP or MCP call, the task executor picks
the user token like this:
- Use the original user token while it verifies and is more than
runCredential.originalTokenMarginSecondsbefore itsexp. - Otherwise, if the run has a LONG registration, use the token exchanged
through light-oauth (
LongAuthority::token_for). - Otherwise use the original token while it still verifies, and fail the call with an authentication error once it does not.
The margin only decides when a LONG run switches to the exchanged token. Only
workflow_start runs register with LONG, because only they can outlive the
original token. A workflow_invoke run never does.
The same selector serves bound MCP and protected executor HTTP calls. An expired Invoke credential fails the outbound call. Retired broker rows have no credential source and fail closed.
For registered protected executor HTTP calls, the target rule follows the selected token. A LONG run using its still-valid original user token may call the approved endpoint directly. Once the selector exchanges that token, the HTTP target must use the configured Gateway origin. A definition whose approved endpoint remains direct needs a Gateway route and revised target to keep working after the exchange margin is reached.
Where the original token is kept. A run can move to another Workflow replica, so the token must be stored, not only held in memory.
workflow_startruns already store it inworkflow_long_credential_tas part of the LONG registration.workflow_invokeruns store it sealed inworkflow_run_credential_t, expiring atdeadline_ts. It is deleted when the run reaches a terminal state, and a sweeper deletes any that outlive their expiry.
The token is sealed before the admission transaction starts. If no vault key
is configured, workflow_invoke is rejected: the run credential store never
writes plaintext. (The LONG store currently writes key id plaintext when no
keys are configured. Those rows are unchanged by this work; the keyring is
configured in the local Compose stack.)
Admission hook. Inside the start_invocation_with_stage transaction:
- on Accepted: take the per-Tool advisory lock, run the admission-limit
counts, then insert the sealed credential and the action authority row with
credential_kind = 'invoke'; - on an in-flight replay (re-attach): replace the stored credential only
if the row exists and the new token’s
expis later; - on a terminal replay: write nothing.
Action authority without a grant. The permit ledger that links nested
calls (workflow_action_authority_t) requires a non-nil grant_id, which
today comes from the LONG registration. For workflow_invoke runs the
authority row uses the run id as grant_id, the run’s deadline_ts and an
action limit from runtimeBounds. credential_kind records where the run’s
token comes from: invoke, long (set by the LONG producer) or broker
(existing rows without a LONG credential). The migration backfills existing
rows and then drops the column default, so every new writer must set it.
Authorization
A user calling a workflow-backed Tool needs permission for the Tool, not the
workflow-admin role.
workflow_start is an administration tool: it can start any saved
definition, so its Gateway ACL is limited to admin, host-admin and
workflow-admin. A workflow-backed Tool exposes one approved workflow
version with fixed limits. The Tool is the capability being granted, and its
ACL (admin, host-admin, genai-admin, or whatever the Tool owner sets)
decides who may call it. Requiring workflow-admin as well would give every
Tool user the power to start any workflow, which is the opposite of least
privilege.
workflow_invoke is not ACL-checked against user roles at the Gateway,
because no client calls it. Workflow authorizes it in three layers:
- Caller. The
x-scope-tokenservice id must be inallowedCallerServiceIds, as forworkflow_start. The Gateway does not listworkflow_invokeintools/listand refuses it as a direct clienttools/call. The manifest permissionworkflow.instance.invokeis a catalog label; there is no separate scope check. - Binding. The Tool must have an active, owner-approved revision, and its definition version must be active.
- Caller policy. If the revision has a
callerPolicy, the user token must satisfy it, for example{"anyRole": ["claims-agent", "genai-admin"]}.
Root calls are admitted with no broker, peer or policy, the same as
workflow_start. Workflow cannot tell a Portal call from a direct client
call by the scope token alone, because the Gateway sends its own token in
both cases; that is why the Gateway refusal in layer 1 matters.
The caller policy exists because the Tool ACL lives in the Portal and the Tool owner can widen it later without new approval. The policy is part of the approval digest, so the workflow owner can pin who may reach the workflow, and Workflow enforces it on every call. An empty policy means the owner accepts the Tool ACL. The review form shows the Tool’s current ACL next to the caller policy.
Input
{
"stableToolRef": "0198a3c2-...",
"expectedBindingDigest": "sha256:...",
"expectedDefinitionDigest": "sha256:...",
"input": { "claimId": "C-1001" },
"idempotencyKey": "optional; required when the revision's key kind is explicit or business"
}
The Gateway sends no deadline, budget, class or depth. Workflow derives them
from the revision. parentActionId is reserved for nested calls and is
refused for now (see Parent action).
Trust split
| Value | Source |
|---|---|
| Tool identity, input, explicit idempotency key | Gateway arguments, after the Gateway ACL and schema checks |
| Binding digest, definition digest | Gateway arguments, used only as pins that must match |
| Caller identity | User token and Gateway scope token, verified by Workflow |
| Mode, wait, deadline, class, budget, delegation, idempotency policy, cancellation, limits | Workflow binding revision |
Admission
Workflow must:
- Load the active revision for
stableToolRefand the caller’s host. - Compare
expectedBindingDigestandexpectedDefinitionDigestto it. On mismatch, returnWORKFLOW_DEFINITION_MISMATCH. This detects a Gateway snapshot that is ahead of or behind Workflow. - Refuse a retired definition version with
WORKFLOW_DEFINITION_RETIRED. - Check
callerPolicy. - Look up the idempotency key. An active key attaches to the running instance, replays its result, or conflicts, before any capacity check (see Idempotency).
- Build the
StartInvocationRequestfrom the revision:modefrominvocation_mode,execution_classfrom the revisiondeadline_ts= now +total_deadline_msbudgetfromruntime_boundspermit_depth0cancellation_policyfrom the revisionbinding_id= the revision id
- Seal the user token, then admit through
start_invocation_with_stagewithAdmissionProfile::WorkflowBacked, soverify_binding, orchestration and dependency validation, approval evidence andenforce_deadline_aware_admissionall run. The admission hook checksadmissionLimitsand stores the credential and authority.
Idempotency
The key scope is:
scope_digest = canonical_sha256({ v: 1, kind, hostId, toolId,
endUserSubject, key })
For kind: derived, key is the normalized input digest and Workflow
computes it; the Gateway no longer does. For explicit and business,
key is the caller’s idempotencyKey, 1–256 characters.
The scope uses the end-user subject only, not the Gateway client. The key is
there to stop one user from starting duplicate runs of the same Tool with the
same input. principal_subject is still stored as the client principal,
because status and list ownership use it. The existing claim function
compares it, so the same user and key through a different client returns
WORKFLOW_IDEMPOTENCY_CONFLICT while the key is active. That is the safe
answer for writes.
workflow_invoke uses a new claim function,
workflow_claim_idempotency_v2; workflow_start keeps v1. “Same” below means
the same Tool, client principal, end user, definition and input:
| Existing reservation | Outcome |
|---|---|
| none | accept |
| run not terminal, same | replay: attach to the run |
| run not terminal, different | WORKFLOW_IDEMPOTENCY_CONFLICT |
run terminal, inside result_replay_until, same | replay the stored result |
run terminal, inside result_replay_until, different | WORKFLOW_IDEMPOTENCY_CONFLICT |
| run terminal, window expired | retire the old reservation and accept a new one |
A nonterminal run is never re-accepted, even after its window has passed, so an expired window can never start a second run beside a live one. After the window expires on a terminal run, a different client or definition gets a new run, not a conflict.
Windows:
- At admission,
in_flight_untilandresult_replay_untilare bothdeadline_ts. - When the run ends,
result_replay_untilbecomes the terminal time plus the revision’sresultReplayMs, except that it is the terminal time for aFAILEDorCANCELLEDrun whoseeffect_stateisnone, so a clean failure can be retried at once.
With the read-only default resultReplayMs: 0, an identical call while the
first run is in flight attaches to it, and a call after completion starts a
new run. A write Tool must set resultReplayMs of at least 600000 for every
key kind, so a client retry within ten minutes gets the stored result instead
of a second write. A different input under the same explicit or business key
returns WORKFLOW_IDEMPOTENCY_CONFLICT.
The attach check runs before admission, so it skips capacity checks and keeps
the original deadline. Two identical calls that arrive together may both miss
it; the loser can get WORKFLOW_CAPACITY_EXHAUSTED, and its retry attaches.
Waiting, timeout and re-attach
Workflow waits for the admitted run for
min(sync_wait_ms, deadline_ts - now). The Gateway sets its HTTP timeout to
the revision’s syncWaitMs from mcp-router.yml plus a 2 second margin, and
treats a transport timeout as WORKFLOW_INVOCATION_UNAVAILABLE.
When the wait ends without output:
- If the run can still finish before
deadline_ts, Workflow returnsWORKFLOW_TIMEOUTwith theworkflowInstanceIdandretryable: true. The run continues. A retry with the same input maps to the same key and re-attaches to the running instance instead of starting another one. - If
deadline_tshas passed, Workflow cancels the run according to the revision’s cancellation policy and returnsWORKFLOW_TIMEOUTwithretryable: false.
With the default before-effects-only policy, a write that has already been
dispatched is not cancelled; the run is left to finish and its result is
kept for the replay window. Workflow does not cancel a run merely because
the caller stopped waiting.
MAX_WAIT_MS stays at 20000, matching the sync binding limit.
Output
Success:
{
"status": "completed",
"workflowInstanceId": "...",
"definitionDigest": "sha256:...",
"output": { "summary": "..." }
}
Error envelopes
There are two kinds of failure result:
- Authentication failure (HTTP 401: a token is missing or does not verify) is a JSON-RPC error, as today.
- Everything else — an authenticated call refused by policy (today’s
403), an admission refusal, a timeout or a failed run — is a normal MCP
result with
isError: trueandstructuredContent:
{
"status": "failed",
"workflowInstanceId": "...",
"error": {
"code": "WORKFLOW_TASK_FAILED",
"message": "task fetchClaim failed",
"retryable": false,
"afterEffect": false
}
}
status is failed, timeout, cancelled or rejected.
workflowInstanceId is present once a run exists. retryAfterMs is added
for capacity errors. details is an optional object with code-specific
data, passed through unchanged by the Gateway and the Portal client; the
equal-revision synchronization conflict uses it for Workflow’s stored
appliedRevision and digest. afterEffect is true when the run’s
effect_state is not none. The text content is "<code>: <message>". For a failed run,
error is the run’s stored normalized error, not a generic message.
mcp_api::handler_error and ApiError change accordingly: they currently
map 401 and 403 to JSON-RPC -32001 and wrap other failures in the detail
text.
Error mapping
| Condition | Code | Retryable |
|---|---|---|
| Input fails the definition schema | WORKFLOW_INPUT_INVALID | no |
| No active revision, or admission refused | WORKFLOW_START_REJECTED | no |
| Binding or definition digest mismatch | WORKFLOW_DEFINITION_MISMATCH | no, until the snapshots converge |
| Definition version retired | WORKFLOW_DEFINITION_RETIRED | no |
Policy refused, invalid user or Gateway identity, revoked binding, or parentActionId supplied | WORKFLOW_POLICY_DENIED | no |
| User fails the revision’s caller policy | WORKFLOW_POLICY_DENIED | no |
| Interactive capacity full, or a concurrency or rate limit hit | WORKFLOW_CAPACITY_EXHAUSTED, with retryAfterMs | yes |
| Same key, different input or different client | WORKFLOW_IDEMPOTENCY_CONFLICT | no |
| Wait ended, run still within deadline | WORKFLOW_TIMEOUT | yes, re-attaches |
| Deadline passed | WORKFLOW_TIMEOUT | no |
| Run cancelled | WORKFLOW_CANCELLED | no |
| Task failed | WORKFLOW_TASK_FAILED | per stored error |
| Output fails the response contract | WORKFLOW_OUTPUT_INVALID or ..._AFTER_EFFECT | no |
| Budget exhausted | WORKFLOW_BUDGET_EXHAUSTED or ..._AFTER_EFFECT | no |
| Workflow unreachable (Gateway only) | WORKFLOW_INVOCATION_UNAVAILABLE | yes |
The Gateway passes Workflow’s code and message through to the MCP client. It
uses WORKFLOW_INVOCATION_UNAVAILABLE only for transport failures, and the
same rule applies to workflow_get_status, workflow_get_result and
workflow_cancel.
Parent action
An action is one outbound Tool call made by a running workflow task. The parent action reference is the ID of that call. It links a child workflow run to the task that caused it, so the child inherits the parent’s authority instead of starting as an unrelated root.
The flow already exists in the code:
- Workflow records the action. Before a task calls a Tool
(
bound_mcp.rs), Workflow installs a permit inworkflow_action_permit_t. The permit records the run, user, grant, Tool, contract digest, depth, maximum depth, execution class, deadline and budget limits (workflow_action::Binding). - The task calls the Gateway. The request carries the user token, the
Workflow app token in
x-scope-token, andx-workflow-action: <action_id>. - The Gateway inspects the action.
dual_identityreads the header andaction_gateway.rsasks Workflow for the stored permit over the mTLS action dispatch. The Gateway rejects the call if the permit’s contract digest does not match the Tool, and uses the permit’s depth and class for its own limits. - The Tool is workflow-backed. The Gateway calls
workflow_invokewithparentActionId = action_idand forwardsx-workflow-action. - Workflow verifies the parent (
rule_api.rs,receiver_parent), through the verified dual-identity peer path and the action settings.
The #415 checkpoint stopped wiring the mTLS action dispatch in production
(load_mcp_router_runtime), so step 3 cannot run. Until it is restored,
Workflow refuses any workflow_invoke with parentActionId with
WORKFLOW_POLICY_DENIED “nested workflow-backed Tool calls require the
workflow action dispatch”, and never downgrades it to a root start. Nested
workflow-backed Tools are deferred work, with their own design for deriving
the child’s depth, class, deadline and budget from the parent permit.
The purpose is containment. Without the link, a workflow could call a Tool that starts another workflow, which calls a Tool that starts another, each with a fresh 30-second deadline and a full budget.
Gateway Changes
execute_workflow_toolcallsworkflow_invokeinstead ofworkflow_startfollowed bywait_for_workflow_start_result. The wait loop is removed.McpWorkflowBindingConfigshrinks to what the Gateway uses: Tool name,stableToolRef, input schema,workflowDefinitionId,workflowVersion,definitionDigest,bindingDigest,invocationMode,syncWaitMsandresultTextMode.bindingDigestis optional in the parser; a Tool without it is refused withWORKFLOW_START_REJECTED“republish this Tool”, and the other Tools keep working. The Gateway keeps the request-size pre-check fromruntimeBounds.maximumRequestBytesas an early rejection; Workflow still enforces the full budget.- A direct
workflow_startsuccess proves the definition can start; it does not publish a workflow-backed Tool revision. The Tool must first receive an approvedworkflow_binding_publishrevision, then a Gateway Tool publication containing itsbindingDigestmust be activated. The Tool page disables Invoke while Portal reports that its binding needs publication. A confirmed Gateway error withafterEffect=falseis shown as a rejection, separate from the warning for an unconfirmed network outcome. - The Workflow Definition list’s binding count is the count awaiting owner review, not the number of Portal source bindings. Its owner binding page lists requested Workflow revisions; a legacy binding with no request time does not appear there until a binding publication succeeds.
- Legacy Portal source bindings can carry
inFlightDedupMsand a zeromaximumCostUnits; neither is accepted by the current binding publication contract. Portal normalizes these values when building a new publication request. A D21 operation already stored with legacy arguments is immutable: repeating Retry cannot change its payload. Such a pending operation needs explicit outcome reconciliation before a new operation for the same binding can be submitted. - The Gateway no longer computes the idempotency key.
- Workflow
isErrorresults and theirstructuredContentpass through as described in Error envelopes. - The Gateway forwards the verified user bearer and supplies its own app bearer
in
X-Scope-Tokenon calls to Workflow. workflow_invokeis hidden fromtools/listand refused as a clienttools/call.loadRuntimeWorkflowToolinConfigPersistenceImpltakes the digests from the Workflow receipts instead of recomputing them.
Other Contract Changes
cancellation_policy,admission_limits,caller_policy,tool_annotations. Added to the normalizer, the Portal binding table and the Workflow revision.- Portal publication columns. The Portal binding row stores
publication_status,published_binding_digest,published_definition_digest,published_revision_id,published_aggregate_version,published_request_digest,publication_comment,publication_status_tsandpublication_decided_by;gateway_tool_binding_tstoresbinding_digest. - Portal synchronization tables.
workflow_sync_state_tholds the desired revision and actor and the acknowledged revision and digest per definition and kind (see Synchronization).scheduler_lock_tgets a seed row for lock id 2.workflow_operation_tis the operation ledger for every other mutating Workflow call (see Lost receipts). Both are excluded from snapshots, likeoutbox_message_t. workflow_wait_result. Removed from the Gateway path. It stays forworkflow_startcallers, but must return the stored normalized error for failed and cancelled runs instead of null, and gets itsexamples.jsonentry sovalidate-contracts.mjspasses.workflow_start. Its semantics do not change. It stays async with thePortalExecutionprofile, gains the optionalexpectedDefinitionDigest, and rejects a caller-suppliedstableToolRef, so a Tool call cannot bypass binding admission through it.
Implementation Plan
The step-by-step plan, with gates and schema, is kept outside this repository in the workflow-invoke implementation plan. In outline:
- Workflow schema and contracts. Migration
0018, manifest entries, schemas and examples, error envelopes. - Workflow publication. Publisher assertion, definition save, publish and retire, grants sync, binding revisions, digests, effect matrix, task evidence, pinned readers, owner approval, carry-over, decisions and concurrency.
- Portal. Database patch, normalizer, Workflow client with the acting user token, save-then-start, per-Tool publication, preview digest, approval commands and queries, portal-view screens.
workflow_invoke. Admission, sealed run credential, admission limits, authority, run credential selection, broker removal, wait, result and envelopes.- Gateway. Route to
workflow_invoke, shrink the binding config, pass errors through, forward the user token, supply Gateway’s app token, hideworkflow_invoke. - Cleanup.
workflow_starthardening, FDW sync removal, documentation, light-portal-test, local Compose image tag, user runbook.
Qualification Checklist
- A second identical sync call after completion starts a new run when
resultReplayMsis 0, and replays within a positive window. - A retry after
WORKFLOW_TIMEOUTre-attaches to the same instance. - A clean
FAILEDrun can be retried at once; a failure after an effect replays for the window. - A run past
deadline_tsis cancelled before effects and reportsWORKFLOW_TIMEOUT. - A failed task returns
WORKFLOW_TASK_FAILEDwith the stored message on the first call, not after the Gateway timeout. - A
workflow_invokewithparentActionIdis refused withWORKFLOW_POLICY_DENIED. - A Gateway snapshot with a stale
bindingDigestgetsWORKFLOW_DEFINITION_MISMATCH; one withoutbindingDigestgetsWORKFLOW_START_REJECTEDfor that Tool only. - A publication call without a valid user
Authorizationor GatewayX-Scope-Tokenis refused before storage access. A mismatched host is refused. - A delayed save or grant set with a lower
sourceRevisionreturnsstaleand changes nothing. - With Workflow stopped, a grant revocation shows “Sync pending” and the Portal and Workflow revision gap. Once Workflow returns, a user presses Sync now; subsequent dispatches then honor the acknowledged revocation.
- A publish whose receipt is lost is shown as unconfirmed; a retry by the
same user with the same
operationIdreturns the stored receipt and does not overwrite a newer Portal projection. Another user cannot retry it. - When revision N+1 is acknowledged before a delayed receipt for N, the Portal keeps N+1 and its digest.
- An expired idempotency window on a terminal run starts a new run; a nonterminal run is never re-accepted.
- Publishing the same payload twice returns
unchanged. Reusing anoperationIdwith a different payload is a conflict. Publishing a new digest for a published version is rejected. - A run admitted before a republish or revocation finishes with its pinned revision.
- A retired definition version refuses new runs and cannot be retired while an active revision pins it.
- Editing a definition in the Portal and starting it, from the definition
list or the editor, runs the edited text; a stale head returns
WORKFLOW_DEFINITION_MISMATCH. - An acknowledged grant revocation stops the next call of a running workflow that needs it.
- A failed binding publish leaves the previous Gateway entry for that Tool active, and the other Tools in the batch are published.
- The local stack works with
workflow-projection-syncremoved and no FDW server configured. validate-contracts.mjspasses.- A revision published by someone other than the definition owner is
pendingApproval, is left out of the Gateway snapshot, and becomes active only after the owner approves the reviewed digest. - A name-only change keeps the approved revision active. A deadline change makes a new pending revision while the previous one keeps serving.
- A revoked revision returns
WORKFLOW_POLICY_DENIEDon the next call. - A user with
genai-adminbut notworkflow-admincan call a workflow-backed Tool, and a user outside a revision’s caller policy cannot. workflow_invokemakes no light-oauth call. The run’s outbound calls carry the original user token, and the stored token is gone after the run ends.workflow_invokeis rejected when no vault key is configured.- A
workflow_startrun whose original token is still valid makes outbound calls with that token and does not call light-oauth. Once the token is within the margin, it switches to the exchanged token. - No outbound call uses the retired grant broker.
- A binding with
invocationMode: asyncor a human task is rejected. - A write Tool with
resultReplayMsbelow 600000 is rejected. A read-only Tool reaching a write is rejected. A write Tool with evidence is admitted. - Concurrent and per-minute limits return
WORKFLOW_CAPACITY_EXHAUSTEDwithretryAfterMs, re-attaches do not count, and parallel admissions cannot exceed the cap. - Re-pinning to a new version with unchanged fields stays active; a version
that exceeds
maximumParallelismis rejected at publish; a version published withreapprovesends the revision to the owner. - A client
tools/callforworkflow_invokeis refused by the Gateway.
Future: Gateway to Workflow mTLS
Customers may run their own Workflow instance and connect it to the hosted Gateway. That needs a zero-trust link: mutual TLS between the Gateway and Workflow, on top of the user and app tokens. Parts already exist:
dual_identityapp profiles accept either pinned leaf fingerprints or a trusted CA.- Workflow registers the Gateway’s mTLS peer fingerprint in
workflow_gateway_owner_tand checks it for nested calls.
Still to design:
- Initial distribution. How a new Workflow instance gets its first key and certificate and learns the Gateway’s identity, without copying secrets by hand. One option: the instance generates its key locally and enrols with a one-time token through controller-rs or the Light Identity Issuer, which signs a short-lived certificate.
- Rotation. Short certificate lifetimes with automatic re-enrolment
before expiry, and fingerprint updates in
workflow_gateway_owner_twithout an outage. - Revocation. How a customer or the operator cuts off a compromised instance.
With mTLS, Workflow can also tell the Gateway apart from other callers by
peer identity, and nested calls can use the restored action dispatch. The
workflow_invoke and publication contracts do not change.
Decisions
- Publication goes through the Gateway, which connects to Workflow. Per-host Workflow discovery through controller-rs comes later and does not change the contract.
- Publication is not a workflow.
- Nested workflow-backed Tool calls are refused until the mTLS action dispatch is restored; they are designed separately.
- Admitted runs keep their binding revision. Grants are read live.
- The idempotency scope uses the end-user subject. The same user through a different client conflicts while the key is active.
- The Gateway is the only caller of
workflow_invoke. Agents use the workflow-backed Tool. Root calls use the existing caller check; there is no separateworkflow.instance.invokescope check. - The workflow owner approves revisions from other Tool owners. The pending revision and the decision live in Workflow; the UI is in portal-view.
- Bindings carry concurrency and rate limits that the owner approves.
- Approval carries over to a new definition version when nothing it covers changed, the owner is unchanged and the version fits the approved limits.
- Write workflows can be Tools, under the effect matrix and a replay window of at least ten minutes.
- Workflow-backed Tools are sync only.
workflow_invokepasses the original user token and the Gateway token. There is no light-oauth registration or token exchange. Token lifetime is set by configuration. - Every run uses the original user token until it is near expiry, then the LONG exchanged token if the run has one. The grant broker is removed.
- The proposed
admissionLimitsdefaults are accepted. - Tool permission is enough to call a workflow-backed Tool. The workflow owner can narrow it with the revision’s caller policy.
- Gateway-to-Workflow mTLS is future work (see Future: Gateway to Workflow mTLS).
- Portal-authoritative publication carries the acting user’s bearer in
Authorization. Gateway verifies the user, forwards that bearer and adds Gateway’s application bearer inX-Scope-Tokenfor Workflow to verify. - Portal keeps Workflow’s saved definition head in step with
workflow_definition_savethrough durable, revisioned delivery.StartWorkflow, including the editor’s Start, runs only the acknowledged revision. - Grants belong to the definition and are replaced with
workflow_definition_grants_sync. A revocation is enforced once Workflow acknowledges it; dispatched operations are not undone. - Every published binding is an immutable Workflow revision; Workflow generates task evidence on approval.
- The Portal records every other mutating Workflow call in an operation
ledger. A lost receipt is recovered only by the requesting user resending
the same request with the same
operationId; there is no unattended recovery.
Open Questions
None at present.
Native Agent Call
Status
Recommended platform boundary.
call: agent is currently a native light-workflow task. It does not invoke a
running light-agent container. The workflow engine loads the portal agent
definition, selected skills, and skill tools from the database, builds a bounded
model prompt, calls the configured model provider directly, validates the JSON
output, and continues the workflow.
Containerized light-agent remains the interactive agent runtime. It serves
chat clients, keeps session memory, loads its effective catalog, and calls MCP
tools through light-gateway.
This page defines how both models should coexist in an enterprise platform.
Problem
The platform has two useful agent execution models:
- native agent tasks inside
light-workflow - containerized
light-agentservices
Both can use the same portal-authored concepts: agent definitions, skills, tools, workflow mappings, and gateway-routed API capabilities. They should not be treated as interchangeable runtime paths.
The main design question is whether a workflow should keep executing
call: agent natively or call a containerized light-agent service for every
agent step.
Current Behavior
When a workflow contains:
do:
- review-offer:
call: agent
with:
agent: com.networknt.agent.offer-1.0.0
skill: offer-decision
input:
customerId: "${ .customerId }"
profile: "${ .profile }"
outputSchemaRef: offerDecision
light-workflow handles the task itself:
- Resolve the agent by
agent_def_idor agent API name. - Load active skills assigned to the agent from
agent_skill_t. - If a skill is specified, narrow the prompt to that skill.
- Load skill tool metadata from
skill_tool_t,tool_t, andtool_param_t. - Build a bounded prompt from workflow context, skill instructions, optional task instructions, and the expected output schema.
- Call the model provider configured on the portal agent definition.
- Parse and validate the model response as JSON.
- Return the structured output to the workflow context.
The native task does not:
- call the
light-agentHTTP or WebSocket endpoint, - use
light-agentsession memory, - let the model run a dynamic gateway tool loop,
- execute tool calls from the model response.
Skill tools are included as guidance and future-routing context. In the current
runtime phase, API orchestration remains explicit workflow tasks such as
call: http, call: mcp, assert, switch, and ask.
Native Agent Tasks
Native agent tasks are best for bounded reasoning where the workflow remains the system of record.
Good examples:
- classify a request,
- normalize user-provided input,
- summarize API results,
- choose between workflow branches,
- draft a customer-facing explanation,
- assess whether human approval is required,
- produce structured output that must match a schema.
Benefits:
- Strong auditability: workflow records input, output, status, retry, and failure state.
- Deterministic orchestration: API calls, approvals, assertions, and retries stay in the workflow definition.
- Easier governance: output schemas and workflow-owned context constrain the model.
- Lower operational coupling: the task does not depend on a separate agent service instance being healthy.
- Better replay and diagnostics: the workflow engine owns the execution state.
Tradeoffs:
- It is not the full
light-agentruntime. - It does not use chat session history or Hindsight memory.
- It can duplicate some prompt/catalog handling from
light-agent. - Model provider scaling is tied to
light-workflow. - Dynamic tool selection is intentionally limited.
Containerized Agents
Containerized agents are independently deployed light-agent services.
They are best for interactive or autonomous agent behavior where the agent runtime itself is the product surface.
Good examples:
- user-facing chat agents,
- long-lived specialist agents,
- agents that need session memory,
- agents that should cache and refresh their effective catalog locally,
- agents that need a dynamic
tools/listandtools/callloop throughlight-gateway, - agents that must scale independently from workflow execution.
Benefits:
- Real agent runtime behavior: memory, chat sessions, local catalog cache, and gateway tool execution.
- Independent deployment, scaling, health checks, and versioning.
- Clear service identity through controller registration.
- Better fit for interactive clients and long-running conversational work.
Tradeoffs:
- Harder workflow audit if the agent internally decides which APIs to call.
- More distributed failure modes: network errors, timeouts, retries, and partial progress.
- Requires strict request and response contracts.
- Requires idempotency, correlation IDs, auth scopes, and timeout policy.
- Can make the workflow less deterministic if the agent is allowed to run an open-ended tool loop.
Recommendation
Keep the mixed approach, but make the boundary explicit.
Use native call: agent for bounded reasoning inside workflow-controlled
processes. Use workflow tasks and subworkflows for API orchestration. Use
containerized light-agent for interactive chat and specialist runtime agents.
The recommended enterprise pattern is:
main workflow
-> call: mcp or call: http for deterministic API access
-> run/start subworkflow for reusable skill-backed API orchestration
-> call: agent for bounded reasoning over workflow-owned context
-> ask/assert/switch/retry/audit in workflow
chat client
-> containerized light-agent
-> effective catalog from portal-query
-> tools/list and tools/call through light-gateway
-> session memory and chat history
Do not route every workflow agent step through a containerized agent by default. That would move too much process control into agent services and make enterprise audit, replay, and approval harder.
Do not remove native call: agent. It is the right primitive for workflow-owned
reasoning steps.
Skill To Workflow Pattern
For skills that require API orchestration, prefer mapping the skill to a workflow or subworkflow.
Example:
skill_t: customer-profile-review
-> skill_workflow_t: customer-profile-enrichment-v1
-> wf_definition_t: workflow that calls gateway MCP tools
In that pattern:
- the skill describes when and why to use the capability,
- the workflow owns the API call sequence,
light-gatewayexecutes MCP tool calls,- native
call: agentcan summarize or classify the results, - the workflow remains the audit boundary.
This is the preferred model for enterprise API access because it prevents an agent from inventing an unreviewed process path.
Demo Guidance
The current demos should be described precisely:
insurance-claim-rest-v1.yamlshows workflow-owned API orchestration with direct HTTP calls plus native agent tasks for bounded reasoning.insurance-claim-mcp-v1.yamlshows the same business flow throughlight-gatewayMCP tools plus native agent tasks for bounded reasoning.insurance-claim-headless-v1.yamlshows the deterministic regression path without human-task pauses.
The demos do not currently prove that light-workflow invokes the
containerized light-agent services. That can be added later as an explicit
runtime integration if the platform needs it.
Future Containerized-Agent Invocation
If workflow needs to call containerized light-agent services in the future,
do not silently change the meaning of native call: agent. Add an explicit
mode or task contract so operators can see which runtime path is used.
Possible options:
call: agent
with:
mode: native
agent: com.networknt.agent.offer-1.0.0
skill: offer-decision
call: agent
with:
mode: service
agent: com.networknt.agent.offer-1.0.0
skill: offer-decision
timeout: PT30S
or a separate task type:
call: agent-service
with:
serviceId: com.networknt.agent.offer-1.0.0
envTag: dev
skill: offer-decision
The service-call contract must require:
- explicit timeout and retry policy,
- idempotency key for side-effecting work,
- correlation and workflow instance headers,
- output schema validation,
- clear failure mapping to workflow status,
- portal/gateway authorization policy,
- observability across workflow, gateway, controller, and agent logs.
Decision Matrix
| Need | Preferred runtime |
|---|---|
| Deterministic API sequence | Workflow task or subworkflow |
| Gateway-routed API access | call: mcp through light-gateway |
| Bounded model reasoning | Native call: agent |
| Human approval or form input | ask task |
| Policy assertion | assert, switch, or rule task |
| Interactive chat | Containerized light-agent |
| Session memory | Containerized light-agent |
| Dynamic tool loop | Containerized light-agent |
| Enterprise audit and replay | Workflow-owned task |
Long-Term Direction
The platform should keep both execution models:
- Native agent tasks for workflow-owned reasoning.
- Containerized agents for interactive, memory-backed, independently scaled agent services.
The enterprise control rule is simple: workflows own durable process state and auditable API orchestration; agents provide bounded reasoning or interactive specialist behavior within contracts defined by the platform.
Execution Backends And Sandbox Execution
Status
Proposed product design.
light-workflow should support multiple execution backends for tenant-authored,
automation-heavy, and developer-local workflows. The workflow engine remains
the durable orchestrator and policy authority. Effectful work is dispatched
through the leased runner boundary defined in the
Light-Workflow Runner design, and the
runner uses a capability-described ExecutionBackend selected by the effective
policy.
Not every backend is a security sandbox. Cube Sandbox and Docker Sandboxes use microVM boundaries. Rootless OCI containers and ordinary Kubernetes Jobs share a host kernel unless a stronger runtime is configured. Fedora Toolbx is a host-integrated developer environment and explicitly is not a sandbox. The policy model must preserve these differences instead of treating every backend as interchangeable.
Problem
Workflows can be created by tenants and can eventually include tasks that run commands, scripts, containers, model calls, MCP tools, browser automation, or release automation. Those capabilities are useful, but they are also the highest-risk part of the workflow runtime.
The platform needs a way to say:
- whether a workflow can request effectful execution,
- where each task is allowed to run,
- which minimum isolation boundary and host-integration limits apply,
- whether sandboxed tasks may share a workspace,
- which command, image, resource, network, filesystem, artifact, and secret policies apply,
- how task claims and remote execution remain correct across crashes and retries,
- how release workflows can keep build state without exposing publish or signing credentials to tenant-controlled code.
Architecture And Ownership
Use one authoritative execution path:
workflow start event
-> light-workflow
- creates workflow and task state
- resolves and persists the effective policy snapshot
- owns branching, retries, cancellation, and audit
-> controller-rs
- authenticates runners
- issues and renews fenced task leases
- rejects stale task reports
-> light-workflow-runner
- validates the lease and effective task policy
- invokes the selected ExecutionBackend
- streams bounded logs and reports normalized results
-> execution backend
- prepares or resumes the execution environment
- enforces its approved isolation, resource, network, workspace,
credential, and lifecycle policy
- executes the approved command specification
Component ownership is:
light-workflow: Workflow state, policy snapshots, task attempts, transition decisions, retry decisions, cancellation state, and durable audit.controller-rs: Runner identity, admission, capabilities, lease ownership, lease renewal, fencing, and quarantine.light-workflow-runner: Effectful task execution, backend selection from the lease, backend API credentials, log streaming, artifact transfer, and result normalization.ExecutionBackend: Capability-described adapter for a microVM sandbox, shared-kernel container, Kubernetes Job, dedicated VM, host-integrated environment, or fixed external action.
light-workflow should not hold backend control-plane credentials in the SaaS
topology and should not implement separate direct protocols for Cube, Docker,
Kubernetes, or other substrates. A local installation may colocate the runner
and backend adapter with
light-workflow, but it must preserve the same durable attempt, lease, fencing,
policy, result, and audit contracts.
Goals
- Keep workflow orchestration outside effectful execution environments.
- Keep tenant-authored code outside the SaaS workflow process.
- Make the effective policy server-owned, immutable, and auditable.
- Support multiple backend purposes without weakening minimum isolation.
- Support per-task and per-workflow execution lifecycles where the backend has those capabilities.
- Support long-running tasks without duplicate execution caused by stale locks.
- Clean up execution resources autonomously when a runner cannot reach the control plane.
- Queue temporary capacity shortages without busy retry loops or consuming a workflow retry attempt.
- Fail closed when a required backend capability or policy control is absent.
- Keep raw release and signing credentials away from arbitrary workflow code.
- Treat agent-generated workspace changes as untrusted output and validate them against a server-owned path policy.
- Generate verifiable build provenance for release artifacts without exposing attestation credentials to tenant-controlled code.
Non-Goals
- Do not expose vendor or backend APIs directly through the workflow DSL.
- Do not let workflow metadata define raw backend network rules or backend credentials.
- Do not describe Toolbx or an ordinary shared-kernel container as equivalent to a microVM security boundary.
- Do not store security policy or execution lifecycle state in workflow context.
- Do not promise exactly-once external side effects. The runtime provides fenced at-least-once execution plus explicit reconciliation.
- Do not treat a fresh sandbox as sufficient authorization for publishing or signing.
First Schema Surface
Use existing metadata fields first so the design can be introduced without an
immediate workflow-core schema break. WorkflowDefinitionMetadata already
has document.metadata, and every task has metadata through
TaskDefinitionFields.
Workflow metadata requests security requirements through an approved profile. It does not directly select a backend credential, mutable image name, vendor, or raw network policy:
document:
dsl: "1.0.3"
namespace: release
name: light-fabric-polyrepo-release
version: "1.0.0"
metadata:
lightWorkflow:
runner:
runnerPool: release
security:
schemaVersion: 1
executionProfile: release-sandbox
profileVersion: 7
placement: runner
isolation:
minimumBoundary: microvm
allowedHostExposure: []
workloadTrust: untrusted
sandbox:
sessionScope: workflow
workspace:
mode: copy-on-write
containerEngine:
access: private-daemon
network:
protocols:
- https
credentials:
delivery: proxy-injected
A task may request stricter isolation and an approved command template:
do:
- publish-github-release:
run:
shell:
command: light-release-publish
arguments:
- "${ .artifactSetId }"
- "${ .version }"
metadata:
lightWorkflow:
runner:
commandTemplateId: light-fabric-release-publish-v1
security:
sandbox:
sessionScope: task
reason: release-credential-isolation
approval:
required: true
bindTo:
- artifactSetDigest
- releaseTarget
credentials:
- github-release-oidc
The command template is the authority. The command and arguments in the workflow must match the approved template after expression resolution. The runner rejects a mismatch rather than executing arbitrary text.
approval.required makes the task ineligible for runner scheduling until
light-workflow has persisted a matching approval. It does not instruct a
runner to claim the task and wait. The eventual lease references the already
validated approval and represents a new fixed-action attempt.
Unknown lightWorkflow.security fields, unsupported schema versions, invalid
types, and unapproved profile versions must fail definition validation. They
must not be ignored.
Later, a first-class field can normalize into the same internal policy object:
security:
schemaVersion: 1
executionProfile: release-sandbox
profileVersion: 7
placement: runner
isolation:
minimumBoundary: microvm
allowedHostExposure: []
workloadTrust: untrusted
sandbox:
sessionScope: workflow
workspace:
mode: copy-on-write
Policy Dimensions
Placement, isolation boundary, backend selection, routing, session scope, and
workspace reuse are separate decisions. They must not be combined into one
mode value.
Isolation Boundary
host-integrated
The environment deliberately shares host facilities such as the user’s home directory, session services, devices, sockets, or host networking. Fedora Toolbx belongs in this category. It is useful for trusted local development and troubleshooting, but it is not a security boundary for tenant-authored code.
shared-kernel-container
The task runs in an OCI container that shares the host kernel. Rootless Docker or Podman and a default Kubernetes container belong here. This boundary is appropriate for trusted build and packaging tasks when capabilities, mounts, syscalls, resources, and networking are constrained. It must not satisfy a profile requiring a separate kernel.
microvm
The task runs with a separate guest kernel in a lightweight VM. Cube Sandbox and Docker Sandboxes belong here. A microVM can satisfy untrusted-code profiles only when workspace, network, credential, lifecycle, control-plane, and cleanup requirements are also enforced.
dedicated-vm
The task runs on a separately provisioned VM dedicated to an approved tenant, workflow, or runner pool. This can support privileged or long-running workloads, but image provenance, teardown, attestation, and network isolation remain required.
external-service
The task calls a fixed service-owned action such as publishing, signing, or deployment. It is not a general command environment. Authorization derives from the action contract and immutable inputs rather than from shell isolation.
The effective profile sets a minimum boundary and an allowedHostExposure
allowlist. An empty allowlist means the backend may expose no host facilities.
The backend’s approved hostExposure set must be a subset of that allowlist.
Boundary names are not the complete security decision. For example, a microVM
with a writable host workspace or raw credentials may be unsuitable for a
high-risk task.
Boundary matching uses a server-owned compatibility relation, not simple enum
sorting. An approved dedicated VM may satisfy a separate-kernel requirement,
but external-service is comparable only to a fixed-action requirement, and a
backend cannot claim a stronger boundary merely by changing its registration.
workloadTrust is trusted or untrusted. Tenant-authored code, generated
code, dynamic agent tools, and content from an untrusted repository default to
untrusted; workflow metadata cannot mark them trusted. An untrusted workload
requires an approved compatibility record with supportsUntrustedCode=true in
addition to the boundary, workspace, network, credential, and lifecycle
requirements.
Backend Taxonomy And Intended Use
| Backend | Boundary | Intended use | Untrusted tenant code |
|---|---|---|---|
| Cube Sandbox | microvm | Remote or clustered tenant sandbox sessions, snapshots, controlled egress | Allowed only with the Cube production baseline |
Docker Sandboxes (sbx) | microvm | Autonomous coding agents and isolated Docker builds | Allowed with clone workspace mode and enforced policy |
| Rootless Docker or Podman container | shared-kernel-container | Lightweight trusted CI, tests, packaging, and tools | Not by default |
| Kubernetes Job | Declared by approved runtime class | Scalable jobs in a tenant or service cluster | Only when the selected runtime and node policy satisfy the required boundary |
| Fedora Toolbx | host-integrated | Trusted developer tooling and host troubleshooting | Never |
| Dedicated VM | dedicated-vm | Privileged, tenant-dedicated, or long-running work | Allowed when its approved profile satisfies the task requirements |
| Publisher, signer, or deployer service | external-service | Fixed irreversible actions over immutable inputs | No arbitrary code surface |
Toolbx usually does not need a per-task backend implementation. A trusted local
runner may itself run inside Toolbx and register a host-integrated backend
with execution session scope none. The policy must record its home, device,
D-Bus, socket, and network exposure and prevent it from claiming isolated or
secret-bearing tasks.
An ordinary Docker container and Docker Sandboxes are different backends. A container shares the host kernel. Docker Sandboxes place an autonomous agent inside a microVM with a private Docker daemon. No backend may mount the host Docker socket for tenant-authored execution.
Backend Capability Contract
Every registered backend has an operator-approved compatibility record. A representative capability document is:
{
"backendId": "docker-sbx-local",
"kind": "microvm",
"implementation": "docker-sandboxes",
"version": "approved-version",
"isolationBoundary": "microvm",
"supportsUntrustedCode": true,
"workspaceModes": ["direct", "clone"],
"hostExposure": [],
"networkEnforcement": ["deny-by-default", "http-l7"],
"supportedEgressProtocols": ["http", "https"],
"credentialDelivery": ["proxy-injected"],
"containerEngineAccess": "private-daemon",
"lifecycle": ["inspect", "reconnect", "cancel", "destroy"],
"sessionScopes": ["task", "workflow"]
}
This is an effective, profile-specific record, not an implementation-wide claim. It is scoped to the backend implementation and version plus the approved template, image, runtime class, node policy, workspace mode, and enforcement configuration that make the capabilities true. Registration and leases bind the digest of that exact record. A configuration or compatibility change creates a new immutable record and digest.
The minimum capability vocabulary includes:
- isolation boundary and supported trust classes,
- direct host exposure such as home, devices, D-Bus, SSH agent, localhost, container sockets, and writable workspace mounts,
- workspace modes: direct, ephemeral, clone, copy-on-write, and workflow reuse,
- network enforcement layer and supported protocols,
- credential delivery: proxy-injected, workload identity, attempt-bound local
broker, or task-unique read-only
tmpfsfile; environment-value delivery is prohibited, - container-engine access: none, private daemon, or prohibited host daemon,
- CPU, memory, disk, process, time, output, artifact, and concurrency controls,
- lifecycle inspection, reconnect, cancellation, snapshots, log cursors, operation lookup, idempotency, and cleanup,
- tenant isolation, data residency, attestation, and audit support.
Runner self-report is not sufficient. Server-owned compatibility definitions and backend conformance tests determine which capabilities are trusted. A backend may claim only task attempts whose effective requirements are a subset of that approved compatibility record.
The workflow DSL does not name a backend implementation. Policy resolution selects an eligible backend from the registered runner pool and persists the selected backend ID, implementation, version, and capability digest in the effective task policy.
Placement
host
The task runs in the trusted light-workflow process. Only control-plane tasks
and explicitly approved native calls can use this placement.
runner
The task is sent through a fenced lease to light-workflow-runner. The runner
executes it through an eligible backend permitted by the effective profile.
Execution Session Scope
none
The runner does not create an additional execution environment. This is allowed only when the effective profile explicitly permits the runner environment itself as the execution boundary.
workflow
One backend execution session is reused by approved tasks in one workflow instance. This supports a shared checkout, build output, and dependency cache. The selected backend must advertise workflow-session support. The session must never be reused across workflow instances, tenants, principals, or incompatible policy snapshots.
agent-session
One backend execution session provides a bounded interactive workspace across turns for one authenticated agent session. Reuse requires identical tenant, host, principal, agent definition, workspace base, policy digest, runtime adapter, backend compatibility, network/model/tool policy, and unexpired cleanup state. Conversation history and memory remain origin-domain state and are never recovered from the sandbox.
task
One fresh backend environment is created for one task attempt. This is stricter isolation and is required for untrusted code that must not share state, raw credential fallbacks, and high-impact operations.
Profile labels such as per-agent-call or per-publish map to task scope
plus an isolation class and additional policy requirements. They are not
separate lifecycle primitives.
Workspace Reuse
Workspace reuse is independently controlled:
ephemeral: Fresh workspace for one task attempt.workflow: Reused only within one workflow instance and policy snapshot.agent-session: Reused only for the same authenticated agent session, principal, immutable base, runtime adapter, and policy snapshot.copy-on-write: A task receives an isolated clone of an approved workspace.
Cross-tenant caches and mutable cross-workflow workspaces are out of scope. Shared dependency caches, if added later, require content-addressing, integrity verification, and separate poisoning controls.
Task Routing
Host execution remains the default for control-plane tasks:
ask
assert
set
switch
workflow context merge
task creation and transition
process state persistence
approved native call.agent without tools or file access
Runner execution is required for effectful or tenant-local task families:
run.shell
run.script
run.container
browser automation
tenant-provided code
filesystem mutation outside workflow context
external MCP server processes
command-line tools
release build and package commands
agent tasks with files, tools, or private network access
Calls that can run in more than one location require policy-based placement:
| Task | Host placement | Runner placement |
|---|---|---|
call.http | Approved SaaS endpoint and host credential boundary | Tenant-private endpoint or sandbox egress boundary |
call.jsonrpc | Approved SaaS endpoint | Tenant-private or backend-local endpoint |
call.mcp | Approved gateway endpoint | External process or tenant-local MCP server |
call.agent | Bounded model call without tools or files | Tools, files, generated code, or tenant-local data |
call.rule | Default for curated local rules | Only when an approved rule profile requires isolation |
Placement depends on endpoint identity, credential source, data boundary, required capabilities, and network policy. Egress reachability alone is not sufficient. Destination validation remains mandatory for host-executed HTTP, JSON-RPC, and MCP calls.
Unsupported task types and task/backend combinations must be rejected before execution. When the task graph can be inspected statically, definition publication should reject the workflow before any instance can partially run. Dynamic destinations and values are validated again for each attempt.
Effective Policy
The runtime computes a workflow policy snapshot at instance creation and a derived effective policy for each task attempt:
{
"policySnapshotId": "019f0000-0000-7000-8000-000000000001",
"requestedProfile": "release-sandbox",
"effectiveProfile": "release-sandbox",
"profileVersion": 7,
"policyDigest": "sha256:...",
"placement": "runner",
"runnerPool": "release",
"approvedTaskTypes": ["run.shell", "call.http", "call.mcp"],
"executionRequirements": {
"minimumBoundary": "microvm",
"allowedHostExposure": [],
"workloadTrust": "untrusted"
},
"executionBackend": {
"backendId": "cube-prod-east",
"kind": "microvm",
"implementation": "cubesandbox",
"version": "approved-version",
"capabilityDigest": "sha256:..."
},
"sandbox": {
"templateId": "tpl-immutable-id",
"templateDigest": "sha256:...",
"sessionScope": "workflow",
"workspaceMode": "copy-on-write"
},
"networkPolicyId": "release-egress-v3",
"networkPolicyDigest": "sha256:...",
"trustBundleRef": "trust-bundle://enterprise-egress-v3",
"trustBundleDigest": "sha256:...",
"credentialPolicy": "brokered-task-scoped",
"artifactPolicy": "release-artifacts-v2",
"provenancePolicy": {
"format": "slsa-provenance-v1",
"mode": "signed",
"policyDigest": "sha256:..."
},
"localCleanupPolicyDigest": "sha256:...",
"resourcePolicy": "release-build-medium-v1"
}
The snapshot is stored in dedicated runtime state, not in workflow context or task output. Audit records reference its immutable ID and digest.
Policy resolution rules are field-specific:
- Operator profile definitions provide the base allowed backend compatibility records, templates, commands, networks, trust bundles, mounts, workspace-change policies, credentials, provenance, local cleanup, limits, and placements.
- Service policy intersects the profiles and capabilities available in the deployment.
- Tenant policy further restricts the allowed set.
- Workflow metadata requests one allowed profile and version.
- Task metadata may request stricter isolation or a subset of capabilities; it cannot downgrade operator-derived workload trust.
- Allowlists are intersected.
- Explicit denies take precedence.
- Numeric resource and duration limits use the lowest permitted maximum.
tasksession scope may strengthenworkflowscope; a task cannot weaken a required execution environment tonone.- The selected backend must meet the minimum isolation boundary and every required capability. A host-integrated or shared-kernel backend cannot satisfy a microVM requirement.
- Backend and template selections must be members of the approved set. They do not have a meaningful “more privileged” ordering within the same boundary.
- Credential access is the intersection of profile, task, command template, approval, and current credential-broker policy.
Profile versions are immutable. A new operator policy creates a new version. An in-flight workflow continues with its recorded snapshot unless an emergency revocation explicitly invalidates it. Revocation must fence new attempts, cancel affected active attempts where possible, revoke credentials, and record why the snapshot was invalidated.
Profile changes that require approval must remain pending and cannot publish an active workflow definition. Runtime approvals for irreversible tasks are separate objects and must bind the approver, task, command template, artifact digest, target, policy snapshot, and expiry.
Durable Runtime State
Remote backend execution adds distributed state and requires dedicated
persistence. Do not put session IDs, leases, policy snapshots, credentials, or
backend operation IDs in process_info_t.context_data.
The controller/runner portion of this state is origin-neutral. A workflow task,
standalone agent turn, or agent action can use the same scheduling, execution
attempt, lease, backend, and cleanup contract. light-workflow and
light-agent keep separate domain tables and are the only services allowed to
advance their respective subjects. See
Light-Agent Execution for agent
session and turn ownership.
The initial storage model should include:
workflow_execution_policy_t
- workflow process and instance IDs,
- tenant, host, trigger principal, and correlation IDs,
- requested and effective profile IDs,
- profile version and policy digest,
- creation and revocation state.
execution_session_t
- authenticated origin and subject scope, including optional workflow or agent session correlation,
- backend ID, kind, implementation, version, and capability digest,
- backend environment and session IDs,
- immutable template or image ID and digest where applicable,
- workflow policy snapshot ID,
- tenant, workflow, and principal scope,
- lifecycle state/version/fence, active action/lease owner, task and lease deadlines,
- optional approval-hold ID/reason, hold expiry, policy/cost binding, pause/checkpoint state, and retained-resource evidence,
- origin idle/max expiry, policy/grant expiry, effective minimum expiry, backend-native expiry, last runner contact, last inspection time, cleanup deadline, attempt count, and cleanup state.
execution_session_cleanup_request_t
- request ID, authenticated origin, origin-session/subject correlation, and execution-session ID,
- close, revoke, expiry, policy-change, quarantine, or operator reason,
- requested, dispatched, cleaned, retryable, or operator-action state,
- attempt/fencing watermark, retry schedule, and cleanup evidence reference,
- unique active request per execution session.
execution_input_t
- immutable input ID, authenticated origin subject and optional attempt or execution session,
- kind such as context, workspace base, skill package, trust bundle, or fixed action input,
- content digest, size, media/package type, storage reference, provenance and scanner bindings, and mount/entrypoint policy,
- staging, verification, retention, and cleanup state without embedded storage credentials.
execution_attempt_t
- authenticated origin service, execution subject kind and ID, and monotonically increasing attempt number,
- optional workflow task or agent turn/action correlation,
- lease ID, runner ID, and fencing token,
- backend operation ID and command idempotency key,
- effective task-policy digest,
- state, heartbeat, started, deadline, and completed timestamps,
- normalized result, error classification, and reconciliation state.
runner_scheduling_request_t
- idempotent scheduling request ID, authenticated origin, execution subject, tenant fairness key, runner pool, and effective requirements digest,
- enqueue time, queue deadline, priority class, and scheduling state,
- short-lived capacity reservation ID, runner and backend slot, reservation expiry, and consumption state,
- cancellation, policy-revocation, and terminal admission reason.
workflow_approval_t
- approval request ID, tenant, workflow, orchestration task, and state,
- artifact set and provenance digests, release target and version, command template, policy digest, and prior-outcome reconciliation state,
- approver identity, decision, reason, creation, decision, and expiry times, and single-use nonce,
- consuming post-approval execution attempt ID or rejection and expiry transition.
Standalone agent approvals remain in agent-domain storage, but use the same immutable binding and single-use post-approval common-attempt contract. The runner never owns either approval table.
workflow_artifact_t
- immutable artifact ID, tenant, workflow, task, and attempt IDs,
- canonical name, size, media type, and trusted digest,
- storage reference, retention class, provenance statement and envelope digests, attestation signer identity, and approval bindings.
The runner also maintains a minimal durable local cleanup journal before it prepares an environment or dispatches an operation. The journal contains the lease, fencing token, backend environment and operation IDs, absolute and monotonic deadlines, backend-native expiry, and cleanup state. It contains no credential values or task payloads. Runner restart recovery and the local watchdog use this journal; the SaaS database remains authoritative for workflow state.
Security and execution audit should be append-only. Mutable status tables can reference the latest state, but must not replace the history needed to explain claims, retries, cancellation, policy changes, and cleanup.
Origin Result Wakeup
The common attempt row is authoritative. The PostgreSQL transaction that
conditionally stores a newly terminal execution_attempt_t also emits a
versioned execution_result_ready_v1 notification containing only attempt ID,
authenticated origin, subject kind, and correlation ID. It carries no result
bytes, tenant content, or authorization.
light-workflow or light-agent uses that notification only to wake a
reconciler, reloads and verifies the authoritative attempt, and conditionally
accepts it into its own domain transaction. Every origin must also run indexed
startup and periodic catch-up scans because notifications can be missed,
duplicated, or reordered. A push callback may be another wakeup later, but
correctness never depends on notification delivery and controller/runner code
never updates origin-domain state directly.
Lease, Attempt, And Fencing Model
Remote backend execution is at-least-once. Exactly-once external effects cannot be guaranteed across the backend and workflow database boundary.
Each attempt receives a short-lived lease and a monotonically increasing fencing token. The runner must renew the lease while the backend operation is active. Every progress, log, artifact, and completion report includes the lease ID, attempt number, and fencing token. The control plane rejects reports from a stale token.
Completion updates must use compare-and-set semantics against the active attempt. A late result from an expired attempt must not overwrite a newer attempt. Workflow transitions occur only after the accepted result and audit records commit.
Backend prepare and execute calls must receive stable idempotency keys when the
backend supports them. When the runner loses contact after dispatch, the
attempt enters UNKNOWN, not immediately FAILED. A reconciler inspects the
backend operation before deciding whether to accept a result, resume waiting,
cancel, or create another attempt.
The DSL idempotencyKey is part of the task contract, but it is not by itself a
guarantee. A side-effecting command template must declare how the target system
honors that key or how the runner queries the external operation before a
retry. Tasks without such a contract must not be automatically retried after
an unknown outcome.
Execution Session Lifecycle
For workflow-session execution:
light-workflowresolves and persists the workflow policy snapshot.light-workflowmarks the task ready and submits an idempotent scheduling request. It remainsPENDING_CAPACITYwithout an execution attempt when no eligible slot is available.controller-rsreserves an eligible runner and backend slot. The control plane idempotently creates the task attempt against that reservation and issues its lease with the task policy, attempt number, and fencing token.- The runner validates the lease, its own capabilities, and the approved command template.
- The runner idempotently prepares, resumes, or inspects the backend execution session scoped to the workflow policy snapshot.
- The runner dispatches the command with a stable backend operation ID and starts lease and operation heartbeats.
- Logs are streamed with bounded chunks and resumable sequence numbers.
- Declared artifacts are safely copied into controlled storage and hashed outside the sandbox trust boundary.
- The runner reports a normalized result with the active fencing token.
light-workflowaccepts the result, persists the transition, and schedules cleanup when the session is no longer needed.
The session identity is scoped to:
tenant id and host id
workflow definition id and version
workflow process and instance id
policy snapshot id and digest
trigger principal or approved service identity
runner pool and execution backend
For workflow and agent sessions, effective physical-session expiry is the earliest of the origin session idle/max expiry, execution policy, credential or broker-grant expiry, and backend-native TTL. The runner must not extend one clock merely because another has time remaining.
When an origin closes, revokes, or expires its logical session, the same durable
origin transaction creates an idempotent
execution_session_cleanup_request_t. controller-rs fences and cancels
active attempts, revokes grants, and dispatches cleanup. The runner destroys
the backend session and records evidence; retries survive controller and runner
restart. A backend-native TTL is the last fail-safe. Leaving a known-abandoned
sandbox alive until that independent TTL is a cleanup defect.
An action lease and a reused execution session are different resources. Ending
an action lease always removes executable authority, model/credential broker
access, and task-scoped grants. It cleans task scope, but it does not by itself
delete a compatible workflow or agent-session workspace.
Under an explicit non-secret retention policy, an origin may put the session in
IDLE_APPROVAL_HOLD with a durable hold ID, reason, policy digest,
holdUntil, retained-resource cost, and verified checkpoint/patch evidence.
The runner pauses or checkpoints where supported. The hold expires no later
than approval expiry, idle/max lifetime, policy/cost limit, grant boundary, or
backend TTL; it cannot be extended by fake action heartbeats. Zero active
attempts is not an abandonment signal while the bounded hold is valid. Close,
revocation, policy mismatch, hold expiry, or cleanup request still destroys the
session.
Do not reuse a sandbox across tenants, unrelated workflow instances, different policy snapshots, or incompatible principals.
A workflow session is single-writer unless the backend profile explicitly supports safe concurrent operations. Parallel workflow branches must use separate task sandboxes or copy-on-write clones unless their workspace access is serialized.
Cancellation fences the attempt first, then requests backend cancellation and environment cleanup. If cleanup cannot be confirmed, the attempt remains in a cleanup-pending state and an orphan reconciler continues inspection. A cancelled or expired attempt can never publish a valid completion afterward.
Disconnected Runner Watchdog
Control-plane reconciliation alone cannot clean resources inside a tenant network that has become unreachable. Every runner therefore has an autonomous watchdog, separate from the task worker, with these rules:
- Write and sync the local cleanup journal before creating a backend resource.
- Stop claiming and starting work as soon as the control-plane session is unavailable.
- A running task may continue only while its server-issued lease remains locally
valid. A connectivity grace period cannot extend
expiresAtor the task deadline. - At lease expiry, task deadline, cancellation observed before disconnect, or maximum environment lifetime, locally fence the attempt, revoke local credential handles, cancel the operation, and destroy or quarantine the environment.
- Use a monotonic deadline derived when the lease is received, bounded by the authenticated absolute deadline, so wall-clock rollback cannot extend execution.
- Tag every backend resource with tenant, workflow, attempt, policy digest, owner runner, and expiry. Configure a backend-native TTL or lifecycle rule where available so cleanup still occurs if the runner host itself fails.
- On startup and periodically, scan the local journal and backend-owned tagged resources, retry expired cleanup with bounded backoff, and preserve minimal evidence for unresolved operations.
- On reconnect, report cancellation, outcome, and cleanup evidence with the original lease and fencing token. The control plane may accept matching resource cleanup evidence for a fenced attempt, but only the current token can transition task outcome; an expired result remains diagnostic.
Destroying an environment cannot undo an external side effect. A disconnected
fixed action or publish operation remains UNKNOWN and must be inspected and
reconciled before retry. The control-plane orphan reconciler, local watchdog,
and backend-native expiry are complementary layers.
Timeouts, Resource Limits, And Admission
The policy must distinguish:
- task wall-clock deadline,
- total workflow deadline,
- backend API request timeout,
- sandbox idle timeout,
- maximum sandbox lifetime,
- lease heartbeat interval and expiry,
- cancellation grace period,
- cleanup and evidence-retention period.
Backend idle timeout is not a task wall-clock timeout. The runner and control plane enforce the task deadline even when backend activity keeps resetting its idle timer.
Each profile must set backend-enforced limits for:
- CPU and memory,
- writable disk and inode or file count,
- process and PID count,
- open files,
- network destinations, connections, and optional bandwidth,
- stdout, stderr, and structured result size,
- artifact count and total bytes,
- maximum concurrent tasks and sandboxes per tenant and runner pool.
Admission checks capacity and tenant quotas before a lease is issued. The runtime must provide backpressure instead of creating unbounded pending sandboxes. Cost and quota exhaustion are explicit non-command failure classes.
Capacity Queueing And Backoff
Temporary saturation is a scheduling state, not a command failure. When no
eligible runner or backend slot is available, the task remains
PENDING_CAPACITY; no task attempt, lease, sandbox, or workflow retry is
created. controller-rs owns a bounded, per-tenant fair queue and atomically
reserves capacity before issuing a lease.
The authoritative workflow task remains in light-workflow; the controller
queue stores only the idempotent scheduling request. When capacity opens,
controller-rs returns a short-lived reservation token. Attempt creation and
lease issuance bind that token through an idempotent, fenced handshake so a
lost response cannot allocate two attempts or two capacity slots.
Runner claims use long polling or server push. Empty claims and transient
backend-capacity responses carry retryAfter and use capped exponential
backoff with jitter. Capacity release wakes only a bounded number of eligible
waiters, and repeated backend admission failures open a short circuit breaker
instead of causing a thundering herd.
Hard tenant quota, policy, or cost-limit violations return
execution_admission_denied and require a policy, quota, or operator change.
Temporary saturation returns execution_capacity_deferred and remains queued
until capacity is available, the queue deadline expires, the workflow is
cancelled, or policy is revoked. Queue time counts toward the workflow deadline
but not the task wall-clock deadline, which begins only after lease acceptance.
Execution Backend Interface
The provider-neutral interface is named ExecutionBackend. It supports
capability discovery and the normalized operations needed for crash recovery:
capabilities
validate effective configuration
prepare environment idempotently
inspect environment
resume or reconnect
execute task idempotently
inspect operation
stream logs from cursor
cancel operation
export artifact safely
report measured execution evidence
create and delete checkpoint
clean up environment
Not every operation applies to every backend. A Toolbx-backed trusted runner may have no separate environment to create. An external signer exposes a fixed action rather than a shell or filesystem. Optional operations are advertised through the approved capability record; required missing operations fail policy resolution.
Capabilities include isolation and host exposure, resource and network controls, credential delivery, private container-engine access, snapshots, cancellation, log cursors, operation lookup, idempotency, measured execution evidence, attestation support, native expiry, and cleanup behavior.
Backend-specific response codes are normalized, but request IDs and raw diagnostic references remain in restricted audit data. Every backend adapter must pass a boundary-appropriate conformance suite before it is enabled.
Cube Sandbox Production Baseline
Cube Sandbox defaults are not the Light platform security policy. The Cube adapter and deployment admission must verify the following baseline:
- CubeAPI is authenticated and authorized. Authorization checks both HTTP path and method.
- CubeAPI, CubeMaster, Cubelet, WebUI, Redis, and database access are restricted to approved private networks and protected with firewall rules.
- TLS or mTLS protects control-plane traffic where it crosses a host boundary.
- Sandbox public traffic is disabled unless an approved task explicitly needs an inbound service, and that service requires its per-sandbox access token.
allow_internet_accessis false for restricted profiles.- L3/L4 allow rules and L7 HTTP/HTTPS rules are generated from the effective profile, not copied from workflow metadata.
- The effective rendered network configuration is inspected after provider template and request merging.
- Non-HTTP traffic, DNS, SSH, registry access, and provider-internal traffic are considered explicitly. L7 HTTP rules do not control arbitrary TCP or UDP.
- Templates contain the required egress CA only when TLS interception is part of the approved profile.
- Provider, control-plane, template, and SDK versions are pinned to a tested compatibility set.
The adapter must not rely on a domain list while leaving default internet access enabled. It must not allow a workflow-supplied rule to precede or weaken operator policy during provider rule merging.
Docker Sandboxes Backend Baseline
Docker Sandboxes (sbx) is a separate microVM product, not an ordinary Docker
container. Each sandbox has its own guest kernel, Docker daemon, filesystem, and
network. It is a strong candidate for autonomous coding agents and workflows
that need to build or run containers without exposing the host Docker daemon.
The approved Docker Sandboxes backend must enforce:
- clone workspace mode for untrusted or autonomous tasks; the default direct workspace mount is not an isolation boundary because changes are immediately applied to the host working tree,
- deny-by-default network policy with only approved HTTP and HTTPS destinations,
- explicit rejection of tasks requiring raw TCP, UDP, ICMP, SSH, or private network access that the backend cannot provide safely,
- host-side proxy injection for supported service credentials,
- no registry, SSH, or custom credential copied into the VM unless the effective task policy explicitly accepts that exposure,
- the sandbox’s private Docker daemon; the host Docker socket is never mounted,
- immutable backend and template or kit compatibility records,
- lifecycle inspection, reconnect, cancellation, disk quotas, and explicit removal after the workflow retention period,
- centrally managed organization policy for managed endpoints, or a locked and audited local policy for approved developer-local runners.
Docker Sandboxes is initially a developer-local or managed-endpoint backend. Before it is used as a headless service backend, the adapter must prove stable machine-to-machine lifecycle control, runner authentication, idempotent operation lookup, audit export, and cleanup through the same conformance suite used by the control plane.
Shared-Kernel Container And Kubernetes Baseline
Rootless Docker or Podman containers and default Kubernetes Jobs share a host
kernel. They are useful for high-volume trusted builds, tests, packaging, and
approved internal tools, but they do not satisfy a microvm or dedicated-vm
requirement.
The baseline requires:
- rootless execution where supported,
- no privileged containers and no host container-engine socket,
- no host PID, IPC, or network namespace,
- dropped Linux capabilities,
no-new-privileges, a restrictive seccomp profile, and the applicable SELinux or AppArmor policy, - read-only root filesystem except for declared ephemeral volumes,
- canonical allowlisted mounts with no user-controlled host paths,
- hard cgroup CPU, memory, PID, disk, time, log, and concurrency limits,
- deny-by-default network policy and tenant-scoped service identity,
- immutable image digests and image provenance verification,
- complete pod, job, container, volume, and credential cleanup.
A Kubernetes backend records its approved runtime class and node isolation.
Kata, gVisor, dedicated nodes, or another hardened runtime may qualify for a
stronger server-owned compatibility record, but the word Kubernetes alone
does not establish the isolation boundary.
Toolbx And Host-Integrated Baseline
Fedora Toolbx exists to provide a convenient mutable development and host
troubleshooting environment. It exposes the user’s identity, home directory,
network, devices, D-Bus, system journal, SSH agent, and other host facilities.
It must be classified as host-integrated, never as a sandbox.
An approved Toolbx runner profile must:
- accept only trusted operator or developer tasks,
- advertise all host integrations in its backend capability record,
- use execution session scope
none, - reject tenant-authored code, dynamic agent tools, and arbitrary scripts from untrusted sources,
- reject publish, signing, platform credential, and cross-tenant tasks,
- rely on the host user’s permissions and audit identity rather than claiming a separate security boundary,
- remain opt-in for local workflows and never be selected as a fallback when a stronger backend is unavailable.
Toolbx can be useful without a dedicated per-task adapter: the trusted
light-workflow-runner process can run inside a Toolbx environment and register
that environment’s approved host-integrated capabilities.
Dedicated VM And External Action Baseline
A dedicated VM backend is suitable for tenant-dedicated, privileged, or long-running workflows when the profile pins the VM image, tenant assignment, network, bootstrap identity, resource limits, attestation, and teardown. A pre-existing mutable VM cannot claim untrusted work solely because it is a VM.
External action backends expose fixed typed operations instead of arbitrary commands. Publishers, signers, and deployers validate immutable inputs, approval bindings, target identity, and idempotency keys. They remain the preferred backend for irreversible actions and high-value credentials.
Release Workflow Example
A Light-Fabric release workflow can use one workflow session for build work:
light-fabric
portal-service
controller-rs
light-example-rs
The execution path remains:
light-workflow
- owns workflow, policy snapshot, attempts, approvals, and transitions
-> controller-rs fenced lease
-> light-workflow-runner
- owns backend interaction and artifact transfer
-> workflow-session sandbox
- checks out repositories
- runs tests and approved build commands
- stores workflow-scoped caches and build output
- exports declared artifacts
Recommended task grouping:
prepare workspace workflow-session sandbox
checkout repositories workflow-session sandbox
run unit tests workflow-session sandbox
build release artifacts workflow-session sandbox
generate release notes workflow-session sandbox or bounded host task
publish release per-task fixed publish action
sign artifacts external signing service or per-task fixed action
Do not mount a host Docker socket into tenant-authored build sandboxes. Container image builds should use an approved rootless builder or remote build service, with pinned builder and base-image digests.
Build and package tasks may share state within one workflow policy snapshot. Publish and signing tasks must not execute arbitrary scripts from that mutable workspace. They receive only immutable artifact records from controlled storage, verify trusted-side digests, and use an operator-owned action. Runtime approval binds the exact artifact set, target, version, command template, policy snapshot, and expiry.
Agent-produced changes must not flow directly into a release. After a patch is accepted and merged, release artifacts are rebuilt from the reviewed immutable commit under a fresh build attempt and provenance record.
Agent Workspace Change Policy
An agent-modified workspace is untrusted output even when the agent ran in a
microVM. Every write-capable agent profile references an immutable
workspaceChangePolicyId and digest in the effective policy and lease. The
policy defines allowed and denied path patterns, repository roots, file types,
maximum changed files and bytes, and whether file creation, deletion, rename,
mode changes, submodule changes, or binary files are allowed. Workflow metadata
can request a stricter subset but cannot weaken this policy.
The default agent-repair policy denies changes to privilege-bearing surfaces, including:
.git/** and repository hooks
.github/workflows/** and reusable CI actions
.gitlab-ci.yml, Jenkinsfile, azure-pipelines.yml, and equivalents
CODEOWNERS and repository approval policy
workflow definitions and runner or execution-policy configuration
publish, signing, deployment, and release-credential configuration
Repository-specific policy adds equivalent paths used by that project. An exception requires a distinct operator-approved profile and explicit human review; an agent cannot request the exception itself.
In-sandbox filesystem enforcement is defense in depth. The authoritative check
happens after export in a trusted runner or control-plane component by diffing
against the immutable base commit or tree in a fresh trusted checkout, without
using repository-provided hooks or mutable Git configuration. It canonicalizes
path separators, case and Unicode according to repository rules; detects
symlink, hardlink, rename, mode, submodule, and nested-repository changes; and
creates an immutable canonical patch whose digest is validated against the path
policy. A violation returns
workspace_change_denied, records restricted evidence, and prevents branch,
pull-request, artifact, publish, or signing actions from consuming the patch.
The agent environment never receives repository push credentials. Branch or pull-request creation is a separate fixed action that consumes only the immutable accepted patch, its base commit, path-policy digest, and human-review requirements, never the mutable agent workspace.
Trusted Input And Skill-Package Staging
External inputs are not fetched by sandbox code. Before creating the backend
environment, trusted runner code resolves the lease’s immutable
execution_input_t records, downloads them with runner authority, verifies
kind, size, digest, signature/provenance and scan bindings where required, and
rejects unsafe archive paths, links, devices, ownership, and expansion ratios.
The runner stages only the accepted context, workspace base, trust bundle, and
skill-package bytes and mounts them read-only with nodev, nosuid, and
noexec unless an approved package entrypoint requires execution.
light-agent-worker may revalidate the mounted package manifest, but neither
the worker nor generated code receives artifact-store credentials or outbound
package-download access. Verification or staging failure prevents sandbox
creation. Staged inputs are attempt/session scoped and are removed by the same
idempotent cleanup and watchdog path as the environment.
Secret And Credential Handling
The execution environment must never receive broad platform credentials. Credential access is server-owned and task-scoped.
Required rules:
- Workflow metadata references only logical credential names.
- The effective task policy and command template must both allow the credential.
- Credentials are delivered through opaque, short-lived redemption handles; values are never included in workflow context, leases, logs, or audit.
- Prefer backend-side or egress-proxy credential injection so arbitrary code cannot read the raw value.
- Prefer workload identity and short-lived OIDC tokens over static release tokens.
- Raw credential values are forbidden in environment variables, command-line arguments, process titles, shell history, workflow context, and persistent configuration files. Environment variables may contain only non-secret references such as a credential socket, metadata endpoint, or mounted-file path.
- When the task process must obtain a token, use a workload-authenticated local metadata service or broker bound to the attempt, execution identity, audience, scope, and short expiry. It must be unreachable from other tasks, must not trust an unauthenticated host-wide localhost caller, and must not log or cache returned values.
- Sandboxed model inference uses a runner-owned broker outside the untrusted payload boundary. Provider keys and reusable proxy bearer tokens are not projected into the worker or generated-code environment.
- Prefer a runner-created preconnected descriptor, peer-credential-authenticated Unix-domain socket, vsock, or backend-equivalent local channel. A socket pathname alone is not authority: the broker binds peer and attempt and enforces the approved model, data-boundary and policy digests, token/cost budget, rate, cancellation, and expiry.
- Run the trusted worker/runtime and generated payload under separate
identities and process/mount namespaces. Deny ptrace and cross-process
/procaccess, and prevent broker-descriptor inheritance or reconnection by payload children. A model adapter that requires an extractable provider key is ineligible for an untrusted profile. - A read-only, memory-backed
tmpfsfile with task-unique ownership and mode0400is the fallback for tools that cannot use a broker. Mount it only for the consuming process or task environment and unmount and overwrite metadata on completion. - The
tmpfsraw-value fallback is allowed only in a fresh task sandbox with narrow egress, no untrusted command, no pause or checkpoint, and mandatory termination after the task. - Secret-bearing profiles disable core dumps, restrict cross-process
/procinspection, and exclude credential mounts from snapshots, artifacts, and diagnostic bundles. - Credential revocation occurs on completion, cancellation, lease expiry, policy revocation, or runner quarantine.
- Log redaction is defense in depth, not the primary secret boundary. Encoded or transformed secrets cannot be reliably redacted after exposure.
A workflow-session sandbox must not receive raw publish or signing credentials. Snapshots and auto-pause can preserve process memory and filesystem contents; secret-bearing task sandboxes must use kill-on-timeout and must not be resumed.
Network Policy
Every sandbox profile defines deny-by-default egress. A release profile may allow destinations such as:
github.com
api.github.com
ghcr.io
crates.io
index.crates.io
registry.npmjs.org
approved container registries
The profile must also define schemes, ports, methods, paths, DNS behavior, and
whether apex and wildcard subdomains are allowed. github.com does not imply
*.github.com, and an HTTPS allow rule does not imply SSH access on port 22.
For Cube Sandbox, restricted profiles set allow_internet_access=false, use
explicit L3/L4 allow targets, and add L7 rules for HTTP/HTTPS method and path
control. The adapter verifies the effective provider policy rather than
assuming the requested policy was installed.
Host-executed HTTP, JSON-RPC, and MCP calls keep destination validation, redirect restrictions, response-size limits, and service credential policy. Sandbox placement is not a substitute for SSRF and destination validation.
TLS Inspection Trust Bundles
TLS interception is an explicit profile capability, never an implicit side
effect of network routing. Operator policy selects an immutable
trustBundleRef and digest; workflow metadata cannot supply a CA or disable
certificate verification. The trust bundle contains public CA certificates
only, never interception private keys.
Prefer installing the approved bundle in the immutable template or image. When runtime projection is required, mount it read-only and let the approved command template select the appropriate adapter, for example:
- the OS trust store for native tools,
- an immutable Java truststore selected with JVM truststore options,
NODE_EXTRA_CA_CERTSpointing to the mounted public bundle for Node.js,SSL_CERT_FILEorREQUESTS_CA_BUNDLEpointing to an approved merged bundle for Python and OpenSSL-based tools.
These variables contain non-secret paths, not credential values. The runner verifies the effective trust-store digest from inside the environment before execution and records it in audit and provenance. Rotation creates a new versioned bundle and policy digest. Tasks using certificate pinning, mTLS, or a runtime that cannot honor the bundle must use a separately approved non-intercepting route or fail closed; they must never fall back to disabling TLS verification.
Artifact And Log Boundary
The sandbox filesystem and console are untrusted input. Tasks declare candidate artifact paths, but a glob match alone does not authorize export:
metadata:
lightWorkflow:
artifacts:
- dist/*.tar.gz
- dist/*.sha256
- target/release/light-workflow
Artifact transfer must:
- resolve paths beneath a canonical workspace root,
- refuse symlinks, hardlinks outside the root, devices, sockets, and path traversal,
- avoid time-of-check/time-of-use races,
- enforce per-file, total-byte, and file-count limits,
- compute the authoritative digest after bytes cross the sandbox trust boundary,
- write to immutable, tenant-scoped storage,
- record media type, provenance, template digest, command template, task attempt, policy digest, and retention class,
- scan or validate artifacts when the artifact policy requires it.
Task output contains references, not raw large artifacts:
{
"artifacts": [
{
"artifactId": "019f0000-0000-7000-8000-000000000010",
"name": "light-fabric-0.3.0-x86_64-unknown-linux-gnu.tar.gz",
"sha256": "...",
"size": 12450000,
"storeUri": "artifact://...",
"provenanceRef": "provenance://...",
"provenanceDigest": "sha256:..."
}
]
}
Logs use bounded, ordered chunks with sequence numbers and resumable cursors. The runner applies output limits before transmission. The artifact store keeps full logs only when policy allows it; workflow context keeps summaries and references. Log access, encryption, retention, and deletion remain tenant scoped.
Build Provenance Attestation
For build and release profiles, trusted-side hashing is followed by automatic
provenance generation. The interchange format is an in-toto Statement v1 with
the SLSA Provenance v1 predicate (https://slsa.dev/provenance/v1). The
ExecutionBackend supplies measured execution evidence, but a trusted runner
supervisor or control-plane attestor constructs and authenticates the final
statement; tenant-controlled build steps cannot choose or rewrite its fields.
The statement binds at least:
- each exported artifact name and trusted digest as an attestation subject,
- the command-template build type, template version, resolved argument digest, and policy-approved external parameters,
- source repository URI, immutable commit and tree digest, input artifact digests, base images, and best-effort resolved dependencies,
- workflow, task, attempt, lease, and policy snapshot identities,
- runner and builder identity and version,
- backend kind, implementation, capability record, template or image digest, runtime class where applicable, and execution session isolation,
- resource, network, workspace-change, credential-delivery, and trust-bundle policy digests,
- start and completion timestamps, outcome, and completeness metadata.
The attestation is stored immutably beside the artifacts and referenced by URI and digest. When signed provenance is required, signing occurs outside the tenant execution environment with a short-lived workload identity or key that tenant code cannot access. Publish and signing actions verify the attestation signature, subject digests, builder identity, policy expectations, and approval bindings before consuming an artifact.
Using the SLSA format does not by itself establish a SLSA Build level. A profile
may declare provenanceMode: unsigned for development or signed for release,
but the platform must be assessed against the applicable SLSA build-platform
and isolation requirements before making a level claim. Host-integrated builds
must not inherit a hosted or isolated level merely because they emitted the
same JSON shape. Provenance records claims and evidence about the build; it does
not by itself prove artifact correctness or that the execution environment was
uncompromised.
Audit
Every sandboxed task attempt records:
- tenant, host, trigger principal, and correlation ID,
- workflow definition ID and version,
- workflow process and instance ID,
- task ID, task name, and attempt number,
- runner ID, lease ID, and fencing token,
- requested profile and immutable effective policy digest,
- backend ID, kind, implementation, version, capability digest, and request ID,
- execution session ID, immutable template or image digest, and backend operation ID,
- command template ID, resolved argv digest, working directory, immutable base commit or tree, workspace-change policy and accepted patch digests, and environment name allowlist,
- credential names and delivery mechanism, never values,
- requested and effective network, trust-bundle, resource, and local-cleanup policy digests,
- approval IDs and the artifact and target digests they authorize,
- artifact metadata, provenance statement and envelope digests, builder and signer identities, and verification result,
- exit status, duration, resource usage, output sizes, and log reference,
- cancellation, disconnect, watchdog action, retry, reconciliation, and cleanup events.
For call: agent, also record model provider scope, model name, prompt profile,
token budget, output schema ID, validation result, tool policy, and data
boundary, plus the canonical changed-file manifest and path-policy result.
Audit records are append-only. Workflow-visible output must not contain backend credentials, control-plane tokens, raw injected secrets, or restricted backend diagnostics.
Failure And Retry Handling
Backend and runner failures map to stable workflow errors:
- hard admission, quota, or cost-policy failure:
execution_admission_denied, - temporary capacity shortage:
execution_capacity_deferred, - capacity queue deadline expired:
execution_queue_timeout, - policy rejection:
execution_policy_denied, - unsupported capability:
execution_capability_missing, - environment startup failure:
execution_start_failed, - task wall-clock timeout:
execution_timeout, - command non-zero exit:
command_failed, - agent patch violates the workspace policy:
workspace_change_denied, - required provenance cannot be generated or authenticated:
provenance_generation_failed, - oversized result, log, or artifact:
execution_output_too_large, - lease expiry:
runner_lease_expired, - cancellation:
execution_cancelled, - backend outcome not yet known:
execution_outcome_unknown, - cleanup not confirmed:
execution_cleanup_pending.
Retry classification is explicit:
- Policy, validation, and unsupported-capability failures are not retried.
execution_capacity_deferredstays in the fair scheduling queue and does not consume a command retry attempt.- Idempotent startup may retry using the same session idempotency key.
- A non-zero command exit follows the workflow retry policy only when the command template declares retry safety.
- Transport loss after dispatch enters
UNKNOWNand is reconciled before a new attempt. - Timeout requests cancellation, fences the attempt, and inspects the backend before retry evaluation.
- External side effects require a target-system idempotency or reconciliation contract. Otherwise, an unknown result requires operator intervention.
The workflow process does not transition until it has accepted a result from the current fenced attempt and committed the result, audit, artifact records, and next-task creation atomically in the workflow database.
Cleanup And Retention
Sandbox termination and evidence retention are separate responsibilities.
- Workflow execution sessions are cleaned after workflow completion, permanent failure, cancellation, maximum lifetime, or policy revocation.
- Per-task backend environments are cleaned after the task result and required evidence are secured.
- Backend snapshots and checkpoints are independent resources and must be explicitly deleted according to policy.
- Failed destroy or delete calls create cleanup-pending records for the orphan reconciler.
- Backend metadata tags include tenant, workflow instance, policy snapshot, session record, and expiry so orphan discovery does not depend only on the workflow database.
- The runner watchdog cleans expired local journal entries and tagged backend resources when control-plane contact is unavailable. Native backend expiry remains required where possible in case the runner host also fails.
- Retained snapshots, logs, and artifacts require encryption, tenant isolation, retention limits, deletion audit, and data-residency policy.
- Secret-bearing sandboxes cannot be paused, checkpointed, or retained for debugging.
Definition And Runtime Approval
Definition approval and task approval solve different problems.
Definition approval covers profiles that enable command execution, external MCP processes, broad network access, mutable mounts, credential access, or high-cost resource classes. The approval produces an immutable workflow definition and profile version.
Runtime approval covers a specific irreversible action. A publish or signing approval includes:
tenant and workflow instance
task and attempt
command template
artifact set and trusted digest
release target and version
effective policy digest
approver and approval time
expiry and single-use nonce
Changing any bound value invalidates the approval. A new attempt after an unknown outcome requires reconciliation and may require a new approval; it must not silently reuse the prior approval.
Approval Is An Orchestration State
Waiting for a person is owned by the authenticated origin service—
light-workflow for workflow tasks or light-agent for standalone agent
actions—never by a runner task:
build or stage task completes
-> artifacts and provenance commit
-> action lease ends and task credentials/model channel are revoked
-> task sandbox cleans, or eligible session workspace enters bounded hold
-> origin persists WAITING_APPROVAL
-> approval is granted and bindings are revalidated
-> origin creates a new numbered domain and common execution attempt
-> controller-rs issues a fresh lease, fencing token, and scoped grants
When approval is known before dispatch, the origin records the bound action
intent but creates no common execution attempt until approval. If a running
agent/runtime discovers the boundary, it returns a known
approval_required terminal result; its lease and grants end and its sandbox
is cleaned or explicitly checkpointed under bounded non-secret policy. Approval
never reactivates that attempt.
The origin transaction that enters WAITING_APPROVAL persists exactly one
session disposition—cleanup or a policy-valid bounded hold. If common session
state is later moved to another database, use an idempotent transactional
outbox. A session reaper must not infer abandonment from the ended action lease
before that disposition is durable.
The runner must not poll for approval, hold or renew a task lease, or keep a secret-bearing environment alive while the workflow waits. A non-secret workflow or agent session may be paused/checkpointed or retained only through the separate bounded hold contract when explicit cost, maximum-lifetime, and retention policy allows it. Important uncommitted work should also be exported as an approved immutable patch/checkpoint so correctness does not depend only on the live sandbox. The preferred release path exports immutable artifacts and provenance, then cleans the build environment.
Approval rejection or expiry transitions the orchestration state without dispatching a runner task. Approval grants authorize only the new fixed-action attempt and are rechecked against its operation, arguments, artifact, provenance, target, policy, command template, expiry, and nonce. The approval is consumed once by the new common attempt, which has a monotonic fencing token. The prior attempt, lease, backend handle, and grants remain immutable and are never returned to the build or agent environment. A held workspace may be reused by the fresh action only after principal/base/runtime/policy/expiry and cleanup-state revalidation; otherwise restore only a verified policy-permitted checkpoint/patch into a fresh environment.
Implementation Plan
Phase 1: Contracts And Persistence
- Define versioned security metadata and strict validation.
- Define field-specific policy resolution and immutable policy snapshots.
- Add dedicated policy, execution session, session-cleanup request, immutable input, common attempt, artifact, approval, and audit persistence.
- Define workspace-change, trust-bundle, credential-projection, local-cleanup, capacity-queue, and provenance policy contracts.
- Persist tenant, trigger principal, correlation ID, and policy snapshot on workflow start.
- Define the identifiers-only transactional result-ready PostgreSQL wakeup plus indexed startup/periodic origin catch-up.
- Keep unsupported
run.*task types disabled.
Phase 2: Runner Lease And Fencing
- Align with the
light-workflow-runnerregistration and lease protocol. - Add attempt numbers, lease renewal, fencing tokens, compare-and-set result acceptance, cancellation, and reconciliation.
- Add normalized result and resumable log contracts.
- Store terminal common results and emit origin wakeups in one transaction; prove correctness when notifications are missed, duplicated, or reordered.
- Add the durable runner cleanup journal, autonomous watchdog, startup scan, and backend-native expiry contract.
- Add bounded per-tenant fair capacity queues, atomic slot reservation, and jittered claim backoff.
- Prove that stale or duplicate reports cannot transition a workflow or agent domain object.
Phase 3: Minimal Per-Task MicroVM Backend
- Define the
ExecutionBackendtrait, capability vocabulary, compatibility records, and boundary-appropriate conformance tests. - Implement Cube Sandbox capability discovery and production-baseline checks.
- Support one approved
run.shellcommand template in a per-task sandbox. - Start with no credentials, no external side effects, deny-all egress, and hard resource limits.
- Add cancellation, orphan cleanup, and backend operation reconciliation.
Phase 4: Additional Backends, Artifacts, And Sessions
- Add safe artifact extraction, trusted-side hashing, immutable storage, and trusted-side in-toto/SLSA provenance generation.
- Add Docker Sandboxes for autonomous developer-local agents, requiring clone workspace mode and an approved policy.
- Add rootless OCI container execution for trusted build and packaging profiles.
- Allow trusted runners hosted in Toolbx to register only host-integrated, no-sandbox capabilities; no per-task Toolbx adapter is required initially.
- Add Kubernetes Jobs only with an approved runtime-class compatibility record.
- Add workflow-session reuse with single-writer enforcement where supported.
- Add bounded agent-session reuse with effective minimum expiry and durable origin close/revoke/expiry cleanup requests.
- Add session state/version/fencing plus idempotent hold/pause/checkpoint,
resume, and cleanup operations.
IDLE_APPROVAL_HOLDis separate from an action lease, bounded by effective expiry and cost policy, and never carries model or credential authority. - Add trusted runner-side download, verification, safe extraction, staging, read-only mounting, and cleanup for immutable skill packages and inputs.
- Add copy-on-write isolation for parallel branches and agent repair tasks.
- Add checkpoint lifecycle and deletion without allowing secret-bearing snapshots.
- Add server-owned agent workspace-change policies and trusted post-export diff validation before branch or pull-request creation.
Phase 5: Network And Credential Profiles
- Add deny-by-default L3/L4 and L7 policy generation and verification.
- Add immutable TLS trust-bundle profiles and language-runtime adapters.
- Add brokered credential handles, authenticated local metadata endpoints,
read-only
tmpfsfallback, and backend-side credential injection. - Add protected runner-owned model-broker transports, separate worker/payload identities, descriptor isolation, and broker-side model and budget policy.
- Add short-lived workload identity and revocation.
- Add runtime approval records bound to immutable inputs.
Phase 6: Release, Publish, And Signing
- Add release build/test/package workflows using workflow sessions.
- Export immutable artifact sets with provenance.
- Add fixed publish actions or a dedicated release service.
- Add external signing or a fixed per-task signing action.
- Require digest-bound human approval and complete unknown-outcome reconciliation before retry.
- Persist approval waiting only in the origin service; end any current action lease before waiting, use only the separate bounded session-hold contract where allowed, and issue a new numbered domain/common fixed-action attempt, lease, fencing token, and grants after approval.
Acceptance And Failure-Injection Tests
The feature is not ready for tenant workloads until tests prove:
- a command lasting beyond the old task-lock interval is not claimed twice,
- an expired runner cannot report success after a new attempt starts,
- a crash after backend dispatch but before database commit is reconciled,
- a terminal result committed while its origin listener is offline is found by indexed catch-up and accepted once; duplicate or reordered wakeups do not duplicate the domain transition,
- repeated create and execute requests do not create duplicate operations,
- a disconnected runner stops new work, locally fences execution at lease expiry, and cleans the environment without control-plane reachability,
- runner restart replays the cleanup journal, and backend-native expiry cleans resources when the runner host does not restart,
- temporary saturation remains queued with bounded jittered backoff and does not create attempts, sandboxes, or a claim storm,
- a
host-integratedorshared-kernel-containerbackend cannot claim a task requiringmicrovm, - a Toolbx runner cannot claim tenant-authored, isolated, or credential-bearing tasks,
- Docker Sandboxes direct workspace mode is rejected for untrusted tasks and clone mode preserves the host repository boundary,
- no tenant-authored backend receives a host container-engine socket,
- a Kubernetes Job cannot claim a stronger boundary than its approved runtime class and node policy provide,
- cancellation fences results and eventually removes the sandbox,
- origin session close, revocation, and expiry fence active work and reclaim a reused sandbox without waiting for backend-native TTL, even across controller or runner restart,
- policy revocation stops new attempts and revokes credentials,
- denied network destinations remain denied for HTTP, non-HTTP, DNS, and direct IP access,
- backend control-plane requests require authorization,
- output, log, artifact, process, disk, and time limits are enforced,
- symlink and path-traversal artifact exports fail,
- agent changes to CI/CD, workflow, approval, release, signing, deployment, or repository-policy paths are rejected before PR or branch creation,
- case, Unicode, symlink, rename, mode, submodule, and nested-repository tricks cannot bypass the workspace-change policy,
- raw credentials do not appear in process memory snapshots, logs, task output, artifacts, or workflow context for the preferred credential path,
- raw tokens are never placed in environment variables or argv, and metadata
service and
tmpfsprojections are isolated to the leased task, - generated payload code cannot inspect or inherit the worker’s model-broker channel, obtain a provider/proxy bearer, impersonate another attempt, choose an unauthorized model, or exceed broker-enforced budget,
- skill-package digest/signature mismatch, unsafe archive content, or staging failure prevents sandbox creation, and sandbox code cannot access the artifact store,
- the effective TLS trust-bundle digest is verified without allowing workflow code to add a CA or disable certificate validation,
- snapshots and failed sandbox creations are found and cleaned by the orphan reconciler,
- required provenance binds the trusted artifact digest, source and input digests, command template, builder and backend identity, and policy digest; tampering or an unexpected signer blocks publish,
- no action lease, model channel, action credential, or secret-bearing task environment remains active while a workflow or agent waits for approval; an eligible non-secret session survives only through a distinct bounded hold/checkpoint, and approval creates a fresh domain/common fixed-action attempt, lease, and fencing token without reopening the prior attempt,
- releasing an action lease does not prematurely clean a valid approval-held session, while hold/session expiry and close/revocation always trigger cleanup; resume cannot extend the fixed maximum or restore unverified state,
- publish approval fails if the artifact digest, target, policy, command, or expiry changes.
Open Decisions
- Whether approved execution profiles live only in service configuration or are also portal-managed immutable records.
- Whether artifact metadata lives in portal tables while bytes live in object storage, or whether another artifact service owns both.
- Whether all publish and signing operations use a separate release service or whether a limited set of fixed runner actions is supported.
- Which Cube and Docker Sandboxes versions and capabilities form the first supported compatibility sets.
- Which hardened Kubernetes runtimes and dedicated-VM attestation mechanisms should be approved for untrusted workloads.
- Whether Toolbx support should remain an operational runner profile or later gain explicit local-runner lifecycle helpers.
- Which watchdog deadlines, backend-native TTL mechanisms, and cleanup evidence are required for each backend compatibility record.
- Which protected runner-local broker transport is supported first for each backend: preconnected descriptor, peer-checked Unix-domain socket, vsock, or backend-native equivalent.
- Which repository paths are protected by the default agent policy and how repository-specific additions are approved.
- Which provenance signer, transparency or timestamp mechanism, storage convention, and target SLSA Build level release profiles require.
- How emergency policy revocation balances immediate termination against preserving forensic evidence.
- Whether a first-class
securityfield should be added toworkflow-coreafter the metadata-based contract proves stable.
References
- Light-Workflow Runner Design
- Cube Sandbox Introduction
- Cube Sandbox Lifecycle
- Cube Snapshot, Rollback, And Clone
- Cube Egress Network Policy
- Cube Security Proxy
- Cube Restrict Public Access
- Cube Authentication
- Cube Network Hardening
- Docker Sandboxes
- Docker Sandboxes Architecture
- Docker Sandboxes Security Model
- Docker Sandboxes Default Security Posture
- Docker Sandboxes Credentials
- Fedora Silverblue Toolbx
- Toolbx
- SLSA v1.2 Build Provenance
- SLSA v1.2 Build Requirements
- in-toto Attestation Framework
Insurance Claim Agentic Workflow
This page describes a product workflow demo for orchestrating multiple agents,
skills, APIs, and human tasks with light-workflow.
The scenario is an auto insurance claim from first notice of loss to a settlement recommendation. It is a useful demo because it is familiar, has clear business states, needs several API calls, and includes human decisions that should not be delegated fully to an agent.
Demo Goal
The workflow should show how a deterministic process can coordinate:
- two or three agents
- multiple skills per agent
- REST API calls
- MCP tool calls through
light-gateway - human input and approval tasks
- branching based on policy, risk, and claim severity
The same business flow should be executable in two variants:
- REST workflow: calls the demo APIs directly with HTTP/OpenAPI tasks.
- MCP workflow: calls the same capabilities through MCP tools exposed by
light-gateway.
The workflow owns the process. Agents work inside bounded tasks and should not invent new process paths outside the workflow definition.
For the agent execution boundary, see
Native Agent Call. In the current implementation,
call: agent is a native light-workflow task. It does not invoke a
containerized light-agent service. API access in this demo is owned by the
workflow through direct HTTP tasks or MCP tool calls routed through
light-gateway.
Execution Model
This demo uses the enterprise workflow-first model:
light-workflowowns the claim process, task state, retries, branching, human tasks, and audit trail.- API access is explicit in the workflow as
call: httporcall: mcp. - Native
call: agenttasks perform bounded reasoning over workflow-owned context and must return structured output. - Skills provide instructions, tool context, and workflow mappings, but they do not give an agent permission to invent unreviewed process paths.
- Containerized
light-agentservices are not invoked by this demo workflow. They remain the runtime for chat clients and future service-agent integration.
Demo APIs
The existing demo APIs can be used as stand-ins for insurance services.
| API | Role in the claim workflow |
|---|---|
demo-customer-profile-api | Policyholder profile, vehicle list, policy status, contact preference, prior claims. |
demo-offer-decision-api | Claim triage, risk decision, settlement or repair recommendation. |
If more realism is needed later, the same workflow can add simulated services for document storage, repair estimates, fraud review, or payment authorization.
Agents
Claim Intake Agent
The Claim Intake Agent owns first notice of loss collection and basic validation.
Skills:
- collect accident facts
- validate required claim fields
- look up customer, policy, and vehicle data
- identify missing information
- summarize the claim for the next agent
Typical tools or API calls:
- get customer profile
- get customer policies
- get covered vehicles
- get prior claims
Human tasks:
- claimant confirms accident details
- claimant answers missing information questions
- claimant uploads or confirms photos, police report, and tow status
Coverage And Liability Agent
The Coverage and Liability Agent checks whether the claim can continue and whether a human adjuster must review it.
Skills:
- coverage eligibility check
- incident date versus policy period check
- vehicle coverage check
- liability and severity classification
- fraud or special investigation flagging
Typical tools or API calls:
- get policy status
- get prior claim history
- run triage decision
- run risk decision
Human tasks:
- adjuster reviews unclear liability
- adjuster confirms coverage exception handling
- special investigation team reviews high-risk claims
Settlement Agent
The Settlement Agent prepares the next action and customer-facing explanation.
Skills:
- repair versus total-loss recommendation
- deductible explanation
- settlement recommendation
- customer message draft
- next-document request
Typical tools or API calls:
- get offer decision
- get customer contact preference
- create settlement recommendation
Human tasks:
- adjuster approves high-value payment
- claimant accepts repair or settlement path
- claimant requests callback or more review
Claim Context And Handoffs
The workflow engine owns the claim state. Agents should be treated as stateless workers that read the current claim context, perform a bounded task, and return structured output.
Each major step enriches a shared claim context:
- intake adds normalized accident facts and missing information status
- customer lookup adds profile, policy, vehicle, and prior-claim data
- coverage review adds eligibility, deductible, liability, and risk signals
- triage adds severity, recommended path, and human-review requirements
- settlement adds the recommendation, explanation, and next actions
Handoffs between agents should happen through this workflow-owned context, not through private agent memory. This keeps the process deterministic, replayable, and auditable.
Workflow Outline
1. Start Claim
Input:
{
"customerId": "CUST-001",
"vehicleId": "VEH-001",
"incidentDate": "2026-05-30",
"accidentDescription": "Rear-ended at an intersection.",
"location": "Ottawa, ON",
"injuryReported": false,
"vehicleDrivable": false
}
The workflow validates that customerId, vehicleId, incidentDate, and
accidentDescription are present.
2. Fetch Customer Context
The workflow calls the profile and policy capabilities to retrieve:
- customer identity
- policy list
- covered vehicles
- contact preference
- prior claim count
Assertions:
- customer exists
- vehicle belongs to customer
- at least one active policy exists
3. Ask For Missing Information
If the input is incomplete, the workflow creates a human task for the claimant.
Example questions:
- Was anyone injured?
- Was another vehicle involved?
- Is the vehicle drivable?
- Was a police report filed?
- Are photos available?
The workflow should be resumable after the claimant answers.
4. Coverage Check
The workflow passes the gathered claim context to a native Coverage and Liability agent task. That task checks:
- policy active on incident date
- covered vehicle
- applicable coverage type
- deductible
- excluded conditions
Branches:
- no matching policy: route to adjuster review
- policy inactive: prepare denial draft for human review
- coverage found: continue to triage
5. Triage Decision
The workflow calls the decision API, either directly with HTTP or through
light-gateway MCP, with normalized claim context.
Expected decision output:
{
"severity": "medium",
"riskLevel": "low",
"recommendedPath": "repair",
"requiresAdjusterReview": false,
"estimatedLoss": 3200
}
Branches:
- low risk and low value: continue automatically
- unclear liability: create adjuster review task
- high risk: create special investigation task
- high value: create approval task
6. Settlement Recommendation
The workflow passes the approved claim context to a native Settlement agent task. That task prepares:
- recommended path: repair, estimate, total-loss review, denial draft, or more information
- deductible explanation
- next documents required
- customer-facing summary
The result should be structured so the UI can render it and the agent can explain it.
7. Human Approval
Approval is required for:
- high estimated loss
- denial recommendation
- special investigation referral
- liability uncertainty
- customer dispute
The task should record:
- approver role
- approval decision
- comment
- timestamp
- whether the workflow should proceed, revise, or stop
8. Customer Response
The claimant chooses one of:
- accept repair path
- request adjuster callback
- upload more documents
- dispute the recommendation
This should be modeled as a human ask task rather than an agent-only step.
9. End State
Possible workflow outcomes:
| State | Meaning |
|---|---|
claim-approved | Claim can proceed to repair or settlement. |
needs-adjuster-review | Human adjuster must review before next action. |
needs-customer-info | Claimant must provide missing information. |
referred-to-siu | Claim is referred to special investigation. |
claim-denied-draft | Denial is drafted but still needs human approval. |
Failure Handling And Fallbacks
The demo should show graceful degradation when an API call or agent task cannot finish automatically.
Recommended fallback behavior:
| Failure | Workflow response |
|---|---|
Customer profile returns 404 | Create a manual customer verification task. |
| Policy or vehicle lookup is unavailable | Retry, then route to adjuster review with the partial claim context. |
| Decision API is unavailable | Create a manual triage task and include the last successful context. |
| Agent output fails validation | Re-run once with validation feedback, then create a human review task. |
| Human task times out | Escalate to the configured role or mark the claim as waiting for follow-up. |
The failure branch should preserve the accumulated claim context and the failed request or response metadata so the human reviewer can continue from the same state instead of restarting the claim.
REST Variant
The REST workflow calls the demo APIs directly.
Use this variant to show:
- deterministic API orchestration
- direct HTTP/OpenAPI task execution
- workflow assertions
- human waiting tasks
- repeatable headless tests with fixed inputs
Example task sequence:
start-claim
get-customer-profile
assert-active-policy
ask-missing-info
run-claim-triage
switch-risk-path
ask-adjuster-approval
prepare-settlement-summary
ask-customer-response
complete-claim
MCP Variant
The MCP workflow invokes the same capabilities through MCP tools exposed by
light-gateway.
Use this variant to show:
- tool discovery with
tools/list - tool execution with
tools/call - agent skill guidance over the selected tool set
- gateway as the runtime MCP data plane
Skills should be treated as guidance and curation for the agent, not as the
runtime transport. The workflow still calls MCP tools through light-gateway.
A skill describes when and how to use tools. For example, the
coverage-review skill can instruct the agent to call evaluate_coverage
before score_claim_risk, explain which fields must be present, and define
what output shape the workflow expects.
Example tool groups:
| Skill | Tools |
|---|---|
claim-intake | get_customer_profile, get_policy, get_vehicle, list_prior_claims |
coverage-review | evaluate_coverage, score_claim_risk, classify_liability |
settlement | recommend_offer, generate_customer_summary, list_required_documents |
Human Task Model
Human work should be explicit and durable.
Recommended task types:
- claimant information request
- adjuster approval
- liability review
- special investigation review
- customer settlement response
Recommended fields:
- prompt
- mode: choice, text, object, file, approval
- assignee or candidate role
- due time
- validation rules
- sensitive flag
- comments
- decision result
The workflow should pause at the human task and resume after a valid response is recorded.
The pause is durable. light-workflow persists the process and task state while
waiting, so the workflow can remain idle for hours or days without consuming
active execution resources. When the claimant, adjuster, or investigator
completes the task, the workflow resumes from the persisted state and continues
with the same claim context.
Minimal First Implementation
Start with a narrow happy path:
- Start with
customerId,vehicleId, and accident details. - Workflow fetches customer profile through HTTP or MCP.
- Workflow asserts active policy and covered vehicle.
- Workflow calls the decision API for triage.
- Workflow asks an adjuster to approve if
estimatedLossexceeds a threshold. - Native Settlement agent task prepares the recommendation.
- Workflow completes with
claim-approvedorneeds-adjuster-review.
This first version is enough to demonstrate multi-agent orchestration without needing every insurance edge case.
Later Enhancements
Add complexity incrementally:
- document upload and OCR simulation
- repair shop estimate comparison
- fraud and special investigation path
- payment authorization
- subrogation when another driver is liable
- scheduled headless regression runs
- customer notification drafting
- analytics for cycle time and approval bottlenecks
Demo Success Criteria
The demo is successful if it shows:
- the same business process running through REST and MCP variants
- agents using skills to perform bounded work
- APIs called through both direct HTTP and MCP tool paths
- at least one human input task
- at least one human approval task
- auditable workflow state transitions
- clear final outcome and explanation
Light Portal Setup
This page describes the portal-side setup required to run the
light-workflow product demos from a local light-portal stack.
For the execution model behind native agent tasks, see Native Agent Call. For the insurance product scenario, see Insurance Claim Agentic Workflow.
Prerequisites
Start the local portal stack with the workflow services, gateway, controller, and Postgres available.
For the Rust local stack:
cd /home/steve/workspace/portal-config-loc
./scripts/deploy-local.sh pg rust
The local stack should include:
- Postgres,
workflow-command,workflow-query,light-gateway,- controller,
- config-server,
demo-customer-profile-api,demo-offer-decision-api.
light-workflow must use the same database as workflow-command:
DATABASE_URL=postgres://postgres:secret@localhost:5432/configserver
Start Light-Workflow
Build and run light-workflow from the light-fabric checkout:
cd /home/steve/workspace/light-fabric
cargo build -p light-workflow --locked
cd apps/light-workflow
DATABASE_URL=postgres://postgres:secret@localhost:5432/configserver \
LIGHT_PORTAL_AUTHORIZATION="Bearer <workflow-service-token>" \
SERVER_ENVIRONMENT=dev \
LIGHT_WORKFLOW_CONFIG_MODE=local \
RUST_LOG=light_workflow=debug,info \
WORKFLOW_LOG_ANSI=false \
./run.sh --debug-binary
For repeated runs, put those values in
apps/light-workflow/light-workflow.env and run:
./run.sh --debug-binary
Import Agent Catalog Data
Native call: agent tasks load portal agent, skill, and tool metadata from the
portal database. Import the demo catalog events before running workflows that
contain agent tasks.
cd /home/steve/workspace/event-importer
./importer.sh \
--filename /home/steve/workspace/light-fabric/apps/light-workflow/examples/agent-catalog-events.json
For a different host or user, pass replacement rules:
./importer.sh \
--filename /home/steve/workspace/light-fabric/apps/light-workflow/examples/agent-catalog-events.json \
--replacement '[
{"field":"hostId","from":"01964b05-552a-7c4b-9184-6857e7f3dc5f","to":"<host-id>"},
{"field":"user","from":"01964b05-5532-7c79-8cde-191dcbd421b8","to":"<user-id>"},
{"field":"operationOwner","from":"01964b05-5532-7c79-8cde-191dcbd421b8","to":"<user-id>"},
{"field":"deliveryOwner","from":"01964b05-5532-7c79-8cde-191dcbd421b8","to":"<user-id>"}
]'
The demo catalog uses modelProvider: mock for deterministic local runs. For
real model execution, update the portal agent definitions to use the desired
provider and apiKeyRef.
Upload API Metadata
For the insurance claim demos, upload or refresh the OpenAPI specs for:
demo-customer-profile-api,demo-offer-decision-api.
The portal catalog should contain endpoint and tool projections for the demo
APIs before the MCP workflow is run. The MCP workflow expects light-gateway
tools/list to expose these tools:
getCustomerProfile
getCustomerPreferences
getCustomerPolicies
getCoveredVehicle
listPriorClaims
triageClaim
recommendSettlement
Verify the tool surface through the gateway:
curl -k -sS -X POST "https://localhost:8443/mcp" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <access-token>" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
Create Workflow Definitions
Create workflow definitions in the portal UI or through the workflow command API. For the insurance claim demo, create these definitions:
insurance-claim-rest-v1.yaml
insurance-claim-mcp-v1.yaml
insurance-claim-headless-v1.yaml
The files live in:
/home/steve/workspace/light-fabric/apps/light-workflow/examples
After creation, capture their ids:
psql "postgresql://postgres:secret@localhost:5432/configserver" \
-c "select host_id, wf_def_id, name from wf_definition_t where active and name in ('insurance-claim-rest-v1', 'insurance-claim-mcp-v1', 'insurance-claim-headless-v1') order by name;"
Roles And Human Tasks
The insurance claim workflow creates durable human tasks. Confirm that the demo host has the roles used by those assignments:
claimant
claims-adjuster
siu-investigator
customer-service
Human tasks remain in the portal database while waiting. The workflow resumes after the task-completion command records a valid response.
Start And Verify
Use the portal UI start action, Postman collection, or curl helper from the examples directory.
cd /home/steve/workspace/light-fabric/apps/light-workflow/examples
ACCESS_TOKEN=<token> \
HOST_ID=<host-id> \
HEADLESS_WF_DEF_ID=<headless-wf-def-id> \
./insurance-claim-demo-curl.sh start-headless
Run the SQL verification helper after each start or task completion:
psql "postgresql://postgres:secret@localhost:5432/configserver" \
-v host_id=<host-id> \
-f /home/steve/workspace/light-fabric/apps/light-workflow/examples/insurance-claim-demo-queries.sql
For the full runbook, see:
/home/steve/workspace/light-fabric/apps/light-workflow/examples/README.md
Troubleshooting
| Symptom | Check |
|---|---|
| Workflow starts but no process appears | Confirm light-workflow uses the same DATABASE_URL as workflow-command. |
| Agent task fails before a human task | Confirm agent-catalog-events.json was imported for the same hostId. |
| MCP tool is not found | Call gateway tools/list and confirm the tool names match the workflow YAML. |
| Human task is not visible | Check task_asst_t, role membership, and task status. |
Input fields resolve as ${ .customerId } | Confirm startWorkflow sends input as a JSON object, not a JSON string. |
Comparison: Light-Fabric vs. AgentGateway
This document provides a high-level comparison between Light-Fabric and AgentGateway to help architects and engineering leaders choose the right foundation for their agentic workflows.
Overview
While both systems aim to facilitate interaction with Large Language Models (LLMs), they operate at different layers of the AI stack and prioritize different architectural outcomes.
| Feature | Light-Fabric | AgentGateway |
|---|---|---|
| Primary Philosophy | Agentic Fabric: Unified Governance & Lifecycle | Agentic Gateway: High-performance Proxy |
| Core Architecture | Integrated Platform (Layer) | Standalone Gateway (Service) |
| Target User | Central IT / Platform Engineering | Application Developers / DevOps |
| Lifecycle Management | APIs, Agents, MCPs, and Gateways | Primarily LLM Request Routing |
| Language | Native Rust (Extreme Performance) | Rust / Go (Variable) |
1. Governance vs. Connectivity
Light-Fabric (Governance)
Light-Fabric is designed as a Single Control Plane. It assumes that in an enterprise environment, “freedom without governance is chaos.” It provides:
- Centralized Registry: Every agent, skill, and tool is registered and governed via the
light-portal. - Fine-Grained Authorization: Deep policy enforcement at the endpoint level, including row and column-level data masking.
- Auditability: A unified audit trail for all agentic interactions across the entire organization.
AgentGateway (Connectivity)
AgentGateway typically focuses on the North-South traffic between an application and multiple LLM providers. Its primary strength is:
- Simplified Routing: Getting a request from Point A to Point B with retries and failover.
- Provider Abstraction: Normalizing different LLM APIs into a single interface.
2. Integrated Intelligence: Hindsight
One of the defining differences of the Light-Fabric is the deep integration of Hindsight Memory.
- Light-Fabric: Memory is not an “add-on.” The platform provides native biomimetic memory banks (World Facts, Experiences, Mental Models) that are automatically managed and scoped (Global, Shared, Private) as part of the fabric.
- AgentGateway: Typically treats memory as external state. The application or a separate vector database must manage context before sending the request through the gateway.
3. Skill & Tool Management
Centralized Skills (Fabric)
In Light-Fabric, skills (tools) are first-class citizens. They are registered, versioned, and governed centrally. An agent doesn’t just “have” a tool; the Fabric grants the agent access to a skill based on its role and the current context.
Standard Tooling (Gateway)
AgentGateway generally passes tool definitions through to the provider. The management of who can use which tool and how those tools are secured is usually left to the application logic.
4. Orchestration: Hybrid Agentic Workflows
Light-Fabric (Integrated Orchestrator)
Light-Fabric treats orchestration as a foundational service. It implements a Hybrid Model:
- Deterministic Process: The overall business logic (e.g., insurance claim steps) is fixed and compliant.
- Autonomous Tasks: Individual steps within the process are delegated to agents.
- Statefulness: The Fabric manages long-running state across days or weeks, ensuring durability.
AgentGateway (Stateless Proxy)
AgentGateway is primarily a stateless component.
- External Orchestration: The workflow logic must reside in your application code or an external engine (like Temporal).
- Proxy Only: It handles the communication but does not “understand” or manage the multi-step business process itself.
5. Security: The Rule Engine
Light-Fabric (Integrated Governance)
Light-Fabric includes an integrated YAML-based Rule Engine (light-rule) designed for fine-grained authorization:
- Data Filtering: Automatically masks or filters response data (column/row level) based on policies.
- Policy Enforcement: Checks permissions before an agent executes a tool or accesses a memory unit.
- Hot-Reloading: Security rules can be updated in real-time without redeploying the platform.
AgentGateway (Basic Middleware)
AgentGateway typically provides basic security features like API key validation or rate limiting.
- Limited Filtering: While it can intercept traffic, implementing complex, context-aware data masking usually requires writing custom middleware or handling it at the application level.
6. MCP Support: Gateway vs. Ecosystem
Light-Fabric (Integrated Tooling)
Light-Fabric treats Model Context Protocol (MCP) as a primary source for agent tools.
- Direct Integration: Agents use the
mcp-clientto directly consume tools from MCP servers. - Registry Management: MCP servers are registered in the
light-portal, allowing for centralized discovery and governance. - Unified Security: The same Fine-Grained Authorization rules apply to MCP tools as they do to native Rust tools.
AgentGateway (Specialized MCP Proxy)
AgentGateway provides a highly specialized MCP Gateway layer.
- Protocol Translation: It excels at translating between different MCP transports (SSE, Streamable HTTP, etc.).
- Exposing Servers: Its primary role is to make MCP servers accessible to external applications through a normalized gateway interface.
- Advanced Networking: Includes features like stream merging and specialized MCP routing.
For a deep dive into the technical differences, see our Detailed MCP Feature Comparison.
Summary: Which to Choose?
Choose Light-Fabric if:
- You are building an Enterprise AI Strategy that requires unified governance, stateful workflows, and integrated security.
- You need to manage the entire lifecycle of agents and the business processes they participate in.
- You require advanced data privacy (masking) and long-term memory (Hindsight) as native platform features.
Choose AgentGateway if:
- You need a lightweight proxy to handle LLM provider failover and basic request normalization.
- You prefer to manage agent logic, workflows, memory, and security entirely within your external application stack.
- You are looking for a simple tool to solve immediate connectivity needs without implementing a comprehensive platform layer.
Detailed Comparison: MCP Gateway Features
This document provides a technical deep dive into the Model Context Protocol (MCP) implementations in Light-Fabric and AgentGateway.
Feature Matrix
| Feature | Light-Fabric | AgentGateway |
|---|---|---|
| Primary Role | Provider/Gateway/Portal: Exposes MCP/API Servers. | Provider/Gateway: Exposes MCP servers. |
| Onboarding | Auto-Discovery: Automatic tools/list sync. | Manual: K8s CRD/Manifest configuration. |
| Data Privacy | Deep: Row/Column level masking. | Basic: Allow/Deny access control. |
| Transports | SSE, Streamable HTTP, WebSocket | SSE, Streamable HTTP, WebSocket |
| Legacy Integration | Native: REST/RPC to MCP transformation. | External: Manual wrappers required. |
| Authorization | Managed: Roles, Groups, Positions, Attributes. | Infrastructure: CEL-based policies. |
| Hot-Reloading | Native: Integrated Control Plane & Registry. | Infrastructure: Istio/xDS sync. |
| Authentication | JWT (End-to-End Propagation) | JWT, Keycloak, OIDC, Passthrough |
| Observability | Distributed Tracing (OTEL) and Integrated Hindsight Memory | Distributed Tracing (OTEL) |
1. Architectural Intent
AgentGateway: The Network Proxy Layer
AgentGateway is designed as a high-availability proxy for MCP servers. Its primary focus is the North-South traffic between an application and multiple MCP backends.
- Multiplexing: Optimized for merging multiple MCP backends into a single upstream connection (
mergestream.rs). - Protocol Translation: Excels at translating between SSE, Streamable HTTP, and WebSocket transports.
- Infrastructure Focus: Operates as a Kubernetes-native component managed via manifests and standard networking policies.
Light-Fabric: The Managed Enterprise Platform
Light-Fabric provides a Unified Governance Fabric that treats AI agents and MCP tools as part of the broader enterprise API ecosystem.
- Unified Gateway: The AI Gateway (Rust/Pingora-based) serves as a single entry point for UI, Agents, and Tools, supporting both MCP and traditional REST/RPC APIs.
- Centralized Portal: Uses the Light-Portal as a control plane for onboarding (auto-discovery), configuration (hot-reloading), and security management.
- Governed Intelligence: Integrates the gateway directly with Hindsight Memory and the Fine-Grained Rule Engine, ensuring that every tool call is governed by corporate compliance rules (e.g., row/column masking).
- End-to-End Security: Maintains a single JWT-based identity from the user’s chat interface all the way to the underlying MCP or API endpoint.
2. Security & Authorization
AgentGateway: Infrastructure-Aware RBAC
AgentGateway uses Common Expression Language (CEL) for its authorization policies.
- Capabilities: High-speed, network-level blocking based on JWT claims and request headers.
- Limitation: Lacks native support for content-aware data masking or organizational hierarchy logic.
Light-Fabric: Content-Aware Managed Auth
Light-Fabric provides a mature Fine-Grained Authorization layer:
- Managed ABAC/PBAC: Supports Role, Group, Corporate Position (Hierarchy), and Attribute-based protection.
- Data Privacy: Supports native Row and Column filtering (data masking), ensuring agents only see data they are authorized to process.
- End-to-End JWT: The same JWT token is propagated from the UI through the Agent to the AI Gateway and MCP tool.
3. Lifecycle & Tool Onboarding
AgentGateway: Configuration-Driven
Onboarding tools in AgentGateway is an infrastructure task:
- Manual Mapping: Requires defining Kubernetes Custom Resources (
HTTPRoute,Backend) to map MCP servers to the gateway. - Scope: Primarily focused on exposing existing MCP servers.
Light-Fabric: Registry-Driven
Light-Fabric provides a “Zero-Effort” onboarding experience via Light-Portal:
- Auto-Discovery: Registering an MCP API triggers an automatic
tools/listcall to populate the registry. - Protocol Transformation: Automatically transforms existing OpenAPI/REST and RPC services into MCP tools without requiring wrappers.
- Centralized Governance: All tools (Native, REST, MCP) are managed in a single unified registry.
4. Control Plane & Configuration
AgentGateway: Kubernetes-Native
- Orchestration: Managed via the Istio/xDS control plane.
- Updates: Configuration changes are applied via Kubernetes manifests (YAML).
Light-Fabric: Portal-Managed
- Hot-Reloading: Uses a dedicated Config Server and Control Plane to update gateway and agent configurations in real-time without restarts.
- Enterprise Management: Business-centric UI for managing tool visibility, agent permissions, and security policies.
5. Conclusion
- Use AgentGateway if you are an infrastructure provider who needs to expose MCP-based tools to multiple external applications securely and reliably.
- Use Light-Fabric if you are building intelligent agents that need to use those tools to solve complex business problems within a governed framework.
Why Light-Fabric Already Covers the MCP Gateway — No Second Gateway Required
This document addresses a recommendation (produced by Grok AI) suggesting that an enterprise should deploy the open-source AgentGateway as a dedicated MCP layer alongside an existing API platform. After performing a side-by-side source code analysis of both projects (see vs-agentgateway.md and vs-agent-gateway-mcp.md), we present the findings below.
1. The Recommendation Was Generated Without Knowledge of Light-Fabric
The Grok-produced analysis operates under a critical blind spot: it has no knowledge of Light-Fabric (Rust-based, open-sourced to customers) or its capabilities. The recommendation frames the choice as “keep your existing REST platform + add AgentGateway for MCP,” because Grok only knows about publicly documented open-source projects. It does not account for the fact that:
- Light-Fabric is already in production and serving agentic workloads today.
- Every feature listed in the recommendation — MCP federation, tool discovery, protocol translation, security, and observability — has already been built, demonstrated, and validated with the project team.
- The comparison is therefore not between “a REST framework” and “an MCP gateway.” It is between two systems that both provide MCP gateway capabilities, where one (Light-Fabric/Light-Gateway) is already deployed and battle-tested in our environment.
2. Source Code Analysis: Light-Fabric Already Does What AgentGateway Does
We conducted a detailed, code-level comparison of both projects. The full results are documented in our High-Level Comparison and Detailed MCP Feature Comparison. The key findings are summarized below.
2.1 MCP Protocol Support
| Capability | Light-Fabric | AgentGateway |
|---|---|---|
| Transports | SSE, Streamable HTTP, WebSocket | SSE, Streamable HTTP, WebSocket |
| Tool Discovery | Auto-discovery via tools/list sync | Manual K8s CRD configuration |
| Protocol Translation | Native REST/RPC → MCP transformation | Manual wrappers required |
| Stream Handling | Supported | Supported (mergestream) |
Both projects support the same MCP transports. Light-Fabric goes further with automatic tool discovery and native protocol transformation from existing REST/RPC APIs — exactly the “OpenAPI-to-MCP mapping” that the Grok recommendation credits to AgentGateway, except Light-Fabric does it without requiring a separate component.
2.2 Security & Authorization
| Capability | Light-Fabric | AgentGateway |
|---|---|---|
| Authentication | JWT (end-to-end propagation) | JWT, Keycloak, OIDC, Passthrough |
| Authorization | Role, Group, Position, Attribute-based (ABAC/PBAC) | CEL-based policies |
| Data Privacy | Row/Column-level masking | Allow/Deny access control |
| Rule Engine | Integrated YAML-based, hot-reloadable | Basic middleware |
The Grok recommendation highlights “tool-level RBAC” and “MCP-compliant OAuth 2.1” as AgentGateway strengths. Our code analysis shows that Light-Fabric’s authorization model is significantly deeper — it supports corporate-hierarchy-aware policies and content-level data masking that AgentGateway simply does not implement.
2.3 Lifecycle & Operations
| Capability | Light-Fabric | AgentGateway |
|---|---|---|
| Onboarding | Portal-driven, auto-discovery | K8s manifest-driven, manual |
| Hot-Reloading | Native (Config Server + Control Plane) | Infrastructure-dependent (Istio/xDS) |
| Observability | OTEL + integrated Hindsight Memory | OTEL + OpenInference |
| Orchestration | Integrated hybrid workflows (deterministic + autonomous) | None (stateless proxy) |
Light-Fabric manages the entire lifecycle — from tool registration through governance to runtime orchestration — while AgentGateway only handles the proxy layer.
3. Two Gateways Is Overkill
The Grok recommendation frames the architecture as a “clean separation of concerns.” In practice, deploying both Light-Fabric and AgentGateway creates redundant infrastructure with real costs:
Duplicated Capabilities
Both systems would be performing the same core functions:
- Receiving MCP requests from agents
- Translating tool calls to backend HTTP requests
- Enforcing security policies on tool access
- Providing observability for agentic traffic
Running two gateways that do the same thing is not “separation of concerns” — it is duplication of concerns. Every MCP request would traverse two proxy layers instead of one, adding latency and operational complexity for zero additional capability.
Operational Burden
- Two deployment pipelines to maintain on EKS
- Two sets of security policies to keep in sync
- Two configuration surfaces (K8s CRDs for AgentGateway vs. Portal for Light-Fabric)
- Two failure domains to monitor and troubleshoot
- Two upgrade cycles to coordinate
The “No Code Changes” Claim Is Misleading
The Grok recommendation states AgentGateway requires “no code changes.” This is true only if you ignore the work required to:
- Write and maintain Kubernetes Custom Resources for every MCP backend
- Build manual wrappers for non-MCP services (Light-Fabric does this natively)
- Implement application-level logic for everything AgentGateway doesn’t cover (stateful workflows, data masking, memory management)
Light-Fabric also requires no code changes to existing backend services — and it provides the governance layer out of the box.
4. Addressing the “Rust Performance” Argument
The recommendation claims AgentGateway has a “performance edge” due to its Rust data plane. This argument does not hold:
- Light-Fabric’s AI Gateway currently runs on the high-performance Java-based light-gateway, and a new Rust-based AI Gateway is also under way, built on the Pingora framework (Cloudflare’s production proxy engine). Even the existing Java gateway delivers exceptional throughput, and the Rust gateway will remove the JVM from the critical path entirely.
- Both systems benefit from Rust’s zero-cost abstractions, memory safety, and lack of garbage collection pauses.
- The performance comparison between the two Rust implementations would be marginal and workload-dependent — not a differentiator.
5. Addressing the “Custom Development” Concern
The recommendation warns against “implementing MCP directly” because it “involves significant custom development.” This concern does not apply:
- Light-Fabric’s MCP support is not custom development — it is a fully implemented, production-ready feature of the platform.
- The MCP client, gateway routing, tool registry, and security integration are all existing, tested components, not a backlog of work to be done.
- The project team has already seen these features demonstrated end-to-end.
6. Summary
| Concern from Grok Recommendation | Reality |
|---|---|
| “Light4j is a REST framework, not an AI proxy” | Light-Fabric is a full agentic platform with an AI Gateway already in production |
| “AgentGateway provides MCP federation and tool discovery” | Light-Fabric provides the same capabilities with deeper governance |
| “Rust performance advantage over JVM” | Light-Fabric’s Java gateway is already very fast, and a Rust (Pingora-based) gateway is coming |
| “Clean separation of concerns” | Two gateways doing the same thing is duplication, not separation |
| “No code changes required” | True for both — but AgentGateway requires extensive K8s manifest management |
| “Custom MCP implementation is risky” | Light-Fabric’s MCP support is already built, tested, and in production |
Conclusion
The Grok-generated recommendation is well-structured but fundamentally flawed because it was produced without knowledge of Light-Fabric’s capabilities. When evaluated against the actual source code and production state of both systems, the case for adding AgentGateway collapses:
- Light-Fabric already provides every MCP gateway capability that AgentGateway offers.
- Light-Fabric goes significantly further with integrated governance, data privacy, memory, and orchestration.
- Adding a second gateway introduces operational complexity and latency with no net-new capability.
The pragmatic, low-risk path is to continue with the platform that is already built, already in production, and already proven to the team.