vgi-rpc Wire Protocol Specification¶
Wire protocol version: 1 Status: Normative Audience: Cross-language implementors (Go, Rust, TypeScript, C++, etc.) Reflects: vgi-rpc 0.45.3
This document specifies the vgi-rpc wire protocol at byte level. A conforming implementation can interoperate with the Python reference without reading Python source code. The protocol is transport-agnostic; specific transport bindings (pipe, HTTP, shared memory) are described in later sections.
1. Overview & Conventions¶
vgi-rpc is an RPC framework where:
- Serialization uses Apache Arrow IPC Streaming Format.
- All integers are little-endian unless stated otherwise.
- All metadata strings are UTF-8 encoded.
- Wire protocol version:
"1"(the single ASCII byte0x31). - Metadata keys and values in Arrow IPC custom metadata are byte strings. Keys in the
vgi_rpc.*namespace are framework-reserved.
Terminology¶
| Term | Definition |
|---|---|
| IPC stream | A complete Arrow IPC streaming-format message sequence: schema message, zero or more record batch messages, terminated by an EOS marker. |
| Batch | An Arrow RecordBatch — zero or more rows conforming to a schema. |
| Custom metadata | Per-batch KeyValueMetadata attached to individual record batches within an IPC stream (distinct from schema-level metadata). |
| Zero-row batch | A batch with num_rows == 0. Used for log messages, error signals, pointer batches, and stream-completion markers. |
| Data batch | A batch with num_rows > 0, or a zero-row batch that lacks log/error metadata keys (e.g., void return). |
2. Arrow IPC Framing¶
Each logical message exchange uses one or more IPC streams written sequentially on the same byte stream (pipe, TCP socket, HTTP body, etc.).
An IPC stream consists of:
- Schema message — describes the columns and their Arrow types.
- Zero or more RecordBatch messages — each optionally carrying per-batch custom metadata.
- EOS marker — the 8-byte sequence
0xFF 0xFF 0xFF 0xFF 0x00 0x00 0x00 0x00(continuation token0xFFFFFFFFfollowed by 4 zero bytes for metadata length).
Multiple IPC streams are written sequentially on the same underlying byte stream. Each reader opens one stream, reads until EOS, and stops. The next reader picks up immediately after the EOS marker.
Refer to the Apache Arrow IPC specification for the byte-level encoding of schema messages, record batch messages, dictionary messages, and the encapsulated message format.
3. Metadata Key Reference¶
All framework-reserved metadata keys, their wire-format byte representations, where they appear, and their semantics:
Request metadata (on the request batch's custom metadata)¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.method |
UTF-8 method name | Target RPC method to invoke. Required. |
vgi_rpc.protocol |
UTF-8 protocol name | Which protocol the method belongs to — the routing key. Required, including against a server hosting exactly one protocol. See Section 3.1. |
vgi_rpc.request_version |
"1" (ASCII 0x31) |
Wire protocol version. Required. |
vgi_rpc.protocol_version |
Canonical semver MAJOR.MINOR.PATCH |
Application protocol surface version. Required when the peer Protocol declares one, absent otherwise. See Section 13. |
vgi_rpc.request_id |
UTF-8 string (16-char hex) | Per-request correlation ID. Optional; if absent, the server generates a new 16-char hex ID. |
vgi_rpc.cancel |
"1" — presence is the signal |
Client-initiated stream cancellation, on a stream input batch. See Section 9. Optional. |
traceparent |
W3C Trace Context string | OpenTelemetry trace propagation. Optional. |
tracestate |
W3C Trace Context string | OpenTelemetry trace state. Optional. |
vgi_rpc.shm_segment_name |
UTF-8 OS name | Shared memory segment name (session-level). Optional. |
vgi_rpc.shm_segment_size |
Decimal integer string | Shared memory segment total size in bytes. Optional. |
vgi_rpc.transport.shm |
"true" / "false" |
Client's shared-memory capability, on the __transport_options__ request. See Section 15. Optional. |
HTTP stream continuation requests additionally carry
vgi_rpc.stream_state#b64 and vgi_rpc.call_state#b64 — see the stream-state
table below.
Response / log / error metadata (on response batch custom metadata)¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.log_level |
One of: EXCEPTION, ERROR, WARN, INFO, DEBUG, TRACE |
Severity level. Present on log and error batches. |
vgi_rpc.log_message |
UTF-8 string | Human-readable message text. |
vgi_rpc.log_extra |
JSON string | Additional structured data. Optional. |
vgi_rpc.error_kind |
UTF-8 token (open set) | Stable machine-readable error category on EXCEPTION batches — the reason. See Section 8. Optional. |
vgi_rpc.error_code |
Canonical code name (closed set) | The error's canonical code, e.g. UNAVAILABLE, on EXCEPTION batches. Required whenever vgi_rpc.error_kind is set, and emitted on every EXCEPTION batch by a server implementing Section 8's error model. |
vgi_rpc.error_details |
JSON array of typed objects | Machine-readable specifics from a fixed catalog (vgi_rpc.RetryInfo, …), on EXCEPTION batches. At most 4 KiB; omitted whole when larger. See Section 8. Optional. |
vgi_rpc.server_id |
UTF-8 string (12-char hex) | Server instance identifier for distributed tracing. |
vgi_rpc.request_id |
UTF-8 string | Echoed request correlation ID. |
vgi_rpc.transport.shm |
"true" / "false" |
Server's shared-memory capability, on the __transport_options__ response. See Section 15. |
Stream state (HTTP transport)¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.stream_state#b64 |
Base64-encoded binary (signed token) | Serialized stream state for stateless HTTP exchanges — the per-turn cursor when the server splits its state. The #b64 suffix signals that the value is base64-encoded binary data. |
vgi_rpc.call_state#b64 |
Base64-encoded binary (signed token) | Optional. The stream's call state — the half fixed for the life of the call. Minted once by /init, never re-issued, and echoed by the client on every subsequent request. Absent from servers that keep everything in the cursor. |
Shared memory pointer batch metadata¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.shm_offset |
Decimal integer string | Absolute byte offset in the SHM segment. |
vgi_rpc.shm_length |
Decimal integer string | Number of bytes of the serialized batch. |
vgi_rpc.shm_source |
UTF-8 SHM segment name | Provenance indicator on resolved batches (diagnostics). |
External storage pointer batch metadata¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.location |
UTF-8 URL | URL to fetch the externalized batch data. |
vgi_rpc.location.sha256 |
UTF-8 hex string | SHA-256 of the uploaded payload before compression. Optional; when present the reader MUST verify it after fetching. |
A pointer batch on the wire carries only the two keys above. It MUST NOT carry either provenance key below: those are stamped by the reader, and a writer that emits one produces a resolved batch whose provenance names the writer's own guess rather than the URL that was actually fetched.
Resolved batch metadata (added by the reader)¶
Neither key below is wire content. Both are added by the resolving reader at resolve time, on the batch it returns — see Resolution (reading).
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.location.fetch_ms |
Decimal float string (e.g. "42.3") |
Elapsed fetch time in milliseconds, measured by the reader. |
vgi_rpc.location.source |
UTF-8 URL | The URL the reader fetched, in full. |
location.source records the URL complete with its query string, and a
presigned URL's query string is a bearer credential. That is deliberate — the
value has to identify the object that was actually fetched, and a redacted form
identifies a different URL — but it has a consequence worth stating, because
one port redacted this key for exactly this reason before the rule was written
down.
The redaction rule under Fetch safety governs diagnostics: errors, logs,
traces and exception chains. It does not govern this key. The same URL is
already on the wire in vgi_rpc.location, so carrying it here exposes it to
the consuming application rather than to a new network observer — but note it
outlives the pointer, which a resolved batch MUST NOT carry. Treat
location.source as credential-bearing: it is safe to compare, attribute and
cache on, and it MUST NOT be written to logs or telemetry without the same
redaction a diagnostic URL gets.
Introspection batch metadata (on __describe__ response batch custom_metadata)¶
| Key (bytes) | Value | Description |
|---|---|---|
vgi_rpc.protocol_name |
UTF-8 string | Protocol class name. |
vgi_rpc.request_version |
"1" |
Wire protocol version. |
vgi_rpc.describe_version |
"4" |
Introspection format version. |
vgi_rpc.protocol_hash |
UTF-8 hex string | SHA-256 digest over the canonical describe payload. |
vgi_rpc.protocol_version |
Canonical semver | Application protocol surface version. Present only when the Protocol declares one. |
vgi_rpc.server_id |
UTF-8 string | Server instance identifier. |
3.1 Protocol routing¶
A server hosts one or more protocols. Every request names the one it
addresses, and dispatch resolves the pair (protocol, method). Method names may
collide across protocols — that is what lets protocols be authored
independently, and a port that merges them into one namespace is not conformant.
Names¶
A protocol name is an identifier, optionally dot-qualified:
At most 255 bytes UTF-8. The vgi_rpc. prefix is reserved for protocols
the framework itself defines; a server MUST refuse to register an application
protocol claiming it, because an application that could claim
vgi_rpc.Reflection.v1 could shadow the one surface a client trusts before it
knows anything else about the server.
The major version is part of the name — vgi_rpc.Identity.v1,
vgi_rpc.Reflection.v1 — following gRPC (AIP-185), Kubernetes API groups and
D-Bus. Two consequences, both deliberate:
- An incompatible major is a different protocol, so addressing it is a routing
failure (
protocol_not_supported, HTTP 404). Every proxy, WAF and load balancer understands that answer without an Arrow parser. .v1and.v2can be served side by side while clients migrate. That is what makes a major change rollable rather than a flag day.
Carriage¶
On HTTP the protocol rides twice: in vgi_rpc.protocol and as a path segment,
{prefix}/{protocol}/{method}.
The metadata field is canonical. It is the only carrier on the stdio, unix and named-pipe transports. The path segment is a required faithful projection, present so an edge device can act on the protocol without parsing Arrow.
A server MUST reject a request whose two carriers disagree
(protocol_not_supported, HTTP 400). Unspecified, this is the
Content-Length/Transfer-Encoding shape: the edge applies policy to one protocol
while the worker dispatches another.
Two further rules close that gap:
- A
%anywhere in the protocol path segment is rejected without decoding (HTTP 404). The name charset never requires percent-encoding, so a percent sign is a bug or an attempt to make the edge and the worker read different strings. Compare raw bytes; never compare decoded-against-raw. - Stream continuations re-verify. A continuation (
/exchange, cancel) must stay on the protocol its stream started on. The required property is that a continuation carrying a token minted under another protocol is refused. Binding the protocol into the AEAD associated data of the cursor and call tokens is how the reference achieves it, and is recommended: there is no comparison code to get wrong, and it covers the call-state cache-hit path, where the call token is never opened at all. An implementation that enforces the property another way is conformant.
The AAD construction itself is not a wire contract. Associated data never crosses the wire, and a sealed token is only ever opened by the implementation that minted it — so the prefix strings, their version numbering, and the byte layout of the identity tail are internal to a worker framework and MAY differ between implementations. Ports are not required to match the reference here, and a port that happens to match today is not promising to keep matching. Do not build an implementation that depends on opening another implementation's tokens; that is not a supported deployment, and nothing in the conformance suite can observe it.
A name that cannot match the grammar is rejected before it is looked up, so a request-supplied string never reaches an error message, a log field or a metric label.
Required, with no single-protocol exemption¶
vgi_rpc.protocol is required even when the server hosts exactly one protocol.
An exemption would let an intermediary that rebuilds a request and drops the
field land silently on whichever protocol happened to be first, rather than
being told. The three answers are distinct and a client depends on the
difference:
| Condition | Answer | error_kind |
|---|---|---|
| No routing key on the request | refuse | protocol_not_specified |
| Named protocol not hosted here | 404 | protocol_not_supported |
| Protocol hosted, method absent | 404 | method_not_implemented |
The last is the documented capability-probe signal: a client testing for an optional method must be able to tell "you do not speak this protocol" from "you speak it but lack this method".
Hosting several application protocols¶
A server MUST let an application register any number of application
protocols when it is constructed, each a (name, version, implementation)
triple — in the reference, RpcServer(primary, impl, extra_protocols=[(P2,
impl2), …]). Hosting is not optional machinery for exotic deployments: a
worker that serves its own surface and also a shared one (a reporting
protocol, a fixture, a later major version beside the current one) needs it,
and a port that can host exactly one application protocol forces every such
worker to fork the framework.
- The protocol is the unit of optionality. There is no API for hosting a
subset of a protocol's methods, and no capability tokens for application
protocols (Section 14 reserves
features). A capability that may be absent is its own protocol, so a client learns whether it is present fromlist_protocolsrather than by calling and reading an error. (vgi_rpc.Identity.v1's hook-driven method narrowing, Section 16, is framework-owned and unaffected.) - The registered set is fixed for the server's lifetime. It may be computed
from configuration or environment when the server is built, but it does not
change afterwards, so reflection output and every
protocol_hashare stable for the life of the process. The set is sealed no later than when the server first starts serving on any transport: an attempt to register a protocol after that MUST fail loudly (an error or exception at the call), never take effect silently or on some transports only. A port may seal earlier — the reference has no registration API after construction at all. - The same set is hosted on every transport the server is offered on.
A protocol reachable over stdio but not HTTP (or the reverse) is a different
server wearing one name. Framework protocols that are transport-scoped by
definition —
vgi_rpc.Identity.v1on transports that authenticate callers — are the only exception, and are not application protocols. - The reserved-prefix rule applies to every registered protocol, however its
name was derived — declared explicitly, taken from a class or interface
name, or synthesised by a binding generator. A port that checks only names
declared one way lets the other way shadow
vgi_rpc.Reflection.v1. - Names are unique. Registering two protocols under one name is a construction-time error, not last-writer-wins.
- Order is registration order, primary first.
list_protocolslists the application protocols in the order they were registered, the primary (the one passed first) first. Framework protocols (vgi_rpc.*) may appear anywhere in the list; clients locate application protocols by filtering the reserved prefix out, and the relative order of what remains is the contract. The client driver'sdescribeop and every client's "describe this server" already rely on "the first protocol whose name does not start withvgi_rpc.".
Each binding is versioned, gated (Section 13) and hashed independently; nothing about a server is "the" protocol except which one is listed first.
4. Type Mapping¶
RPC method parameters and return values are serialized as Arrow columns. The following table defines the canonical mapping from abstract types to Arrow types. Cross-language implementations MUST use these Arrow types for interoperability.
| Abstract type | Arrow type | Serialization notes |
|---|---|---|
string |
utf8 |
UTF-8 encoded. |
bytes / binary |
binary |
Raw byte sequence. |
int / integer |
int64 |
64-bit signed integer. |
float / double |
float64 |
IEEE 754 double precision. |
bool |
bool |
— |
list[T] |
list(T) |
Recursive. |
dict[K, V] / map |
map(K, V) |
Serialized as list of (key, value) tuples. Deserialized back to map/dict. |
frozenset[T] / set[T] |
list(T) |
Serialized as list (order undefined). Deserialized back to set. |
enum |
dictionary(int16, utf8) |
Serialized as the enum member name (string). Deserialized by name lookup. |
optional[T] / T? |
Same as T, but with nullable = true on the Arrow field. |
null represents the absent value. |
dataclass (nested) |
binary |
Serialized as a complete Arrow IPC stream (schema + 1-row batch + EOS) in a binary column. See note below. |
Nested dataclass type context: This
binarymapping applies at the RPC method parameter/return level — each dataclass parameter or return value is a binary blob containing a serialized IPC stream. Within that IPC stream, the dataclass's own Arrow schema uses the same type mapping for primitive fields (string→utf8, int→int64, etc.), but sub-dataclass fields use Arrowstructtype (notbinary), since they are embedded inline rather than serialized as separate IPC streams.
Serialization transforms¶
When writing a value to an Arrow column:
- Enum → write the member's name as a UTF-8 string (not its value).
- dict → convert to a list of
(key, value)tuples, then write asmap(K, V). - frozenset/set → convert to a list, then write as
list(T). - Nested dataclass → serialize to Arrow IPC bytes, write as
binary.
When reading:
- Enum → look up the string by member name (
Enum["NAME"]). Name-based lookup is the normative wire format. (The Python reference implementation also supports a value-based fallback internally for nested dataclass fields, but cross-language implementations need only implement name-based lookup.) - map → convert list of tuples back to dict.
- list (when target is set) → convert to frozenset/set.
- binary (when target is dataclass) → deserialize from Arrow IPC bytes.
5. Request Batch Format¶
Every RPC request is a single IPC stream containing exactly one batch with one row:
IPC Stream:
Schema message:
- One field per method parameter, named after the parameter
- Field types per the type mapping (Section 4)
- Optional parameters have nullable = true
RecordBatch message:
- Exactly 1 row
- custom_metadata:
vgi_rpc.method = "<method_name>" (REQUIRED)
vgi_rpc.request_version = "1" (REQUIRED)
vgi_rpc.shm_segment_name = "<name>" (optional, SHM transport)
vgi_rpc.shm_segment_size = "<size>" (optional, SHM transport)
traceparent = "<W3C trace context>" (optional)
tracestate = "<W3C trace state>" (optional)
EOS marker
For methods with no parameters, the schema has zero fields and the batch has one row with zero columns.
Row count validation: The server only enforces
num_rows == 1when the schema has one or more fields. For zero-field (parameterless) methods, the server accepts batches with any row count (including 0). Conforming clients SHOULD send 1 row for consistency.
Worked example¶
Given an RPC method add(a: float, b: float) -> float:
Schema:
Batch (1 row, calling add(a=1.0, b=2.0)):
Column "a": [1.0]
Column "b": [2.0]
custom_metadata: {
"vgi_rpc.method": "add",
"vgi_rpc.request_version": "1"
}
Default values: When a parameter has a default and the caller omits it, the client merges the default into the kwargs before serialization. The server sees a complete row in all cases.
6. Response Format (Unary)¶
A unary response is a single IPC stream on the result schema:
IPC Stream:
Schema message:
- For methods returning a value: single field named "result"
- For void methods (-> None): zero fields (empty schema)
0..N log batches (zero-row, with log metadata — see Section 8)
1 result or error batch:
- Result: 1-row batch with the return value in column "result"
- Void: 0-row batch on empty schema
- Error: 0-row batch with EXCEPTION-level log metadata (see Section 8)
EOS marker
Log batches MUST appear before the result/error batch. They share the same schema as the result batch (the zero-row log batches conform to the response stream's schema).
Void return¶
When the method has no return value (-> None), the response schema is
empty (pa.schema([])) and the result batch has zero rows and zero columns.
7. Batch Classification Algorithm¶
When receiving any batch from a response stream, classify it using this decision tree:
receive(batch, custom_metadata):
IF custom_metadata is NULL:
→ DATA batch
IF batch.num_rows > 0:
→ DATA batch
// At this point: num_rows == 0 AND custom_metadata exists
IF custom_metadata contains "vgi_rpc.log_level"
AND custom_metadata contains "vgi_rpc.log_message":
level = custom_metadata["vgi_rpc.log_level"]
IF level == "EXCEPTION":
→ ERROR batch → raise RpcError (see Section 8)
ELSE:
→ LOG batch → deliver to on_log callback
IF custom_metadata contains "vgi_rpc.location":
→ EXTERNAL POINTER batch → resolve via URL fetch (see Section 12)
IF custom_metadata contains "vgi_rpc.shm_offset":
→ SHM POINTER batch → resolve via shared memory (see Section 11)
IF custom_metadata contains "vgi_rpc.stream_state#b64":
→ STATE TOKEN batch → stream continuation (see Section 10)
// Zero-row batch with unrecognized metadata
→ DATA batch (e.g., void return, stream-finish marker)
Note: Log-level keys take priority. A zero-row batch that has both
vgi_rpc.log_levelandvgi_rpc.shm_offset(orvgi_rpc.location) is classified as a log batch, not a pointer. This is by design — pointer detection explicitly excludes batches with log-level keys. A batch cannot be both an external pointer and an SHM pointer simultaneously, so the check order between those two does not matter functionally.
8. Log & Error Batch Format¶
Log batches¶
A log batch is a zero-row batch on the response stream's schema, with the following custom metadata keys:
| Key | Required | Value |
|---|---|---|
vgi_rpc.log_level |
Yes | One of: EXCEPTION, ERROR, WARN, INFO, DEBUG, TRACE |
vgi_rpc.log_message |
Yes | Human-readable message text (UTF-8) |
vgi_rpc.log_extra |
No | JSON object with additional structured data |
vgi_rpc.error_kind |
No | Stable error category — the reason; EXCEPTION batches only (see below) |
vgi_rpc.error_code |
On EXCEPTION | Canonical code name; EXCEPTION batches only (Error model) |
vgi_rpc.error_details |
No | JSON array of typed details, ≤ 4 KiB; EXCEPTION batches only (Error model) |
vgi_rpc.server_id |
No | Server instance identifier |
vgi_rpc.request_id |
No | Request correlation ID |
Error batches (EXCEPTION level)¶
When vgi_rpc.log_level is "EXCEPTION", the batch represents a server-side
error. The client MUST raise/throw an error with the following fields
extracted from the metadata:
- error_type:
log_extra.exception_type(string) or the level string"EXCEPTION"as fallback. - error_message:
vgi_rpc.log_messagevalue. - remote_traceback:
log_extra.traceback(string) or empty string. Empty when the operator turned tracebacks off (Tracebacks). - request_id:
vgi_rpc.request_idvalue or empty string. - error_code:
vgi_rpc.error_codevalue, or empty string when absent (a server that predates the error model). - error_kind:
vgi_rpc.error_kindvalue, or empty string when absent. - error_details: the decoded
vgi_rpc.error_detailsarray — every element, in order, unknown types included — or an empty list when absent.
Each of the last three is read from its top-level key first and from the
log_extra mirror (error_code, error_kind, error_details) when the
top-level key is absent. A client MUST surface all three on every decode
path — unary, stream init, stream exchange, an externalized error batch — and
expose an is_retryable check (Retryability). Dropping one
on any path is the defect this list exists to prevent: three clients dropped
error_kind because the client-driver contract never asked for it.
Error kinds¶
vgi_rpc.error_kind carries a stable identifier for errors a client is
expected to branch on, so callers pattern-match a token instead of
substring-searching a human-readable message. It is emitted as a top-level
metadata key, mirroring log_extra.error_kind — a reader may take either, but
the top-level key means the JSON blob need not be parsed to dispatch.
The set is open: a client MUST treat an unrecognised value as an unclassified error rather than rejecting the batch. A kind is unique within the protocol that raised it; the pair (protocol, kind) is gRPC's (domain, reason). Each kind names exactly one canonical code, fixed where the kind is defined, and the code is required whenever the kind is set. Well-known values:
| Value | Code | Details | Meaning |
|---|---|---|---|
method_not_implemented |
UNIMPLEMENTED |
The named protocol is hosted but has no such method (old server vs. new client, or a method that was removed). The intended signal for capability detection with fallback. | |
protocol_not_specified |
INVALID_ARGUMENT |
The request carried no vgi_rpc.protocol routing key. Required even against a single-protocol server — see Section 3.1. |
|
protocol_not_supported |
UNIMPLEMENTED |
This server does not host the named protocol, or the path and the metadata named different ones. Also the answer for an incompatible major, since the major is part of the name. | |
protocol_version_mismatch |
FAILED_PRECONDITION |
PreconditionFailure |
The client's vgi_rpc.protocol_version is incompatible with the server's for the protocol the resolved method belongs to (see Section 13). The violation has type "protocol_version" and subject the protocol's name. |
session_lost |
ABORTED |
An HTTP sticky-session token could not be honoured — expired, evicted, misrouted, or presented under a different principal (see Section 17). Retry the whole session, not the call. | |
server_draining |
UNAVAILABLE |
RetryInfo |
The server is shutting down and refuses new sticky-session opens. |
identity_unavailable |
UNAVAILABLE |
RetryInfo (required) |
See Section 16. |
stale_auth |
UNAUTHENTICATED |
See Section 16. | |
introspection_refused |
PERMISSION_DENIED |
See Section 16. | |
grant_refused |
PERMISSION_DENIED |
See Section 16. | |
token_unresolved |
NOT_FOUND |
See Section 16. |
Normative client behaviour on protocol_not_supported: a client SHOULD
call vgi_rpc.Reflection.v1/list_protocols to learn what the server does host,
and MUST surface both what it asked for and what it was told. Unspecified, six
ports invent six answers to the same question, and the one thing a user needs
in that moment — the two names side by side — is the thing most likely to be
dropped.
A batch carrying error_kind is otherwise an ordinary EXCEPTION batch: the
key adds classification and removes nothing.
Error model¶
Errors take the shape of gRPC's google.rpc.Status — canonical codes, a
reason, and typed details, with AIP-193 as the
usage guide — because that shape has held up across many languages for years.
Three layers ride on every EXCEPTION batch:
| Layer | Key | Set | Purpose |
|---|---|---|---|
| Code | vgi_rpc.error_code |
Closed: the sixteen below | Generic handling — retry or not, how to show it, and the HTTP status a proxy maps it to |
| Reason | vgi_rpc.error_kind |
Open; unique within the raising protocol | What a client branches on |
| Details | vgi_rpc.error_details |
Fixed catalog | Machine-readable specifics: retry delay, which field, which resource |
vgi_rpc.log_message stays developer-facing English, as in gRPC.
Codes¶
The wire value is the code's name — "UNAVAILABLE", never 14 — so a
log line, a proxy rule and a client switch read the same string. gRPC's
sixteen non-OK codes, no more:
CANCELLED, UNKNOWN, INVALID_ARGUMENT, DEADLINE_EXCEEDED, NOT_FOUND,
ALREADY_EXISTS, PERMISSION_DENIED, RESOURCE_EXHAUSTED,
FAILED_PRECONDITION, ABORTED, OUT_OF_RANGE, UNIMPLEMENTED, INTERNAL,
UNAVAILABLE, DATA_LOSS, UNAUTHENTICATED.
A client that receives any other value treats it as UNKNOWN.
Detail catalog¶
Each detail is a JSON object naming its type in @type, mirroring protobuf's
JSON form for Any. Field names follow gRPC's, in snake_case:
@type |
Fields | Use |
|---|---|---|
vgi_rpc.ErrorInfo |
metadata: {string: string} |
Extra context for the reason. The reason and domain are already error_kind and the protocol, so they are not repeated |
vgi_rpc.RetryInfo |
retry_delay_seconds: number |
How long to wait before retrying; finite, ≥ 0 |
vgi_rpc.BadRequest |
field_violations: [{field, description}] |
Which inputs were wrong |
vgi_rpc.PreconditionFailure |
violations: [{type, subject, description}] |
What state must change first |
vgi_rpc.QuotaFailure |
violations: [{subject, description}] |
Which limit was hit |
vgi_rpc.ResourceInfo |
resource_type, resource_name, owner, description |
Which object the error concerns |
vgi_rpc.Help |
links: [{description, url}] |
Where to read more |
vgi_rpc.LocalizedMessage |
locale, message |
Text safe to show an end user |
DebugInfo is deliberately absent — see Tracebacks. String
fields absent from an object read as ""; arrays as [].
Rules¶
- The code is required whenever
error_kindis set, and each kind maps to exactly one code (the table above; protocols define the codes of their own kinds where they define the kinds). A server implementing this model emits a code on every EXCEPTION batch: errors with no classification getUNKNOWN. A code may be sent without a kind.INTERNALis reserved for faults an implementation itself detects in its own machinery — an invariant violated, a response it cannot encode — and MAY be emitted for them; it is never the default for an unclassified error, and conformance does not require any particular fault to produce it. (The reference has no such classified site today and emitsUNKNOWNfor every unclassified error, including framework-raised ones without a kind.) - Details come only from the catalog. A protocol may define its own detail
type only under its own protocol name (
vgi.reports.v1.SomeDetail), and should preferErrorInfo.metadata. A type in the reservedvgi_rpc.space that is not in the catalog, or an unqualified type name, is not a detail. - Each type appears at most once in a details array, as AIP-193 requires.
- Clients ignore detail types they do not know, and never require details. "Ignore" means typed access skips them; the client still reports the whole array as received.
- Bounded. The serialized
vgi_rpc.error_detailsvalue is at most 4096 bytes of UTF-8, measured as emitted. A server whose array would exceed it omits the whole array — both the top-level key and thelog_extramirror — and never sends a prefix or a subset, because a client cannot tell a partial list from a complete one. A server likewise omits an array that breaks the uniqueness or catalog rules. The code and kind are sent regardless. - No secrets. Details MUST NOT carry credentials, tokens or user data.
- Mirrored.
log_extracarrieserror_codeanderror_kindas strings anderror_detailsas a JSON array (not a string) with the same elements. The top-level keys are canonical.
Retryability¶
Retryability follows the code. UNAVAILABLE is retryable; so is
RESOURCE_EXHAUSTED when it carries RetryInfo. When RetryInfo is
present a retry waits at least that long. ABORTED means retry the whole
operation at a higher level (re-open the session, re-read and re-apply), not
this call. Everything else is final.
Clients expose this as an is_retryable check and do not retry RPC errors
automatically: as in gRPC, automatic retry is opt-in, because a method may
not be idempotent. (Transport-level retry of replay-safe HTTP responses — 429,
502, 503, 504 before any RPC body — is a separate mechanism and is unchanged.)
Tracebacks¶
A server has a setting that decides whether EXCEPTION batches carry
log_extra.traceback and with it frames, cause and context. It is on
by default, on every transport, and an operator may turn it off for the
whole server (every transport it answers on — not per transport). The
exception type, message, code, kind and details are sent either way.
The default is "include" everywhere because at least one client puts the
remote traceback into the error a user sees: the DuckDB extension does, and an
earlier draft of this section that omitted tracebacks on HTTP and TCP hid
chained causes from its users. An operator who does not want stack traces to
leave the process — they name files, functions and sometimes values, which is
why gRPC keeps DebugInfo out of production responses — turns the setting
off.
A port without exception stacks still sends one. When the setting is on,
log_extra.traceback MUST be a non-empty string. A language that has no
stack trace for an error (C++, Rust) sends a synthesized trace: at minimum
<ErrorType>: <message> and the <protocol>/<method> that raised it, one per
line, plus whatever context the port has (a chained cause, a backtrace when
one is captured). Conformance asserts non-empty only; the shape is for humans.
frames may be an empty array when there are no frames to report.
Test vectors for all of the above — JSON forms, the cap boundary, the rules —
are in tools/cross-port/specs/MULTI_PROTOCOL_HOSTING.md.
log_extra JSON structure for EXCEPTION¶
{
"exception_type": "ValueError",
"exception_message": "invalid input",
"error_code": "INVALID_ARGUMENT",
"error_kind": "widget_invalid",
"error_details": [
{"@type": "vgi_rpc.BadRequest", "field_violations": [{"field": "size", "description": "must be positive"}]}
],
"traceback": "Traceback (most recent call last):\n ...",
"frames": [
{
"file": "/path/to/module.py",
"line": 42,
"function": "my_method",
"code": "raise ValueError('invalid input')"
}
],
"cause": "Traceback ... (optional, from __cause__)",
"context": "Traceback ... (optional, from __context__)"
}
| Field | Type | Description |
|---|---|---|
exception_type |
string | Exception class name. |
exception_message |
string | str(exception). |
error_code |
string | Mirror of vgi_rpc.error_code. |
error_kind |
string (optional) | Mirror of vgi_rpc.error_kind. |
error_details |
array (optional) | Mirror of vgi_rpc.error_details, as a JSON array. Absent whenever the top-level key is. |
traceback |
string (optional) | Formatted traceback. Truncated at 16,000 characters with "\n… <traceback truncated>" suffix. Absent when the server omits tracebacks; frames, cause and context are then absent too. |
frames |
array of objects | Last 5 stack frames (most recent at end). |
frames[].file |
string | Source file path. |
frames[].line |
integer | Line number. |
frames[].function |
string | Function/method name. |
frames[].code |
string or null | Source code at that line. |
cause |
string (optional) | Formatted __cause__ traceback. Truncated at 16,000 chars. |
context |
string (optional) | Formatted __context__ traceback (only when not suppressed). Truncated at 16,000 chars. |
Non-exception log batches¶
For levels other than EXCEPTION, the log_extra JSON structure is
freeform — it contains whatever key-value pairs the server method attached.
Clients should deliver these to the on_log callback without attempting to
parse them as error structures.
9. Stream Protocol (Pipe / Subprocess Transport)¶
Streaming methods use a multi-phase exchange over a bidirectional byte stream (two pipes: client→server and server→client).
Phase 1: Request parameters¶
Identical to a unary request (Section 5):
Phase 1.5: Optional header stream¶
Only present when the stream method declares a header type. Sent by the server immediately after reading the request, before the main data exchange.
The header is a single-row batch containing serialized header data. If the method does not declare a header type, this phase is skipped entirely.
The header batch is externalizable like any other batch. A server MAY
upload it when it exceeds the server's externalization threshold, in which case
the header stream carries a zero-row pointer batch in its place, bearing
vgi_rpc.location as described under External storage.
The obligation is deliberately asymmetric: externalizing a header is optional for servers and resolving one is mandatory for clients. A server that externalizes only in its data path is conformant, and most are. A client that cannot read an externalized header is not, whether or not its own server ever produces one. Clients MUST resolve pointer batches in the header stream through the same resolution path they use for the data stream, and MUST test for a pointer before classifying a zero-row batch as a log or control batch. A pointer is zero-row by construction, so a reader whose header loop treats zero rows as "log, skip" discards the header and then reports it absent — the header does not fail to parse, it fails to exist.
This is the one resolution path an implementation cannot exercise against itself unless its own server externalizes headers, and most do not: they write the header batch directly and externalize only in the data path. A port whose header reader has never seen a pointer has dead code there, not working code, and MUST verify it against a peer that externalizes headers.
If the server encounters an error during method initialization, it writes an
error stream (with EXCEPTION-level metadata) on the empty schema in place
of the header stream. The empty schema is used because the header schema may
not be available when the error occurs (e.g., the method raised before
returning a Stream object). Clients MUST be prepared to receive an
empty-schema error stream where a header stream was expected.
Phase 2: Lockstep data exchange¶
Both directions use a single long-lived IPC stream each:
Client → Server: IPC stream (input_schema, batch₁, batch₂, ..., EOS)
Server → Client: IPC stream (output_schema, [log*+data]₁, [log*+data]₂, ..., EOS)
The exchange is lockstep: the client writes one input batch, then reads the server's response (zero or more log batches followed by exactly one data batch). This repeats until termination.
Producer streams¶
- Input schema: empty (
pa.schema([])) — the client sends zero-row "tick" batches as timing signals. - Output: the server produces one data batch per tick.
- Termination: the server signals completion by not writing a data batch
after the final log batches — the output IPC stream reaches EOS. The
client detects this as
StopIteration. - Client-initiated close: the client closes its input IPC stream (writes EOS). The server detects this as end of input and stops producing.
Client Server
| |
|--- tick (0-row, empty) ------->|
|<------ log* + data batch₁ ----|
|--- tick (0-row, empty) ------->|
|<------ log* + data batch₂ ----|
|--- tick (0-row, empty) ------->|
|<------ log* + [EOS] ----------| (server called finish())
|--- [EOS] --------------------->|
Exchange streams¶
- Input schema: a real schema matching the exchange input type.
- Output: the server produces one data batch per input batch.
- Termination: the client closes its input stream (EOS). The server drains remaining input and closes the output stream.
Client Server
| |
|--- input batch₁ ------------->|
|<------ log* + output batch₁ --|
|--- input batch₂ ------------->|
|<------ log* + output batch₂ --|
|--- [EOS] --------------------->|
|<------ [EOS] -----------------|
Client-initiated cancellation¶
Either stream kind may be terminated early by the client with an explicit cancel batch, distinct from simply closing the input stream:
- num_rows: 0
- Schema: the stream's input schema (the empty schema for producers).
- Custom metadata:
vgi_rpc.cancel— the reference sets the value"1", but presence of the key is the signal; a reader MUST NOT depend on the value.
On receipt the server MUST NOT invoke process() / produce() for that turn.
It invokes the stream state's optional on_cancel hook — a failure in the hook
is swallowed, never surfaced to the client — and then ends the stream cleanly
by closing the output stream. A cancel is not an error: no EXCEPTION batch is
written, and the exchange terminates normally.
Cancellation is best-effort on the client side: transport errors while sending
the cancel are swallowed, since the session is being abandoned regardless. After
cancelling, the session is closed — subsequent exchange() / tick() calls
fail locally.
The same key applies over HTTP, where it rides on the /exchange request batch
alongside the state tokens (Section 10). Because a
cancel is not a method dispatch, the server suppresses dispatch hooks for it;
the access log still records the call, marked cancelled.
Error during streaming¶
If the server encounters an error during process(), it writes an
EXCEPTION-level log batch on the output stream, then the output stream
reaches EOS. The client reads the error batch, raises RpcError, and
the session is closed.
10. HTTP Transport¶
The HTTP transport maps the pipe-based protocol to stateless HTTP request/response pairs. Streaming state is serialized into signed tokens passed between exchanges.
Content type¶
All requests and responses use:
The server MUST reject requests with any other Content-Type with HTTP 415 (Unsupported Media Type).
Endpoints¶
Given a configurable URL prefix (default /vgi):
RPC surface — Arrow IPC in, Arrow IPC out:
| Endpoint | HTTP Method | Description |
|---|---|---|
{prefix}/{protocol}/{method} |
POST | Unary RPC call |
{prefix}/{protocol}/{method}/init |
POST | Stream initialization (producer and exchange) |
{prefix}/{protocol}/{method}/exchange |
POST | Stream continuation / exchange / cancel |
{prefix}/__upload_url__/init |
POST | Upload URL generation (only when an upload-URL provider is configured) |
RPC paths are namespaced by protocol; see Section 3.1
for the name grammar, the % ban, and the rule that the path must agree with
vgi_rpc.protocol. Introspection is not a special path — it is
{prefix}/vgi_rpc.Reflection.v1/list_protocols and
{prefix}/vgi_rpc.Reflection.v1/describe, reached the same way as anything
else. Reserved framework endpoints below are not namespaced: they belong to
the server rather than to any one protocol.
A GET to a two-segment path whose first segment cannot be a protocol name is
404, not 405. Without that rule any unrelated two-segment path — a
/.well-known/... document among them — matches the RPC route and is answered
"method not allowed" by it.
Framework endpoints — not Arrow IPC:
| Endpoint | HTTP Method | Description |
|---|---|---|
{prefix}/health |
GET, HEAD, OPTIONS | Health check + capability discovery. JSON body on GET; capability headers on all three. |
{prefix}/__session__ |
DELETE | Sticky-session teardown (only when sticky sessions are enabled). See Section 17. |
Optional, human- and IdP-facing — present by default in the reference but
carrying no wire contract; a port may omit them entirely:
GET {prefix} (landing page), GET {prefix}/describe (HTML introspection
page), /.well-known/oauth-protected-resource (OAuth resource metadata), and
{prefix}/_oauth/{callback,logout,token} (OAuth PKCE browser flow).
Capability discovery¶
Capability headers are stamped on every response, so a client that has
already made a call needs no separate probe. The dedicated discovery target is
{prefix}/health, because it is present in every implementation and exempt
from authentication. HEAD is the probe to use: GET, HEAD and OPTIONS
all carry the same headers, and HEAD is the only one of the three that works
from every client. There is no Arrow IPC body on any of them.
A browser client MUST probe with
HEAD, neverOPTIONS. The probe carriesVGI-Accept-Max-Response-Bytes, and a custom request header makes the request non-simple, so the browser preflights it.
HEADsurvives that preflight everywhere, becauseGET,HEADandPOSTare CORS-safelisted methods: they pass the preflight's method check even when absent from the server'sAccess-Control-Allow-Methods. A probe withHEADtherefore needs nothing from a server's CORS configuration, and servers advertising onlyPOST, OPTIONSanswer it correctly today.
OPTIONSis not safelisted, so it must be named inAccess-Control-Allow-Methodsto be permitted — and a server that derives that list from the route's own responders namesGETandHEAD, notOPTIONS. The probe is then refused before it is sent, and the failure surfaces as a CORS error naming a method the client author never wrote, on the first request of the connection. The Python reference does exactly this, so anOPTIONS-probing browser client cannot talk to it at all.The reference C++ client has always probed with
HEAD. The TypeScript client probed withOPTIONSuntil@query-farm/vgi-rpc0.25.3, and every browser consumer of it had to wrapfetchto rewrite the method.Note for implementors migrating from an earlier draft of this document: the discovery endpoint is
{prefix}/health, not{prefix}/__capabilities__. The reference client has never probed the latter.
| Header | Type | Emitted | Description |
|---|---|---|---|
VGI-Max-Request-Bytes |
Integer | when configured | Maximum request body size the server accepts inline. Exceeding it is 413 (see Section 13). |
VGI-Max-Response-Bytes |
Integer | when configured | Decoded Arrow IPC body cap, before HTTP content coding. Hard for every response shape. |
VGI-Accept-Max-Response-Bytes-Support |
"true" |
always | Server honors the strict client response-limit request header. |
VGI-Max-Externalized-Response-Bytes |
Integer | when configured | Cap on total bytes uploaded to external storage during one response. Always hard. |
VGI-Externalization-Enabled |
"true" / "false" |
always | Whether a storage backend is wired up, i.e. whether the client should expect pointer batches at all. |
VGI-Supported-Encodings |
Comma-separated codec tokens | always | Content codings this server will produce. See Content-encoding negotiation. |
VGI-Upload-URL-Support |
"true" |
when enabled | The upload-URL endpoint is available. |
VGI-Max-Upload-Bytes |
Integer | when enabled + configured | Maximum upload size for externalized batches. |
VGI-Proxy-Proof-Required |
"true" |
when required | This worker rejects requests lacking a valid proxy proof. |
VGI-Sticky-Enabled |
"true" |
when enabled | Sticky sessions are available. |
VGI-Sticky-Default-TTL |
Integer seconds | when sticky enabled | TTL applied when a method opens a session without specifying one. |
VGI-Sticky-Echo-Headers |
Comma-separated header names | when configured | Headers the client must replay for the life of a session. |
VGI-Token-Introspection |
"true" |
when enabled | The token-introspection route is live. |
Clients request a per-response hard limit with
VGI-Accept-Max-Response-Bytes. Its normalized field value is one ASCII
positive decimal integer from 65536 through 2^53-1; combined duplicates,
commas, signs, leading zeroes, non-ASCII digits, smaller values, and larger values are a
400 invalid request before method dispatch. The effective limit is the minimum
of that value and the configured application and hosting response caps. See
HTTP response budgets for worker-visible fields and
the error envelope.
A capability header that is absent means "not configured / not supported",
with one deliberate exception: an absent VGI-Supported-Encodings means a
server predating the header, for which a client assumes {zstd}. A
present but empty value is that server positively stating it speaks no
compression — the two are not interchangeable.
The server MAY include Cache-Control: max-age=N on the discovery response;
a client that honours it should refresh on expiry.
Upload URL generation¶
When the server has an upload_url_provider configured, the
POST {prefix}/__upload_url__/init endpoint generates pre-signed
upload/download URL pairs for client-side externalization.
Request: Standard unary request with vgi_rpc.method = "__upload_url__".
| Parameter | Arrow type | Default | Description |
|---|---|---|---|
count |
int64 |
1 | Number of URL pairs to generate (1–100). |
Response schema:
| Column | Arrow type | Nullable | Description |
|---|---|---|---|
upload_url |
utf8 |
No | Pre-signed URL for uploading batch data. |
download_url |
utf8 |
No | Pre-signed URL the server uses to fetch the uploaded data. |
expires_at |
timestamp("us", tz="UTC") |
No | Expiration time of the pre-signed URLs. |
The response has one row per requested URL pair.
Request headers¶
| Header | Description |
|---|---|
Content-Type |
MUST be application/vnd.apache.arrow.stream |
X-Request-ID |
Optional. Correlation ID echoed on response. If absent, server generates one. |
Content-Encoding |
Optional. Coding applied to the request body. An unsupported coding is 415. |
Accept-Encoding |
Optional. Codings the client accepts on the response. |
X-VGI-Accept-Encoding |
Optional. Same, but takes precedence — see Content-encoding negotiation. |
VGI-Proxy-Proof |
Optional. Per-request HMAC proof that the request arrived through a trusted proxy. See Proxy Proof. |
VGI-Session-Accept |
Optional. "true" opts the client in to sticky sessions. See Section 17. |
VGI-Session |
Optional. Resumes an existing sticky session. |
Response headers¶
Every response carries the capability headers from Capability discovery above. In addition:
| Header | Emitted | Description |
|---|---|---|
X-Request-ID |
always | Echoed or generated request correlation ID. |
X-VGI-RPC-Error |
on server-side errors | "true" marks a 200 response whose Arrow IPC body carries an EXCEPTION batch. See Section 13. |
Content-Encoding |
when the response body is compressed | The coding applied. |
X-VGI-Content-Encoding |
instead of the above | Used when the client negotiated via X-VGI-Accept-Encoding. |
VGI-Auth-Reason |
on 401 only |
Machine-readable reason code from the closed set in docs/unauthorized-spec.md. |
VGI-Auth-Proxy-Required |
on 401 only, when applicable |
"true" when this service's auth depends on headers a reverse proxy must inject. Derived from server configuration, so it is identical on every 401 and discloses nothing about the individual request. |
VGI-Session |
when a session was opened | The token the client echoes on subsequent requests. |
VGI-Session-Close |
when a session was closed | "true" tells the client to drop its captured token. |
VGI-Echo-<name> |
on a session-opening response, when configured | Instructs the client to send <name>: <value> on every subsequent request in the session. |
Servers that enable CORS expose WWW-Authenticate, X-Request-ID,
X-VGI-Content-Encoding, X-VGI-RPC-Error, VGI-Auth-Reason, and every
advertised capability header, so a browser client can read them cross-origin.
A server need not name GET or HEAD in Access-Control-Allow-Methods for
the capability probe (Section 10) to reach it: both are
CORS-safelisted and pass a preflight regardless. OPTIONS is not, which is why
the probe is HEAD.
Content-encoding negotiation¶
Request and response compression are independent: a server may decode a compressed request while producing only uncompressed responses.
The portable CI conformance profile is intentionally stricter than that
deployment-level flexibility: its primary HTTP worker MUST advertise and
implement both zstd and gzip for requests and responses. A dedicated
compression-disabled fixture verifies the valid empty-advertisement mode.
Codec tokens are the usual HTTP ones — zstd, gzip, and identity.
identity is the no-op transform, not a compressor: it exists so a client can
explicitly ask for an uncompressed response, which is otherwise only
reachable by accident when nothing it offers happens to be producible. It is
deliberately excluded from VGI-Supported-Encodings, since every
implementation can always do it and advertising it carries no information.
Requests. A client may compress the request body and name the coding in
Content-Encoding. A server that does not support the named coding MUST
answer 415, not fall through to identity — the body would otherwise reach
the Arrow reader as garbage. A body that names a supported coding but fails to
decompress is 400. VGI-Max-Request-Bytes, when advertised, applies both to
the encoded HTTP body and independently to the decoded body passed to Arrow;
either limit is a 413. Implementations MUST enforce the decoded limit while
decompressing (or from a trustworthy frame-size precheck) so a compressed body
cannot force an allocation larger than the advertised cap.
Responses. The client offers codings in Accept-Encoding and/or
X-VGI-Accept-Encoding. The server picks the first offered coding it can
produce, honouring client preference order, with X-VGI-Accept-Encoding
taking precedence over the generic header — both in choosing the codec and in
deciding which response header to stamp. If the first match is identity, the
server MUST honour that and send an uncompressed body rather than continuing
down the list. No overlap means an uncompressed body.
The custom header exists because general-purpose HTTP clients inject their own
Accept-Encoding (frequently listing gzip before zstd) that a caller
cannot suppress, which silently overrides the order vgi-rpc states. The
difference is not cosmetic — for large Arrow bodies gzip compression measured
roughly an order of magnitude slower than zstd end-to-end.
The chosen coding is stamped on Content-Encoding, or on
X-VGI-Content-Encoding when the client negotiated through the custom header.
Nothing is stamped for identity: an untransformed body is just a body.
Unary call (HTTP)¶
POST {prefix}/{protocol}/{method}
Request body: IPC stream (params_schema, 1 request row, EOS)
Response body: IPC stream (result_schema, 0..N log batches, 1 result/error batch, EOS)
HTTP 200: Success — and also server-side errors, which carry
X-VGI-RPC-Error: true plus an EXCEPTION batch in the body
HTTP 400: Protocol error (bad IPC, missing metadata, param validation failure)
HTTP 401: Authentication failure (JSON envelope or HTML page, NOT Arrow IPC)
HTTP 404: Unknown method
HTTP 413: Request body exceeds VGI-Max-Request-Bytes
HTTP 415: Wrong Content-Type, or an unsupported Content-Encoding
See Section 13 for the full mapping, including why
a server implementation error surfaces as 200 rather than 500.
The method name in the URL path MUST match the vgi_rpc.method value in
the request batch's custom metadata. A mismatch is a 400 error.
Stream initialization (HTTP)¶
The response depends on whether the stream is a producer or exchange stream:
Producer stream init response¶
The /init request also drives the producer's first lock-step turn. The
response body contains the output of that single process() invocation:
Response body:
[IPC stream: header_schema, 0..N log batches, 1 header row, EOS] (if header declared)
[IPC stream: output_schema, 0..N log batches interleaved with
0..1 data batch, then 0..1 continuation batch, EOS]
If the invocation does not finish the producer, the server appends a
continuation batch: a zero-row batch with vgi_rpc.stream_state#b64 in
its custom metadata. The client then follows up with an /exchange request
carrying that token, which drives exactly one more turn.
The effective response limit and server-preferred target are exposed to the
producer. They do not change the number of invocations in the response. A
single emitted batch that remains oversize after ordinary externalization is
replaced by a cursor-free ResponseTooLargeError. The companion cap
max_externalized_response_bytes governs external-channel uploads
independently and is also hard.
Exchange stream init response¶
Response body:
[IPC stream: header_schema, 0..N log batches, 1 header row, EOS] (if header declared)
[IPC stream: output_schema, 0..N log batches, 1 zero-row batch with state token, EOS]
The zero-row batch carries the signed state token in
vgi_rpc.stream_state#b64 custom metadata.
A server MUST split a stream's state in two, by lifetime. The half that is
fixed for the life of the call — the init request and the resolved
input/output schemas — is sealed into a separate call token under
vgi_rpc.call_state#b64, carried on this same zero-row batch;
vgi_rpc.stream_state#b64 then holds only the per-turn cursor.
The split is required, not an optimisation. Packing both halves into one token makes every continuation re-serialize, re-seal, re-open and re-parse a payload that cannot have changed — for a typical stream, the overwhelming majority of the token. Splitting also lets each half be compressed on its own lifetime, and lets a server cache the resolved call so a warm process skips opening the call token altogether. A server that keeps everything in the cursor forces every peer and intermediary onto the expensive path permanently, and is not conformant.
The call token is minted once, by /init, and is never re-issued — a
continuation response carries only a cursor. /init MUST emit both keys on
the same zero-row sentinel, for producer and exchange streams alike, and MUST
do so even when the method declares no call state of its own: the token still
carries the frozen schemas and the call id that binds the pair.
Stream exchange (HTTP)¶
POST {prefix}/{protocol}/{method}/exchange
Request body: IPC stream (input_schema, 1 input batch with state token in metadata, EOS)
Response body: IPC stream (output_schema, 0..N log batches, 1 data batch with updated state token, EOS)
The request batch's custom metadata MUST contain vgi_rpc.stream_state#b64
with the current state token (base64-encoded).
The request MUST also echo vgi_rpc.call_state#b64, unchanged, on every
subsequent request — continuations, exchanges, and cancels alike. A server
may resolve the call from a per-process cache, so a client that omits the
token still works while that cache is warm; it fails as soon as the cache is
not — a restarted worker, an evicted entry, or a request balanced onto a node
that never saw the /init. Since the cursor names a call the server minted
but no longer carries its payload, the client's copy is the only one such a
node has. A server MUST reject a continuation it cannot resolve with
400 Bad Request rather than guessing.
The same obligation extends to intermediaries: a proxy that forwards a continuation must carry both tokens, not just the cursor.
Resolution order (normative)¶
A server MUST resolve the two tokens in this order:
- Open and authenticate the cursor token first. Its AEAD tag covers the
call_id; its AAD covers the caller's(domain, principal). - Only then use that now-authenticated
call_idto look up any cached resolved call. - On a cache miss, open the client-supplied call token and require its
embedded
call_idto equal the one the cursor named.
The ordering is a security property, not an implementation detail. A client
cannot name a call_id the server did not mint for it, so a cache hit can
never hand back another principal's call state. A server that opens the
client-supplied call token first, or that keys a cache on any value the
client controls directly, has a cross-principal disclosure bug even though
every functional test still passes.
For producer continuation, the input is a zero-row batch on empty schema
with the state token. The server invokes process() exactly once. The
response may contain at most one data batch and, if the producer is not
finished, ends with another continuation token.
That zero-row input batch is a tick, and its custom metadata is the
tick's metadata. A server MUST deliver it to the FIRST process() call of
that turn, exactly as the pipe transport delivers the metadata riding a real
tick batch — with vgi_rpc.stream_state#b64, vgi_rpc.call_state#b64 and
vgi_rpc.cancel stripped first, since those are transport bookkeeping that
never appears on a pipe tick and the state token is a sealed cursor that
application code must not be able to read. Each HTTP producer request carries
exactly one tick and drives exactly one process() call.
This is what carries between-tick updates: a client that revises its
per-tick metadata mid-stream — DuckDB tightens vgi_pushdown_filters on each
tick as a Top-N boundary moves or a join-key IN set resolves — has no other
channel to reach the producer. A server that builds every continuation tick
from a constant empty batch strands those updates silently: the stream still
completes, the rows are still correct, and the worker simply keeps filtering
on the /init snapshot for the life of the stream. VGI-Max-Response-Bytes
is a hard per-turn limit; it never permits a server to consume more than one
tick or invoke process() more than once in a request. An oversize producer
response is replaced with a cursor-free ResponseTooLargeError envelope.
For exchange, the input carries real data plus the state token. The response data batch carries an updated state token for the next exchange.
To cancel, the client sends a zero-row input batch carrying
vgi_rpc.cancel alongside both tokens (Section 9).
The server ends the stream without dispatching the method.
The client MUST strip vgi_rpc.stream_state#b64 and vgi_rpc.call_state#b64
from the batch metadata before exposing it to application code.
State token binary format¶
Both tokens are opaque AEAD-sealed blobs, base64-encoded for UTF-8 safe metadata storage. The envelope is XChaCha20-Poly1305 (libsodium IETF variant) — confidential (state is not visible to anything between client and server) and authenticated (any tampering, including cross-principal replay, fails decryption).
The two carry independent version lines, because they change for independent reasons. After base64-decoding, both share this envelope:
Offset Size Field
0 1 version: uint8 (cursor token: 5; call token: 1)
1 24 nonce: random per-token (XChaCha20-Poly1305 NPUB)
25 ... ciphertext: AEAD(sealed_payload, AAD) — includes the
16-byte Poly1305 tag at the end
Compression happens inside the seal (normative)¶
sealed_payload is not the framed plaintext directly. It is:
Offset Size Field
0 1 codec: uint8 — self-describing codec tag
(reference: 0x00 raw, 0x01 zstd)
1 ... the framed plaintext below, compressed per `codec`
The order matters and is the point of the exercise: compress, then encrypt. Once a token is sealed it is ciphertext, so the HTTP body codec can no longer find any redundancy in it — measured, zstd over a sealed token recovers only the slack base64 added (to ~76–80%) and never the state's own structure. Compressing inside the seal reaches the real redundancy: a 7,800-byte call state packs to 1,872, turning a 10,820-byte token into 2,552.
This is the second half of what the lifetime split buys. Splitting lets each half be compressed and cached on its own lifetime; a server that packs everything into one token pays full freight on every turn even if it compresses.
The requirements on a server are:
- It MUST compress the payload inside the seal whenever the compressed form is smaller than the raw one. Compressing outside the seal is not an alternative — it accomplishes nothing.
- It MUST prefix a self-describing codec tag, so the reader never guesses, and MUST emit the raw tag and skip compression when compression does not pay. A small token must never grow.
- It MUST bound the decompressed size (the reference caps output at 64 MiB). The payload is authenticated before it is decompressed, so this guards against a framework bug rather than an attacker, but an unbounded decompress on a request path is not worth having.
- It MUST reject an unknown codec tag, or a payload that fails to
decompress, as the same uniform
400 Bad Requestas every other token failure — both mean a token this server did not mint.
The codec is the port's choice: zstd where the runtime has it, deflate or gzip where it does not. Tag values are per-port, since a token never round-trips across ports. The reference uses zstd level 3 for both token kinds — the same speed as level 1 at these payload sizes and slightly smaller, where the levels that compress materially better (9, 19) cost 8× and 84× the CPU for a few hundred bytes.
Because token internals are opaque by design, none of this is observable from outside and the shared conformance suite cannot check it. Each port should assert it in a language-local test over its own seal/open path — that compression engages on a large payload, that a tiny payload stays raw, and that a corrupt or unknown-codec payload surfaces as a 400.
Cursor token plaintext (encrypted, never on the wire) — v5 carries only the advancing state plus the call id binding it to its call token. Everything the pre-split v4 token also carried (both schemas, the stream id) moved into the call token:
Offset Size Field
0 8 created_at: uint64 LE (seconds since Unix epoch)
8 16 call_id: the call token this cursor belongs to
24 4 state_len: uint32 LE
28 N state_bytes: serialized StreamState
Call token plaintext — v1, minted once at /init:
Offset Size Field
0 8 created_at: uint64 LE (seconds since Unix epoch)
8 16 call_id: random, minted at /init
24 4 call_len: uint32 LE
28 N call_state_bytes: serialized call state (empty when the
method declares none)
... 4 type_len: uint32 LE
... M call_state_type: UTF-8 class name, or empty
... 4 schema_len: uint32 LE
... P schema_bytes: serialized output pa.Schema
... 4 input_schema_len: uint32 LE
... Q input_schema_bytes: serialized input pa.Schema
... 4 stream_id_len: uint32 LE
... R stream_id_bytes: UTF-8 chain-correlation id
call_state_type is carried because a stream method may return a union of
state classes whose members declare different call-state types. A reader MUST
resolve that name against the set the method itself declares, and reject an
unrecognised one — never look up a class by a client-supplied name.
The AEAD AAD (authenticated, not encrypted) is:
cursor token: b"vgi_rpc.state.v4\x00" || identity_tail
call token: b"vgi_rpc.call.v1\x00" || identity_tail
where identity_tail is b"\x01" || domain || b"\x00" || principal for
authenticated requests and the literal b"\x00anonymous" otherwise. The
prefixes differ deliberately, so a call token and a cursor token are not
interchangeable even for the same principal: presenting one where the other
is expected fails the tag check rather than decoding into a payload the
reader would misinterpret.
As elsewhere, the plaintext framing above is the reference implementation's;
a port may choose its own encoding for what goes inside each token, and its
own compression codec (see the porting guide). What is normative is that
there are two tokens, split by lifetime, bound by an authenticated call_id,
and that each is compressed inside its seal under a self-describing codec
tag.
This binds every token to its issuing identity: a token sealed for one
principal cannot be opened with another principal's AAD even if both
share the master token_key.
version: Format selector. Not part of the AAD; mismatches are rejected before doing crypto work, but tampering with the byte to point at a different algorithm still fails the subsequent decrypt because the server uses the format-fixed algorithm constants.created_at: Token creation time as seconds since the Unix epoch. Used by the server to enforce a configurable TTL (token_ttl). Whentoken_ttl > 0, tokens older thantoken_ttlseconds are rejected with HTTP 400 ("State token expired"). Settoken_ttlto0to disable expiry checking. The default TTL is 3600 seconds (1 hour).state_bytes: The stream state dataclass serialized as a complete Arrow IPC stream (schema + 1-row batch + EOS).schema_bytes: The output Arrow schema serialized viapa.Schema.serialize().input_schema_bytes: The input Arrow schema serialized viapa.Schema.serialize(). For producer streams, this is the serialized empty schema.
Verification order: The version byte is checked first (cheap
rejection), then the AEAD decrypt is attempted. Any authenticity
failure (bad key, bad AAD, tampered nonce or ciphertext) surfaces as a
uniform HTTP 400 "State token signature verification failed". TTL
enforcement only runs after authenticity is established — the timestamp
is inside the ciphertext, so it cannot be tampered with independently.
Authentication (HTTP)¶
When the server has an authenticate callback configured:
- The callback receives the HTTP request and returns an
AuthContext. - On failure (
ValueErrororPermissionError), the server returns HTTP 401. The body is NOT Arrow IPC — no method has been resolved yet, so no output schema is available. Its shape is the standardized envelope ofdocs/unauthorized-spec.md: a JSON object carrying areasoncode from a closed set, mirrored on aVGI-Auth-Reasonheader, or the styled HTML page when the request'sAcceptasks fortext/html. - Other exceptions from the callback propagate as HTTP 500.
- Clients MUST detect 401 responses before attempting to parse Arrow IPC.
Proxy proof (optional)¶
A worker may additionally require that a request arrived through a trusted proxy. The proxy mints a
per-request HMAC-SHA256 proof in a VGI-Proxy-Proof header; the worker verifies it against a shared
per-worker secret.
- It is a precondition ANDed with the
authenticatecallback above, never an alternative credential — the caller'sAuthorizationheader is untouched and still carries the end user. - Failure maps to the same HTTP 401 as any other authenticate failure, carrying the
proxy_requiredreason code. The body MUST NOT echo the verifier's reason or the claimed key id — every proof outcome collapses onto that one code. OPTIONS,/.well-known/, and{prefix}/healthare exempt in all modes, so load-balancer probes and capability discovery keep working.- A worker requiring proofs advertises
VGI-Proxy-Proof-Required: trueon every response. - Opt-in: an unconfigured worker reads no header, emits none, and is byte-identical to a worker built before the feature existed.
The token format, canonical MAC input, verifier algorithm, reason codes, and rotation procedure are normative in the Proxy Proof Specification.
11. Shared Memory (SHM) Transport¶
The shared memory side-channel enables zero-copy batch transfer between co-located processes. It is used alongside a pipe transport — the pipe carries control messages and small batches; large batches are written to shared memory and replaced with pointer batches on the pipe.
Negotiation prerequisite. Over the pipe / subprocess / AF-UNIX transports, SHM MUST be negotiated via
__transport_options__(Section 15) before use: a client only advertises a segment (and writes SHM pointer batches) to a server that has confirmedvgi_rpc.transport.shm = "true". A server that cannot do SHM (non-POSIX host, missing runtime support, or a server predating the method) reports"false"(or errors), and the client falls back to the inline pipe transport. HTTP servers advertise their deployment capabilities on every response; the dedicated discovery probe isHEAD {prefix}/health(Section 10).
Segment header format¶
The shared memory segment begins with a 64 KiB (65,536 byte) header, followed by a data region. All integers are little-endian.
Offset Size Field
0 4 magic: bytes "VGIS" (0x56 0x47 0x49 0x53)
4 4 version: uint32 = 1
8 8 data_size: uint64 (segment size minus 65536)
16 4 num_allocs: uint32 (number of active allocations)
20 4 padding: uint32 = 0
24 N*16 allocations: array of (offset: uint64, length: uint64)
sorted by offset, where N = num_allocs
- Maximum allocations:
(65536 - 24) / 16 = 4094. - Offsets are absolute — measured from the start of the shared memory segment (not from the data region start).
- Data region starts at byte offset 65,536 (immediately after the header).
Allocation strategy¶
The allocator uses a first-fit strategy with implicit coalescing:
- Scan the sorted allocation list for the first gap that fits the requested size.
- Gaps are computed as: before the first allocation (from offset 65536), between consecutive allocations, and after the last allocation (to segment end).
- New allocations are inserted to maintain sorted order.
- Freeing an allocation removes its entry; adjacent free space coalesces implicitly since only occupied regions are tracked.
Batch serialization in SHM¶
For non-dictionary-encoded batches: A complete Arrow IPC stream (schema + record batch + EOS) is written directly into the allocated SHM region.
For dictionary-encoded batches: The IPC stream is written to a temporary buffer, then the schema message and EOS marker are stripped — only the dictionary messages and record batch message are stored in SHM.
Dictionary batch reconstruction (reader side):
To deserialize a dictionary-encoded batch from SHM:
- Serialize the pointer batch's schema into a schema message by creating a temporary IPC stream writer (which emits a schema message + EOS), then strip the trailing 8-byte EOS marker. This yields the schema message bytes.
- Concatenate:
schema_message_bytes+shm_stored_bytes+EOS_marker(8 bytes:0xFF 0xFF 0xFF 0xFF 0x00 0x00 0x00 0x00). - Open the concatenated buffer as a standard Arrow IPC stream and read the batch.
Non-dictionary batches do not need this reconstruction — they are stored as complete IPC streams and can be read directly.
SHM pointer batch¶
A batch stored in shared memory is replaced on the pipe with a pointer batch:
- num_rows: 0
- Schema: Same as the original batch's schema.
- Custom metadata:
vgi_rpc.shm_offset: Absolute byte offset in the segment (decimal string).vgi_rpc.shm_length: Number of bytes written (decimal string).
SHM segment identity in request metadata¶
When a client owns a shared memory segment, it advertises the segment in the request batch's custom metadata:
vgi_rpc.shm_segment_name: OS name of the shared memory segment.vgi_rpc.shm_segment_size: Total segment size in bytes (decimal string).
The server dynamically attaches to the segment (read-only, untracked by the resource tracker) for the duration of the request.
Resolution algorithm¶
resolve_shm_batch(batch, custom_metadata, shm_segment):
IF shm_segment is NULL:
return (batch, custom_metadata, null)
IF batch.num_rows != 0:
return (batch, custom_metadata, null)
IF custom_metadata is NULL:
return (batch, custom_metadata, null)
IF "vgi_rpc.shm_offset" NOT IN custom_metadata:
return (batch, custom_metadata, null)
IF "vgi_rpc.log_level" IN custom_metadata:
return (batch, custom_metadata, null) // log batch, not pointer
offset = int(custom_metadata["vgi_rpc.shm_offset"])
length = int(custom_metadata["vgi_rpc.shm_length"])
buffer = shm_segment.read(offset, length)
resolved_batch = deserialize_ipc_stream(buffer, batch.schema)
// Strip pointer keys, add provenance
resolved_metadata = remove_keys(custom_metadata, "vgi_rpc.shm_offset", "vgi_rpc.shm_length")
resolved_metadata["vgi_rpc.shm_source"] = shm_segment.name
release_fn = () => shm_segment.free(offset)
return (resolved_batch, resolved_metadata, release_fn)
12. External Storage Pointer Batches¶
When batches exceed a configurable size threshold, they can be externalized to remote storage (e.g., S3, GCS) and replaced with pointer batches.
Pointer batch format¶
- num_rows: 0
- Schema: Same as the original batch's schema.
- Custom metadata:
vgi_rpc.location: URL to fetch the batch data (typically a pre-signed URL).vgi_rpc.location.sha256: Optional. SHA-256 hex digest of the payload before compression.- Must NOT contain
vgi_rpc.log_level(to distinguish from log batches).
Externalization (writing)¶
When a data batch's total buffer size exceeds the threshold:
- Serialize all batches from the current output cycle (log batches + data batch) as a single IPC stream.
- Compute the SHA-256 of those bytes, before any compression.
- Optionally compress with zstd.
- Upload to external storage via the
ExternalStorage.upload()interface. - Replace the entire cycle with a single zero-row pointer batch containing
vgi_rpc.location(andvgi_rpc.location.sha256when the digest was computed).
Pre-published references¶
A server MAY answer a unary call with a pointer to an object it published
earlier (Python: a method returning an ExternalRef built by
publish_external). The wire is unchanged: one zero-row batch with the
method's result schema carrying vgi_rpc.location, and
vgi_rpc.location.sha256 only when the publisher recorded a digest. The
object is exactly what the externalizer would have uploaded for that result —
an Arrow IPC stream of the result schema holding one 1-row data batch,
optionally Content-Encoding-compressed. Such a pointer is written regardless
of the server's externalization threshold or storage configuration, is never
inlined or routed through shared memory, and uploads nothing during the call.
Readers resolve it like any other pointer.
Integrity¶
vgi_rpc.location.sha256 is optional on the wire — a pointer batch without it
is valid, which keeps readers compatible with writers that predate the key.
When the key is present, a reader MUST verify the digest against the fetched
payload (decompressed, if a coding was applied) and MUST fail the resolution
on mismatch rather than handing the batch to application code.
Fetch safety¶
Readers that resolve external pointers MUST apply their configured URL policy to the initial URL and to every redirect target before issuing that hop. A reader MUST bound redirect following; rejecting redirects entirely is also valid. This prevents an allowed public URL from redirecting the reader to a loopback or private service.
The encoded bytes received from storage and the decoded bytes passed to the
Arrow reader have independent operator-configured limits. The encoded limit
MUST be enforced while streaming the response, even if Content-Length is
absent or false. The decoded limit MUST be enforced after applying the named
Content-Encoding. Implementations MAY keep their historical decoded limit
as the default when adding the separate setting.
Diagnostic errors, logs, traces, and exception chains MUST NOT expose URL
userinfo, query strings, or fragments. In particular, signed query parameters
are bearer credentials. Diagnostic URLs retain only scheme, host, port, and
path. This applies to rendering a URL for a human, not to the metadata keys
themselves: the vgi_rpc.location pointer keeps its URL unchanged as
application metadata, and vgi_rpc.location.source — which is not wire
content, but stamped by the reader from the URL it fetched — likewise records
that URL in full rather than its redacted diagnostic form.
Resolution (reading)¶
resolve_external_location(batch, custom_metadata, config):
IF config is NULL:
return (batch, custom_metadata)
IF NOT is_external_pointer(batch, custom_metadata):
return (batch, custom_metadata)
url = custom_metadata["vgi_rpc.location"]
// Fetch with retries
// Default: max 3 total attempts (max_retries=2, capped at 2)
// Retry delay: 0.5s fixed between attempts
// Retryable errors: network/OS errors, Arrow parse errors, HTTP client errors
// Validate the initial URL and every redirect target before requesting it.
// Enforce independent encoded and decoded byte caps.
data = fetch_url(url, config.fetch_config, config.url_validator)
// Decompress if needed (zstd)
// Open as IPC stream, dispatch EVERY log batch, extract the one data batch.
// The object is a whole IPC stream, not a single batch: reading only its
// first batch, or stopping at the first data batch, discards the turn's logs.
reader = open_ipc_stream(data)
data_batch = NULL
FOR each batch in reader:
IF is_log_or_error(batch):
deliver to on_log callback
CONTINUE
IF has "vgi_rpc.location":
ERROR: redirect loop detected
IF data_batch is not NULL:
ERROR: multiple data batches
data_batch = batch
IF data_batch is NULL:
ERROR: no data batch in payload
// Validate schema match
IF data_batch.schema != expected_schema:
ERROR: schema mismatch
// The resolved metadata is the INNER data batch's, never the pointer's.
resolved_metadata = data_batch.custom_metadata
// Add fetch provenance metadata
resolved_metadata["vgi_rpc.location.fetch_ms"] = elapsed_ms
resolved_metadata["vgi_rpc.location.source"] = url
return (data_batch, resolved_metadata)
The three lines worth spelling out, because independent implementations have guessed all of them wrong:
- Provenance is stamped by the reader, at resolve time.
vgi_rpc.location.sourceMUST be the URL that was fetched, andvgi_rpc.location.fetch_msthe elapsed fetch time. Neither is written by the producer, and a pointer batch on the wire MUST NOT carry either. Two failures follow from getting this wrong, and both are invisible to a port testing against itself. A writer that stampslocation.sourceonto the pointer, paired with a reader that passes the pointer's metadata through, agrees with itself and disagrees with every correct peer — and the value it propagates is whatever the writer chose rather than a URL anyone fetched. A reader that stamps neither key returns a batch with no provenance at all, which against a correct peer is indistinguishable from a resolver that silently did not run. - Resolved metadata is the fetched data batch's metadata, merged with the
provenance keys above — it is not the pointer batch's. The pointer
carries only
vgi_rpc.locationandvgi_rpc.location.sha256; everything the writer attached to the data batch is inside the fetched object. A reader that returns the pointer's metadata loses whatever rode on the data batch, and on HTTP that includes the continuation cursor (vgi_rpc.stream_state#b64) — so the first turn succeeds and the second cannot be addressed. A resolved batch MUST NOT carryvgi_rpc.locationorvgi_rpc.location.sha256. - Every log batch in the payload MUST be dispatched. The uploaded object holds the turn's log batches followed by its data batch, so a reader that returns at the first data batch drops logs that were delivered inline before externalization was enabled — a regression no data assertion can see.
Stream externalization¶
For stream methods, the server may externalize an entire output cycle (log batches + data batch) as one IPC stream. The pointer batch replaces the entire cycle. On resolution, the client reads back all batches, dispatches log batches, and returns the data batch.
Transport independence¶
Externalization is not an HTTP feature. Any transport that carries record
batches carries pointer batches, and the persistent byte-stream transports
(pipe, subprocess, Unix socket, TCP) resolve them through the same reader
path. A port whose pointer resolution lives only in its HTTP client has an
untested half; the shared TestExternalByteStream conformance group exists to
find it (see
cross-language-conformance.md).
13. Version Negotiation & Error Handling¶
Three mechanisms, three failures¶
Three independent things can be skewed between a client and a worker, and each has its own mechanism. None substitutes for another, and a port that implements two of the three has a silent failure mode:
| Skew | Caught by | Answer |
|---|---|---|
| Incompatible major version | routing — the major is part of the protocol name, so this is a different protocol | protocol_not_supported, HTTP 404 |
| Method absent on the server | method resolution within the named protocol | method_not_implemented, HTTP 404 |
| Signature changed under an unchanged method name | the protocol_version gate |
ProtocolVersionError |
The third row is the one nothing else covers, and it is why the gate exists despite no mainstream RPC framework having one. That advice reasons by analogy to protobuf's forgiving evolution: field numbers and unknown-field semantics make a renamed or retyped field self-detecting. Arrow has neither. A parameter renamed between releases, or retyped under the same name, produces a request the server will happily coerce through its own declared type and act on. The gate is what turns that into a directional error.
Two supporting rules make the gate meaningful rather than decorative:
- A server MUST reject a request carrying a parameter its protocol does not declare, naming the declared set. A strict version gate sitting on a lenient deserializer is two policies in one codepath.
- The gate is per binding. A server hosting several protocols has a version
per protocol and no single "server version"; gating against the primary
rejects correct callers of a secondary and names the wrong protocol when it
does. The error message MUST name the protocol — with N bindings,
Server: 1.0.0alone does not say which server.
The reflection protocol is exempt from the gate. It is what a version-mismatched client calls to learn what mismatched, and gating it would deny the client the diagnosis it came for. The exemption is a property of the binding, not of a method name.
Version checking¶
Every request batch MUST carry vgi_rpc.request_version in its custom
metadata with the value "1".
| Condition | Error |
|---|---|
vgi_rpc.request_version missing |
VersionError — server writes an error stream on the empty schema. |
vgi_rpc.request_version != "1" |
VersionError — server writes an error stream on the empty schema. |
vgi_rpc.protocol_version missing or mismatched (when enforced) |
ProtocolVersionError — see below. |
vgi_rpc.method missing |
RpcError (ProtocolError) — server writes an error stream. |
| Unknown method name | RpcError (AttributeError) — error stream includes available method names. |
| Request batch has wrong row count (not 1, on non-empty schema) | RpcError (ProtocolError). |
| Non-optional parameter is null | TypeError — error stream on the method's result schema. |
Protocol version negotiation¶
vgi_rpc.request_version versions the framing in this document, and has
been "1" throughout. vgi_rpc.protocol_version is a second, independent
line: it versions the application's RPC surface — the set of methods, their
parameters, and their schemas — so a client and worker built against different
releases of a service fail with a directional message instead of a schema
error deep inside a call.
It is opt-in by declaration. A Protocol that declares a version (in the
Python reference, a protocol_version: ClassVar[str] on the Protocol class)
turns the check on for both peers; a Protocol that declares none disables it
entirely, and the key never appears on the wire.
- Format: canonical semver
MAJOR.MINOR.PATCH— non-negative integers, no leading zeros, no prereleases and no build metadata.1.0.0-rc1and1.0.0+build3are malformed, not merely unusual. - Client obligation: when the bound Protocol declares a version, the client MUST send it on every request batch.
- Server obligation: when its own Protocol declares a version, the server MUST check the client's at the dispatch boundary — before parameter deserialization, so a mismatch cannot be mistaken for a schema problem.
- Comparison rule: exact major and minor match. Patch is ignored, so a
1.4.0client and a1.4.9server interoperate. __describe__is exempt. It is the diagnostic path a version-mismatched client uses to discover what the server actually speaks, so gating it would make the mismatch undiagnosable.
Every failure — absent key, undecodable bytes, malformed semver, or a genuine
major/minor difference — raises ProtocolVersionError, a subclass of
VersionError, and is written as an ordinary error stream carrying
error_kind = "protocol_version_mismatch". The message MUST state both
versions and which side to upgrade; a bare "mismatch" leaves the reader to
guess, which is the whole failure this key exists to prevent.
The server also emits its protocol_version in the __describe__ response
metadata, so a client can read it without triggering a failure.
Relationship to
protocol_hash: the hash is a fingerprint of the protocol's decoded description, canonicalised as RFC 8785 JSON, so it is a cross-language contract. All seven implementations produce the same digest for the same protocol;tests/golden/protocol_hash_vector.jsonships the value and its preimage so a disagreeing port diffs JSON rather than guessing.This paragraph previously said the opposite -- that the hash fingerprinted the Python describe payload's bytes and was not comparable across runtimes. That was true of the old definition and is why the hash was worth little: a field comparable only against itself detects no drift between implementations. The definition changed; the disclaimer outlived it.
protocol_versionand the hash answer different questions. The version is a declared, gated attribute; the hash is derived from the surface itself, so it moves when the surface moves whether or not anyone remembered to bump anything.
Error stream format¶
Protocol-level errors are written as a complete IPC stream on the appropriate schema (empty schema for version/method errors, result schema for parameter validation errors):
IPC Stream (error):
Schema message (empty or result schema)
1 zero-row batch with EXCEPTION-level log metadata
EOS marker
HTTP status code mapping¶
Two distinct classes of failure are deliberately not conflated. A transport
or protocol failure — the request never became a valid RPC call — carries a
4xx status. An application failure — the method was dispatched and raised —
is reported in band, as a 200 whose Arrow IPC body carries an EXCEPTION
batch.
| Error condition | HTTP status |
|---|---|
| Bad IPC, missing metadata, request-version mismatch, param validation | 400 Bad Request |
protocol_version mismatch |
400 Bad Request |
| Expired, tampered, or unresolvable state token | 400 Bad Request |
| Request body fails to decompress | 400 Bad Request |
Path and vgi_rpc.protocol name different protocols |
400 Bad Request |
| Authentication failure (including proxy proof) | 401 Unauthorized |
| Unknown method on a hosted protocol | 404 Not Found |
| Protocol not hosted, or a path segment that cannot be a protocol name | 404 Not Found |
% anywhere in the protocol path segment |
404 Not Found |
| GET to a two-segment path that is not a protocol route | 404 Not Found |
Request body exceeds VGI-Max-Request-Bytes |
413 Payload Too Large |
Wrong Content-Type, or unsupported Content-Encoding |
415 Unsupported Media Type |
| Any error raised by the method implementation | 200 OK + X-VGI-RPC-Error: true |
| Response overshoots a hard response cap | 200 OK + X-VGI-RPC-Error: true |
Why implementation errors are 200¶
A server implementation error never reaches the client as 500. The server
translates it to 200 and marks it with X-VGI-RPC-Error: true, because
intermediaries and HTTP client libraries routinely discard or replace response
bodies on 5xx — and the body is precisely where the typed error lives. A 500
would strip the exception type, message, traceback, and error_kind and leave
the caller with a bare status code.
Clients MUST therefore treat 200 as "a response arrived", not "the call
succeeded", and classify by inspecting the body — the batch-classification
algorithm in Section 7 already does this,
so X-VGI-RPC-Error is a fast path and a diagnostic aid, not a second source
of truth. A client that branches only on status code will silently treat
failures as successes.
This also means the status code no longer varies with the class of exception
the method raised: a TypeError from inside a method body is 200 like any
other. Caller-supplied shape errors are validated before dispatch, and so
remain 400.
Note: For 400 and 413 responses the body is still a valid Arrow IPC stream containing an error batch. The exceptions are 401 (JSON or HTML, per
docs/unauthorized-spec.md), 415 (framework default response), and the non-Arrow framework endpoints in Section 16 and Section 17.
14. Introspection (vgi_rpc.Reflection.v1)¶
Introspection is an ordinary co-hosted protocol, not a special method name. That is what lets every port generate it from the same pipeline as any other method rather than hand-maintain a bespoke format — which is how the six ports drifted before.
A server that offers introspection hosts vgi_rpc.Reflection.v1 and it appears
in that protocol's own output. A server that does not simply does not host it,
and a client asking gets the ordinary protocol_not_supported — not a bespoke
"introspection is disabled" to special-case.
Methods¶
list_protocols is the cheap question — what is here, and has it changed — and
is the only one a client needs on a warm path, because protocol_hash answers
"has it changed" without transferring any schema. describe is the expensive
one, asked once. A client that does not already know a protocol name needs both,
in that order.
Both are ordinary unary methods: they carry vgi_rpc.protocol =
"vgi_rpc.Reflection.v1", their replies ride as serialized bytes in a single
result column, and a server that externalizes payloads externalizes these too.
Payload¶
ProtocolList:
| Field | Type | Notes |
|---|---|---|
server_id |
utf8 |
Server instance identifier. |
server_version |
utf8 |
Build version string. |
request_version |
utf8 |
The framing version from Section 3. |
protocols |
list<ProtocolSummary> |
Every hosted protocol, reflection included. |
ProtocolSummary:
| Field | Type | Notes |
|---|---|---|
protocol |
utf8 |
Wire name — the routing key. |
protocol_version |
utf8 |
Declared semver, or "" when the protocol opts out. |
protocol_hash |
utf8 |
64 lowercase hex. See below. |
deprecated |
bool |
Whether callers should migrate off. |
deprecation_message |
utf8 |
What to migrate to. Empty unless deprecated. |
features |
list<utf8> |
Reserved. MUST be emitted as an empty list in this version; clients MUST ignore its contents. |
ServiceDescription is ProtocolSummary's fields plus methods:
list<MethodInfo>, sorted by name. It deliberately carries no server
identity: two processes serving one protocol must describe it identically, or
the description is not a property of the protocol. Server identity lives on
ProtocolList, which is a statement about a server.
MethodInfo:
| Field | Type | Notes |
|---|---|---|
name |
utf8 |
|
method_type |
utf8 |
"unary" or "stream". |
has_return |
bool |
|
has_header |
bool |
|
stream_kind |
utf8 |
"" for unary; otherwise unknown / producer / exchange. |
params_schema_ipc |
binary |
Arrow IPC. |
result_schema_ipc |
binary |
Arrow IPC; empty when has_return is false. |
header_schema_ipc |
binary |
Arrow IPC; empty when has_header is false. |
idempotency |
utf8 |
unknown / no_side_effects / idempotent. |
deprecated |
bool |
|
deprecation_message |
utf8 |
Schemas travel as serialized Arrow IPC rather than as a structural description: a client's whole purpose in asking is to get a schema it can hand to its own Arrow implementation, and IPC is the one representation every port already reads. The hash is what compares across ports, and it is defined over the decoded structure precisely so these bytes need not match.
stream_kind is a string rather than a nullable bool because the state is
genuinely three-valued. A port that can determine a stream's kind MUST state
it: this is the only field in a description that says whether a stream accepts
input — the description carries parameter, result and header schemas, but a
stream's input schema arrives at init time, so a client, a code generator or a
human reading a describe page has exactly one place to learn whether a method is
send-and-receive or receive-only.
unknown is for the methods a port genuinely cannot classify: one whose
producer-vs-exchange shape is decided from the returned stream rather than
declared, or a registration whose output schema is computed at run time and does
not carry the shape. It is a real answer, not a default to fall back on — "I
cannot say" is different from "producer", and a port that reports unknown for
a method it could have classified has made its description useless for the
question it exists to answer.
An absent schema is empty bytes rather than null, so no port pays a null check on a value it will only ever treat as absent.
features is reserved rather than open because the protocol is the unit of
optionality (Section 3.1): a
capability that may be absent is hosted as its own protocol, which reflection
already reports, rather than advertised as a token whose meaning every port
would have to agree on. A later version may define tokens; until it does a
server sends [] and a client that reads anything else ignores it, so a
future token cannot change how a current client behaves.
idempotency follows gRPC's idempotency_level. With an HTTP transport and a
policy proxy in the path, retries will happen; without this nothing on the
wire says what is safe to retry. unknown is the default and means a caller
must assume the worst.
Decoding is tolerant — normative¶
A decoder MUST:
- read fields by name, not by position;
- ignore columns it does not know;
- default columns that are absent and have a default.
This is what makes minor skew survivable in both directions, and it is load bearing: a strict decoder fails at exactly the moment a client most needs a good answer, which is when it is talking to a server it does not fully understand. Conformance asserts it against a deliberately extended schema.
A decoder MUST NOT default a field that has no default — it errors instead. Silently zero-filling a required field hands a client a description that is wrong rather than absent. One rule follows, and it binds every port: a field added in a minor version MUST carry a default, or the addition is a breaking change wearing a minor version number.
protocol_hash — normative¶
The hash is a fingerprint of a protocol's wire surface, so a client and a
worker can say which one they have without transferring the whole description.
It is defined over what Arrow decodes to, not what an encoder emits: each
language's Arrow implementation may legitimately produce different bytes for the
same logical schema, so a hash over schema.serialize() is comparable only
against itself.
Canonicalisation profile: RFC 8785 (JCS), chosen for its published test vectors. The structure is restricted to objects, arrays, strings and booleans. Every numeric parameter is folded into a type token (below), so the preimage contains no numbers and JCS's number-canonicalisation rule — its hardest and the likeliest place for six ports to diverge — never applies. Keep it that way.
The preimage:
{"protocol":"vgi_rpc.Identity.v1","methods":[
{"name":"introspect_token","type":"unary","has_return":true,
"has_header":false,
"params":[{"name":"token","nullable":false,"type":"utf8"}],
"result":[{"name":"result","nullable":false,"type":"binary"}]}]}
- Methods are sorted by name; a port iterating a hash map must still produce this order.
- Field order within a schema is declaration order and is significant.
resultandheaderare omitted when the method has none. Absent and empty must not hash alike: a method returning nothing is not a method returning an empty struct.- Server identity, docstrings, parameter defaults, language-specific type names,
and
request_versionare not in the preimage. They vary across processes, builds and ports without changing what is on the wire, and folding the framing version in would rotate every protocol's hash on a framework release that changed no protocol. stream_kindis not in the preimage either, and for a different reason: not that no port can determine it — every port can, for most methods — but that which methods a port can classify depends on how that port's registration works. So two ports can disagree about a method while neither is wrong, and a field one port can state and another cannot is not a contract. It still reaches clients on the description, whereunknownis a sayable answer; a hash has no such option.- The
v1in the domain tag is the only version the hash carries, and it moves only when the hash definition moves.
Conformance asserts that every port produces the identical hash for the conformance service, and ships the canonical preimage as a test vector beside the digest so a failing port diffs JSON rather than guessing.
Type tokens¶
A type appears in the preimage as a lowercase ASCII token. Parameters go in
parentheses; children in angle brackets. A child is name:token when
non-nullable and name?:token when nullable.
| Arrow type | Token |
|---|---|
| null, bool | null, bool |
| signed ints | int8 int16 int32 int64 |
| unsigned ints | uint8 uint16 uint32 uint64 |
| floats | float16 float32 float64 |
| decimal | decimal128(p,s), decimal256(p,s) |
| string | utf8, large_utf8, utf8_view |
| binary | binary, large_binary, binary_view, fixed_size_binary(n) |
| date | date32, date64 |
| time | time32(unit), time64(unit) |
| timestamp | timestamp(unit), timestamp(unit,tz=Z) |
| duration | duration(unit) |
| interval | interval_months, interval_day_time, interval_month_day_nano |
| list | list<item?:T>, large_list<…>, list_view<…>, large_list_view<…>, fixed_size_list(n)<…> |
| struct | struct<a:T,b?:U> |
| map | map<key:K,value?:V>, with ,keys_sorted appended when set |
| union | dense_union<code=name?:T,…>, sparse_union<…> |
| dictionary | dictionary<index:T,value:U>, with ,ordered appended when set |
| run-end encoded | run_end_encoded<run_ends:T,values:U> |
| extension | extension(name)<storage> |
unit is Arrow's own spelling: s, ms, us, ns. A timezone is carried
verbatim — UTC and +00:00 are distinct Arrow types and must not
collapse.
What is normalised, and why. Arrow's own type equality ignores the name of
a list's child field and of a map's key/value fields. pyarrow names the list
child item; some Parquet producers name it element. Those names are
normalised to item / key / value, because keeping them would give two
ports different hashes for a protocol Arrow itself calls identical — exactly the
divergence the canonical preimage exists to prevent. Everything Arrow does
treat as part of the type is kept: child nullability, struct field names, union
child names and type codes, dictionary index/value types and orderedness, and
keys_sorted.
A type with no token in this table MUST raise rather than fall back to the Arrow
implementation's own to_string, whose output differs between ports and across
Arrow releases. Adding a token is a change every port makes at once: a one-sided
addition changes only that port's hash.
15. Transport Capability Negotiation (__transport_options__)¶
__transport_options__ is a built-in synthetic unary method (parallel to
__describe__) through which a client and server discover each other's
transport capabilities — chiefly whether the shared-memory side-channel
(Section 11) may be used. HTTP deployment capabilities use response headers;
the dedicated discovery probe is HEAD {prefix}/health (Section 10).
It is mandatory before SHM is used: a client that has SHM available MUST NOT
write SHM pointer batches (or advertise a segment) to a server unless that server
has confirmed SHM support via this method. A server that cannot attach SHM (e.g.
a non-POSIX host, or a runtime without the required FFM support) reports
shm = "false" and the client falls back to inline transport. A server that does
not implement the method at all returns a method_not_implemented error (or any
error), which the client treats as "no SHM".
Capabilities are negotiated once per worker and may be cached for the life of the worker process (they are process-level, not per-connection), so there is no per-call overhead.
Request¶
Standard unary request with:
- vgi_rpc.method = "__transport_options__"
- Empty params schema (zero fields, one row)
- The client's own capabilities as request metadata under the
vgi_rpc.transport.* namespace (e.g. vgi_rpc.transport.shm = "true")
Response¶
An IPC stream with an empty batch (zero fields). Capabilities ride as the
response batch's custom_metadata under the vgi_rpc.transport.* namespace:
| Key | Value | Description |
|---|---|---|
vgi_rpc.transport.shm |
"true" / "false" |
Whether the server can use the SHM side-channel |
vgi_rpc.server_id |
UTF-8 | Server instance identifier |
vgi_rpc.request_version |
"1" |
Wire protocol version |
The capability set is open-ended: keys are matched by the vgi_rpc.transport.
prefix and unknown keys are ignored, so future capabilities (e.g. compression,
AEAD) can be added without a protocol-version bump. A feature is used only when
both peers advertise it.
Negotiation rule¶
shm_enabled = client.advertises("vgi_rpc.transport.shm" == "true")
AND server.advertises("vgi_rpc.transport.shm" == "true")
A server that has not attached a segment but still receives an inbound SHM pointer batch (a negotiation violation) MUST fail loudly rather than silently treat the zero-row pointer as empty input.
16. Identity (vgi_rpc.Identity.v1)¶
Two methods, optional and independently so, hosted as an ordinary co-hosted
protocol. It was previously an HTTP JSON route, POST
{prefix}/__introspect_token__, which meant it existed on one transport only and
had to be hand-written in every port.
introspect_token(token: utf8) -> TokenIdentity
issue_grant(purpose: utf8, scopes: list<utf8>, ttl_seconds: int64) -> IssuedGrant
A method whose hook the deployment did not configure is not hosted, and the
binding's method set — and therefore its protocol_hash — narrows accordingly.
A worker that resolves credentials but does not mint grants hosts
introspect_token alone, and a client learns that from reflection rather than
by calling and reading an error. Absent beats routed-and-refusing: it is what
keeps a dependency upgrade from growing a credential-to-identity oracle on every
existing worker.
introspect_token — an oracle, and guarded as one¶
Resolves an opaque bearer credential to the principal it authenticates as, for a reverse proxy that terminates the only public listener and must know the caller's identity before it can authorize anything.
The answer is an identity assertion made by the thing being protected, which the asker then acts on using credentials the worker does not hold — storage credentials, entitlement lookups, policy-tier selection. "Trust it as much as you trust the worker" is the wrong frame: it must be trusted more. Hence four guards, all normative:
- An introspector allowlist with no permissive default. A server whose resolver is configured without one MUST refuse to start. Authentication and introspection are different capabilities: "any authenticated caller" lets any user test guesses of any other user's credential at unlimited rate, and resolve a stolen one to its owner.
- Authorization is checked before the credential is touched, so an unauthorized caller learns nothing about it — including how long looking at it took.
- Rejections are uniform. Unknown, expired, malformed and over-long are one answer. Distinguishing them confirms that a guessed credential exists.
- A JWS-shaped subject is refused before the resolver runs. Three dot-separated base64url segments are validated locally against a key set; routing one onward hands a third party a token the asker may itself have rejected for being expired or wrong-audience.
introspect_token is not rate limited. The allowlist is the control. An
earlier revision also capped each caller at 20 introspections a second, to bound
what an allowlisted-but-compromised caller could do. It bounded the wrong thing
and cost the right one:
- It bounded only guessing, which a random credential defeats at any rate, and not the harm a leaked introspector credential actually does -- resolving a stolen credential to its owner takes one call.
- The caller is the asker, and the asker introspects on behalf of every client that presents a bearer. A per-caller limit was therefore one budget for every user's first login, and unauthenticated clients drained it simply by sending the asker junk credentials.
Throttle untrusted traffic where it arrives -- at the asker, per client. An
implementation or deployment that throttles introspection anyway MUST answer
with a transient kind (identity_unavailable), never introspection_refused:
that kind is definitive and MAY be cached, so a throttled answer reported as it
negative-caches valid credentials, which is how the cap used to lock users out.
TokenIdentity carries principal, token_name and ttl_seconds, and
never claims. A pass-through claims field would let a worker choose its
caller's tenant routing, row scope and policy branch; the asker derives what it
needs from the principal alone. ttl_seconds is how long the answer may be
cached — treat it as an authorization window, and therefore as the revocation
lag.
issue_grant — not an oracle, so not guarded as one¶
Mints a standing delegation credential for the calling user. OAuth cannot express durable delegation: it fuses the grant, the credential and the session into one refresh token, so an IdP shortening session lifetime shortens the grant. This is the durable record — minted while the user is present, presented later by unattended automation as an ordinary bearer.
There is no subject parameter. The subject is always the caller's
authenticated principal, so cross-subject minting is closed by construction
rather than by a check one of six ports can forget. That is also why this method
needs no allowlist while introspect_token has one.
A credential with no verifiable auth_time cannot mint. That single rule is
what stops a grant being used to mint another grant and escaping the identity
provider permanently — a grant is not IdP-issued, so it carries no auth_time —
and it makes subprocess and unix transports fail closed for free, since there is
no authenticated principal there at all. A static bearer proves a machine holds
a secret, never that a human just authenticated, so it is refused here too.
auth_timeis an OIDC claim meaning when this session began, which can be arbitrarily old while still present and cryptographically valid. Requiring it is not the same as requiring a recent login: the deployment must sendmax_age(or an appropriateacr) at the authorize endpoint for this guard to mean what it says.
Unlike the introspection rejections, these are deliberately actionable: they are always about the caller themselves, so naming the reason leaks nothing, and it is the only way a console learns to re-prompt.
IssuedGrant carries token, expires_at and grant_id. The token format
is the worker's entirely — a sealed envelope, a database row, or a credential
brokered from the IdP are equally valid and equally invisible here. It is never
parsed and never logged. expires_at is required because the framework cannot
enforce it: the real lifetime lives inside the opaque token, so this is a
declaration, and a worker that must state a lifetime has thought about one.
Accepting identity credentials¶
A grant is minted to be presented later as an ordinary bearer, and
resolve_token resolves opaque bearers — so both MUST feed back into HTTP
authentication. The byte-level contract is IDENTITY_V1_SPEC.md §9 (with
vectors in vgi_rpc/conformance/grant_token_vectors.json); in summary:
- Sealed grants, opt-in. When a grant key is configured
(
VGI_RPC_GRANT_KEYS, base64 of exactly 32 bytes each, first mints / all verify;--grant-key), the framework providesmint_grantunless the worker supplies one, and accepts its own grants as bearers. Not configured, nothing changes. A malformed key stops the worker at startup. - Format.
"vgig1." base64url_nopad(kid(8) ‖ envelope), the envelope being the state-token XChaCha20-Poly1305 envelope (version0x01), AAD"vgi_rpc.grant.v1"‖0x00‖kid‖audience, payload a fixed little-endian binary record ofissued_at, expires_at, grant_id, principal, purpose, scopes. Lifetime ≤ the configured maximum; 60 s clock skew. - AuthContext.
domain "grant", the grant's principal, claims{grant_id, scopes, purpose}and noauth_time— soissue_grantrefuses a grant-authenticated caller (stale_auth): grants never mint grants. resolve_tokenas a bearer authenticator. Consulted for bearers the earlier authenticators did not accept: an identity authenticates (domain "token");Nonefalls through (401 if nothing accepts); an unavailable error is 503 withRetry-After, never 401. Never consulted for avgig1.token, a JWS-shaped one, or one over 4096 bytes.- Order. The deployment's authenticators (JWT, static) → sealed grants
(prefix check; a non-
vgig1.bearer never reaches the verifier, and avgig1.bearer that fails is 401 and stops the chain) →resolve_token. - Errors. A bad, tampered, wrong-key or expired grant is a 401 on the auth
path (
invalid_credential, orexpired_credentialonce authentic) — canonical codeUNAUTHENTICATED. - Revocation. Sealed grants are not individually revocable: keep the maximum lifetime short and re-issue; removing a key revokes all its grants.
Error kinds¶
These were an HTTP route whose callers classified definitive-versus-transient on
the status code (404 vs 503). As protocol methods every handler exception
surfaces the same way, so error_kind carries the whole distinction and is
load-bearing rather than decorative:
error_kind |
Code | Meaning | Caller |
|---|---|---|---|
introspection_refused |
PERMISSION_DENIED |
The caller may not introspect -- it is not on the allowlist. Never a throttle. | Definitive; MAY cache. |
token_unresolved |
NOT_FOUND |
The subject credential did not resolve. | Definitive; MAY cache. |
stale_auth |
UNAUTHENTICATED |
The caller has not authenticated recently enough to mint. | Definitive, and actionable — re-prompt. |
grant_refused |
PERMISSION_DENIED |
The worker declined to mint. | Definitive. |
identity_unavailable |
UNAVAILABLE |
The answer is not knowable — a store is down. | Transient; MUST NOT negative-cache. |
identity_unavailable MUST carry vgi_rpc.RetryInfo with the delay the
failing component asked for. Every port had a retry_after on its
identity-unavailable error and none put it on the wire, so a caller learned
the failure was transient and then had to guess when to ask again.
A hook that raises the transport-auth "unavailable" error is translated, not
passed through. A resolve_token or mint_grant hook typically calls the
same backing store an authenticator does, and so raises what an authenticator
raises when that store is down — the port's "authentication unavailable" error
(AuthUnavailableError in the reference; the 503-with-Retry-After signal of
unauthorized-spec.md). The framework MUST emit that as
identity_unavailable carrying that error's own retry hint as RetryInfo,
never unclassified and never with a substituted default. Untranslated it
reaches the wire with no kind, and a caller can no longer tell an outage from a
refusal — the one distinction this table exists to carry. TypeScript and Rust
translated before the rule was written; the other five ports did not.
A caller that negative-caches a transient failure locks out valid users; one
that retries a definitive rejection hammers the worker. identity_unavailable
is deliberately not a ValueError in the reference, because
chain_authenticate advances to the next authenticator on ValueError — a
sidecar outage raised as one is read as "not my credential, try the next" and
ends up a 401 from the end of the chain, restarting every session in the fleet
over a thirty-second blip.
17. Sticky Sessions (HTTP, optional)¶
HTTP-only. Optional. Opt-in on both sides.
Sticky sessions let a method bind a handle-bearing object — an open database cursor, a loaded model, a file handle, an in-progress generation — to the worker process that created it, keyed by a short-lived AEAD-sealed token the client echoes on subsequent requests. The state lives in process memory; it is never serialized onto the wire.
The other transports are single-process, so sticky is meaningless there: a runtime that exposes the session API at all MUST raise on a non-HTTP transport rather than silently no-op. When neither side opts in, the wire is byte-identical to a framework built before the feature existed.
The full normative contract — token envelope, principal binding, TTL and
eviction, drain semantics, concurrency, and the TestSticky conformance
group — is docs/sticky-sessions-spec.md. What
follows is the wire surface only.
Request headers¶
| Header | Required | Purpose |
|---|---|---|
VGI-Session-Accept: true |
when a method may open a session | Client opt-in. A server MUST refuse to open a session for a request lacking it — otherwise it leaks sessions to clients that are not tracking them. |
VGI-Session: <token> |
when resuming | The token minted on a prior response. |
Response headers¶
| Header | Emitted | Purpose |
|---|---|---|
VGI-Session: <token> |
when a session was opened this request | Token for the client to echo. Base64url, no padding. |
VGI-Session-Close: true |
when the session was closed this request | Client drops its captured token and any echo headers. |
VGI-Echo-<name>: <value> |
once, on the session-opening response | Client MUST strip the VGI-Echo- prefix and send <name>: <value> on every subsequent request in the session. Used for client-driven routing on platforms that steer by header. |
Capability headers (VGI-Sticky-Enabled, VGI-Sticky-Default-TTL,
VGI-Sticky-Echo-Headers) are listed under
Capability discovery.
Teardown endpoint¶
Idempotent and best-effort. 204 No Content when the entry was found and
evicted; 200 OK on any failure — missing header, malformed token,
identity mismatch, registry miss. The two are deliberately not
distinguishable, so a stolen token cannot be used to probe whether a session
exists.
Failure surfacing¶
A token that cannot be honoured — expired, evicted, routed to a worker that
never saw it, or presented under a different principal — surfaces as an
ordinary EXCEPTION batch with error_kind = "session_lost". A session open
refused because the server is shutting down surfaces as server_draining.
Neither is retried transparently by the framework: the client is told, and
decides whether to reopen or fail.
Appendix A: IPC Stream EOS Marker¶
The end-of-stream marker is the 8-byte sequence:
This signals to the IPC stream reader that no more messages follow.
Appendix B: Empty Schema¶
The "empty schema" referenced throughout this specification is an Arrow
schema with zero fields: pa.schema([]). When serialized, it produces a
small fixed-size blob. Batches on the empty schema have zero columns.
Appendix C: Empty Schema Serialized Form¶
The empty schema (pa.schema([])) serializes to a fixed 56-byte blob via
pa.Schema.serialize(). This is useful for cross-language implementations
that need to produce or compare serialized empty schemas (e.g., for the
input_schema_bytes field in state tokens for producer streams):
ff ff ff ff 30 00 00 00 10 00 00 00 00 00 0a 00
0c 00 06 00 05 00 08 00 0a 00 00 00 00 01 04 00
0c 00 00 00 08 00 08 00 00 00 04 00 08 00 00 00
04 00 00 00 00 00 00 00
This is a Flatbuffers-encoded Arrow Schema message with zero fields. Implementations MAY hard-code this constant rather than generating it at runtime. The serialized form is stable across Arrow versions.
Appendix D: Request Metadata Location¶
A common implementation question: the vgi_rpc.method key appears in the
request batch's custom metadata (per-batch metadata), not in the
schema-level metadata. This is by design — schema-level metadata is part of
the IPC stream schema message and cannot vary between batches, while
custom metadata is per-batch and can carry request-specific values.
vgi_rpc.request_version is also in batch custom metadata.