Adapting an ASR system to Standard ASR (engine authors)
Authoritative reference:
docs/content/specification/protocol.md. Entry-point rules:plugin-entry-points.md.
You implement one class. The standard layer gives you audio-input negotiation, conversion, resampling, parameter gating, diagnostics, the CLI, the reference server, and the compliance suite — for free.
The contract
Subclass EngineBase and provide:
properties: ClassVar[BaseProperties]— static identity and I/O boundaries (accepted_input,native_sample_rate,accepted_sample_rates,selectable_languages, …).declared_capabilities: ClassVar[DeclaredCapabilities]— what you support, per mode (batch/streaming). Omit what you don't support (fail-closed).provider_params_type: ClassVar[type[ProviderParams] | None]— your typed escape-hatch model, orNone. Publish your own terminal subclass: the bareProviderParamsbase, a non-subclass, or a model that is not closed (extra="forbid") is a compliance error (provider_params_type_is_bare_base/_not_subclass/_not_closed).__init__— capture config only. Keep it pure: no filesystem, GPU, or network (spec IC.9). Load weights lazily in_ensure_model_loaded._transcribe(prepared, params) -> TranscriptionResult— run your model on already-negotiated audio (prepared.kindis one of youraccepted_input).- (Streaming) override
_start_transcription(*, gated_params, audio_format, prepared_audio)returning aTranscriptionSessionsubclass. config_type: ClassVar[type[BaseConfig]]— your config class. It is read from the class, so a settings UI can render the schema without constructing the engine. Without it,registry.config_schema()returnsNone,GET /v1/config-schema/{model}returns an empty schema,standard-asr showreports no init config, and compliance warns (missing_config_type).- If your
propertiesdeclareselectable_languages, your config MUST carry a usabledefault_language(spec IC.6). InheritLanguageConfigMixinto get the field. Without it, everytranscribe()raisesEngineContractErrorand compliance reports the errorlanguage_config_invalid. - If you set
effective_capabilities, it MUST narrowdeclared_capabilities, never widen it. Compliance reports a widening as the erroreffective_widens_declared. - If your engine loads weights, override
prepare()(spec IC.11). Keep it synchronous and zero-argument, and apply the sameallow_downloads()gate as transcription. Compliance rejects a coroutine or an argument-taking hook (prepare_hook_is_coroutine/prepare_hook_requires_args).
Minimal batch engine
from typing import ClassVar, Literal
from standard_asr.engine import (
BaseConfig,
BaseProperties,
BatchCapabilities,
DeclaredCapabilities,
EngineBase,
FlagCap,
InputKind,
LanguageCaps,
LanguageConfigMixin,
PreparedAudio,
RuntimeParams,
TranscriptionResult,
)
class MyConfig(LanguageConfigMixin, BaseConfig[Literal["my-engine"]]):
engine: Literal["my-engine"] = "my-engine"
default_language: str = "en" # IC.6: required once you declare a language axis
class MyProps(BaseProperties):
engine_id: str = "my-engine"
model_name: str = "base"
protocol_version: str = "1.0.0"
accepted_input: set[InputKind] = {InputKind.ARRAY}
native_sample_rate: int = 16000
accepted_sample_rates: list[int] = [16000]
selectable_languages: list[str] = ["en", "auto"]
detectable_languages: list[str] = ["en"]
class MyEngine(EngineBase):
properties: ClassVar[BaseProperties] = MyProps()
config_type: ClassVar[type[BaseConfig]] = MyConfig # schema without instantiation
declared_capabilities: ClassVar[DeclaredCapabilities] = DeclaredCapabilities(
batch=BatchCapabilities(
language=LanguageCaps(runtime_override=FlagCap(supported=True)),
)
)
def __init__(self, **kw: object) -> None:
self.config = MyConfig(**kw) # extra="forbid": a mistyped option fails loudly
self._model = None
def _transcribe(self, prepared: PreparedAudio, params: RuntimeParams) -> TranscriptionResult:
audio = prepared.array # 16 kHz float32 mono, per Properties
text = my_model_infer(audio) # your code
# Report what the recognizer actually decided. Never echo
# params.language: with "auto" selectable it can be the reserved
# "auto", which detected_language rejects.
return TranscriptionResult(text=text, detected_language="en")Map parameters
- The standard layer gates the portable standard set against your
effective_capabilities(which default todeclared_capabilities) before it calls_transcribe:language,candidate_languages,word_timestamps,diarization,prompt,phrase_hints. Map them onto your model's native arguments. Fordiarization, presence means enable: mapparams.diarization is not Noneonto your native enable switch (the v1DiarizationRequestmarker carries no fields). An engine that declaresdiarization.supported=TrueMUST actually diarize when the request passes the gate. The standard layer cannot verify this. Silently ignoring a gated-and-passed request is the cardinal sin: a silent wrong result. - Engine-specific parameters → a
ProviderParamssubclass set asprovider_params_type. Wrong-engine params raiseInvalidProviderParamError. - The base resolves the language axis for you on both paths: the
params.languageyour_transcribereceives is already the effective value. Only a structural engine that bypassesEngineBasecallsstandard_asr.contract.language.effective_language(...)itself. word_timestamps.granularitiesdeclares what you can honestly deliver, not which native API switch exists. Declare every granularity your engine can serve — including ones that come for free. If your model emits per-segment start/end on every run (most do), declare"segment"even when there is no separate "segment mode" switch. Otherwise the standard layer rejects the cheapest, always-satisfiable request as a false incompatibility. Then map each granularity precisely (for example, only"word"enables your forced-alignment pass; a"segment"request MUST NOT back-fill word-level data —words=Nonemeans "not requested").
Audio you receive
prepared is already in one of your accepted_input shapes:
InputKind.ARRAY→prepared.array(float32,prepared.sample_rate)InputKind.ENCODED_FILE→prepared.pathInputKind.ENCODED_BYTES→prepared.dataInputKind.FETCHABLE_URL→prepared.urlInputKind.STORAGE_URI→prepared.storage_uri(a provider storage URI such ass3://orgs://, for engines that read from cloud storage)
You never write decode/resample/encode glue — declare accepted_input and the
standard layer delivers the right shape (and attaches conversion diagnostics).
Streaming
Declare the transport axis first. Set streaming_input=FlagCap(supported=True)
if you accept incremental PCM frames, streaming_output=FlagCap(supported=True)
if you return results incrementally, or both. These are engine-global flags, and
either one may be supported only when you also declare a streaming domain. A
streaming domain with neither axis is a streaming engine nobody can call: every
start_transcription() raises UnsupportedFeatureError, and compliance reports
the error streaming_domain_without_axis.
Subclass TranscriptionSession, implement async _produce() (read fed audio via
self.audio_chunks(), yield TranscriptionEvent objects). The base provides
feed/send_audio/end_audio, backpressure, the done-timeout, and the sync
bridge — you only write async. See the spec §ST for the event model
(partial/final/supersede/progress/done/error) and the stable_until
rules.
Session establishment — the base does the gating for you
You override _start_transcription(...), not the public start_transcription.
The base start_transcription is a template method (symmetric to
transcribe / _transcribe): it runs the standard streaming pipeline and then
calls your hook. Before your hook runs, the base has already:
- enforced the
audio_format/audiomutual exclusion (ST §3.1) viaensure_stream_inputs_exclusive; - validated the language config (LANG R1 / IC.6);
- run the fail-closed wire-format check (
ensure_stream_format_supported) on the encoding, the channel count, and the sample rate. It rejects anaudio_format.encodingnot in your declaredwire_encodings, so an undeclared encoding is not misframed as PCM and silently mistranscribed. Declarewire_encodings: leaving it unset means "unconstrained", and the encoding check is then skipped, so an encoding you never declared reaches you unchecked. This is the one fail-open concession in the check; compliance warns about it (streaming_input_without_wire_encodings). It rejects anaudio_format.channelsother than 1: v1 streaming wire input is mono-only, because the standard layer does not downmix incremental frames the way the batch path does. Downmix to mono before feeding. It also rejects a wiresample_rateyou do not accept. Per spec R7's v1 note, the standard does not resample streaming wire frames in v1 (only the batchtranscribepath resamples), so an unreachable wire rate is a loud error. Whenrequired_input_sample_rateis set, the wire rate MUST equal it — even when another rate appears inaccepted_sample_rates. That list describes the batch path, which resamples to the required rate before your engine; unresampled wire frames at any other rate would be misread. Otherwise the standard accepts the rate whenaccepted_sample_ratesis"any"or when it is in that concrete list. (Standard-layer streaming resampling is a deferred capability; this guard becomes a resample once it lands.) - gated the runtime parameters against your
streamingcapabilities, and resolved the language axis. Gating coversprovider_paramsswap-safety (Runtime R3: a wrongprovider_paramstype always raisesInvalidProviderParamError), capability gating (R2), and guidance degradation (R4). The base attaches the gating and language diagnostics to the returned session; they surface throughsession.diagnostics(). - for the whole-input path (
audio=..., for example, OpenAI-style streaming output), run that complete input through the same audio negotiation/conversion pipeline as batchtranscribe, and hand your hook the result asprepared_audio. Theprepared_audiois aPreparedAudioalready in one of youraccepted_inputshapes, with its conversion diagnostics attached to the session. For the incrementalaudio_format=...path there is no whole input, soprepared_audioisNone.
Your hook receives the already-gated, frozen gated_params (spec R5: streaming
params are frozen at start_transcription and MUST NOT change mid-stream, except a
guidance channel declared mutable_mid_stream — a reserved declaration in v1). Use them
directly — do not re-gate or re-accept raw params. The signature is
keyword-only: gated_params, audio_format (the wire format, or None), and
prepared_audio (the negotiated whole input, or None).
def _start_transcription(self, *, gated_params, audio_format, prepared_audio):
# Guards + gating + (whole-input) audio prep already ran in the base.
# gated_params is frozen (R5); prepared_audio is None for the incremental
# audio_format path and a PreparedAudio for the whole-input audio path.
return MySession(gated_params, ...)Credentials & environment fallback (IC.4)
Build your config with Config.from_env(engine_id, **explicit) instead of the
bare constructor. Unset fields fall back to STANDARD_ASR_<ENGINE>__<FIELD>
environment variables. Note the double underscore that separates the engine
and field segments; explicit args win. Credentials are wrapped in their masking
carrier by construction, never passed around as plaintext. Put secrets
(api_key, tokens) in SecretStr fields — or SecretBytes for byte
credentials — via secret_field(). Declare exactly one carrier per field,
optionally with None. Keep non-secret routing (base_url, region) plain.
A structured field (a list, a mapping, a submodel, a TypedDict, a dataclass)
takes its env value as JSON ('["en","ja"]'). A scalar field — including
SecretStr and Path — takes the raw string, byte for byte. The field's own
schema decides which of the two applies, so any shape the config guards accept
is reachable through the env convention. One shape is refused at class
definition: a field accepting BOTH (str | list[str]) has no defined reading.
"123" is either that string or that JSON number, and either choice would
disagree with the explicit constructor, which always takes the string. Declare
one shape, or model the alternatives as a named submodel.
The config's serialization surface is closed: public_dump() emits your
declared input fields, serialized by pydantic's own machinery, and nothing
else. Definition-time guards hold that closure by enumeration, not proof
(the accident model — see the trust model in AGENTS.md): they walk your
declared annotations and your serialization decorators, at every nesting
depth, and reject the hazards an honest author actually writes. Anything
that would make model_dump run author code — or emit something other than
your declared inputs — is rejected at class definition: @computed_field,
@model_serializer, @field_serializer,
PlainSerializer/WrapSerializer metadata, a SerializeAsAny[...] field
(its dump follows the runtime object, so the declared type no longer
bounds the output), an undeclared value shape (Any, object, an
unparametrized container, dict[str, Any] — same duck-typing, reached from
the type instead of a marker; spell a heterogeneous mapping as
dict[str, str | int | bool | None] or a named submodel), a nested
submodel carrying any of those, and Field(exclude=True). An enumeration
has a boundary: a serializer installed through a custom
__get_pydantic_core_schema__ slips past it, and the guards do not chase
it — installing one means actively smuggling code past the enumeration,
which is an adversary, out of scope for a trusted plugin.
Keep extra="forbid" too (BaseConfig's default).
extra="allow" stores undeclared caller data and dumps it verbatim past the
secret mask. extra="ignore" silently swallows a mistyped credential key, so
it reads as an absent credential rather than a loud error. The reason:
public_dump() is documented safe for /v1/models, persistence, and
telemetry. That holds only while nothing author-defined can rematerialize a
credential inside it. Its output must also stay the declared input surface, so
persisting and reloading a config round-trips.
The input surface stays closed at every depth, not only on the config
itself. Every nested input container your schema reaches — an options submodel,
a TypedDict, a dataclass — must forbid undeclared keys, and one that does not
is rejected at class definition. pydantic's default for all three silently
drops an unknown key. So a user's typo'd nested option
({"decode": {"baem": 8}}) would read as applied while your engine runs on the
field's default — a silent wrong result. The rule reads the effective policy
from the core schema, so pydantic's config propagation is honored. A bare
TypedDict or stdlib dataclass inherits the config's extra="forbid" and is
closed for free. A nested BaseModel needs
model_config = ConfigDict(extra="forbid"). A pydantic dataclass (which owns
its config) needs
@pydantic.dataclasses.dataclass(config=ConfigDict(extra="forbid")).
Use a plain @property for derived in-process values (an authorization
header belongs in your engine code, not in the config dump), and keep config
fields to plain typed inputs.
One boundary is yours to keep, because no schema-level guard can hold it for
you: never copy a secret out of its carrier. The guards bound what the
schema installs in the dump, not the contents of values your own code
builds. A validator that writes get_secret_value() into a plain field or onto
an object's display state (say, a Path subclass whose __str__ embeds the
token) emits that plaintext through public_dump() — and would through any dump
mechanism. Read the credential with reveal_dump() at the point of use in your
engine code, and let it live nowhere else.
Provider-native wire names map onto standard fields with plain string
aliases (Field(alias="xi-api-key"), or an all-string AliasChoices).
AliasPath — and any AliasChoices carrying one — is rejected at class
definition. The flat env convention and the absent-vs-invalid config classifier
(what makes a missing credential a compliance skip instead of a fail) both
resolve fields by single string tokens, which a nested path alias cannot
provide. If a value is genuinely nested, declare it as a submodel field (its env
value arrives as JSON).
def __init__(self, **kwargs):
self.config = MyConfig.from_env("my-engine", **kwargs) # IC.4Wire-visible values: extra, diagnostics
Every slot the wire can see — a TranscriptionResult / Segment / Word /
TranscriptionEvent extra, and emit_diagnostic's provided / effective
— holds JSON values only (JsonValue: null, bool, int, finite float,
str, and lists/str-keyed dicts of those). The Python objects and the JSON
documents are the same protocol seen twice. So a value with no JSON form is
rejected at construction, naming the field, instead of failing later in the
transport — after the server has already committed to a response. Non-finite
floats (NaN, Infinity) are excluded for the same reason: they are Python
floats but not JSON, and a conforming parser rejects the whole document.
emit_diagnostic projects a structured value (a pydantic submodel) into
its JSON form itself, so provided=my_request_model just works. For any
other wire-visible slot — or to absorb the list-invariance complaint a
type checker raises when a list[str] variable meets a list[JsonValue]
parameter (a static-analysis artifact, not a real mismatch) — use
to_json_value from the engine surface:
from standard_asr.engine import to_json_value
hints: list[str] = [...]
event = TranscriptionEvent.final("s1", text, extra={"hints": to_json_value(hints)})Runtime validation is unchanged either way: a value that is genuinely not JSON is rejected loudly at construction, naming the field.
If you genuinely need an arbitrary in-process object, keep it in your own engine/session state: a standard protocol object's whole contract is that both layers can express it.
Streaming responsibilities (what the base does vs you)
The base TranscriptionSession owns the pump, backpressure (bounded buffers),
the done-timeout/idle deadlines, the sync bridge, lifecycle suppression
(strict_lifecycle=True to raise instead of diagnose), and stable_until
monotonicity clamping. You must: emit cumulative/replace text; set
stable_until conservatively (0 if you have no right-context); and for
reconnect, detect the disconnect, re-establish, replay self.replay_buffer(),
keep segment_id/timestamps/language continuous, and call
self.note_reconnect(gap_start, gap_end, content_lost=...). The base always
emits the progress(reconnect) event. It emits a trailing non-terminal
content_lost error (recoverable=true — a fidelity warning; the session
stays alive and events keep flowing) only if you pass content_lost=True.
That is your own determination that the reconnect and replay could not cover the
gap, and that unreplayable audio was permanently lost. The base does not
infer loss from
rolling-buffer eviction (a live ring is always evicting, so that would falsely
claim loss on every long session); you decide, because only you know whether the
replay actually bridged the gap.
error events fail closed to terminal. An error event with recoverable
unset is normalized to recoverable=false (terminal) at construction: unknown
recoverability must not leave consumers waiting on a stream that may never
continue. If you emit an advisory, non-fatal error (the session keeps going),
set recoverable=True explicitly — otherwise your event ends the session.
Surface non-fatal notes via emit_diagnostic. Call
self.emit_diagnostic(code=..., message=..., level="info"|"warning") from
_produce to report a best-effort degradation, an assumed parameter, or a lossy
fallback through the session's diagnostics() channel — the streaming
counterpart of the batch path's result.diagnostics. It is bounded (spec
ST.6.4, like the guard's own diagnostics) and the server forwards it to a WS
client as a mid-stream diagnostics frame. Keep error events for fatal
conditions.
Security: a diagnostic is engine-authored, client-facing output, like the transcript itself. Its
message/param/provided/effectivefields are forwarded to (possibly unauthenticated) clients verbatim and unredacted. Never put a credential, API key, authenticated URL, or raw exception text in a diagnostic — route sensitive operator detail tologginginstead. (The server does scrub anerrorevent'sextra, because that is auto-captured exception detail — pre-summarized input-echo-free by the standard layer, but still operator-only content; a diagnostic is content you chose, so its safety is yours.)
Sequence invariants the guard enforces for free
Beyond lifecycle transitions and monotonic stable_until, the base _LifecycleGuard
also enforces two further per-stream invariants on every event you yield, so a
slipped engine still cannot emit a wrong transcript:
- Monotonic audio cursor — a decreasing
audio_processed_untilis clamped to the prior value (the cursor never moves backwards; ST §4.4), with anaudio_cursor_decreaseddiagnostic (or a raise instrict_lifecycle). - Frozen-prefix immutability — an event that rewrites a segment's
already-frozen prefix (
text[:stable_until]changed) is suppressed with afrozen_prefix_rewrittendiagnostic (ST §4.2). One exemption, from that same section: a terminalclosedfinal MAY restate frozen text once (post-processing punctuation / ITN / casing) and MAY shrinkstable_until; the guard admits it and never clamps it back.
The full set of standard-layer diagnostic codes the guard can emit (read them off
the session with session.diagnostics()):
stable_until_clamped— a decreasing or invalidstable_untilwas clamped.audio_cursor_decreased— a decreasingaudio_processed_untilwas clamped.frozen_prefix_rewritten— an event rewriting a frozen prefix was suppressed.frozen_speaker_rewritten— an event changing (X→Y) or retracting (X→None) a frozen segment's already-acceptedspeakerwas suppressed. First assignment after freezing (None→X, the delay-to-final strategy) stays legal, and aclosedfinal is exempt (terminal correction).lifecycle_after_terminal— apartial/finalafter the segment becameclosed/supersededwas suppressed.lifecycle_partial_after_final— apartialafter the segment'sfinalwas suppressed.lifecycle_final_after_final— a secondfinalfor an already-final segment was suppressed.lifecycle_closed_superseded— asupersederetiring aclosedsegment was suppressed.lifecycle_retired_resuperseded— asupersederetiring an already-superseded segment was suppressed (an id retires exactly once).supersede_unknown_old_id— asupersedewhoseold_idscontain a never-announced segment was suppressed.supersede_reintroduces_segment— asupersedewhosenew_idsreuse an already-known id was suppressed.supersede_noncontiguous_old_ids— asupersedewhoseold_idsdo not form a contiguous block of the live reading order (in reading order) was suppressed: the replacements would have no defined placement, and an untimestamped transcript's word order would silently change.supersede_cross_speaker_merge— asupersedethat would merge segments carrying distinct non-null speakers into fewer segments was suppressed (someone's words would be silently mis-attributed).supersede_deletes_frozen_text— a pure-deletionsupersede(emptynew_ids) that would destroy a frozen prefix was suppressed.frozen_prefix_rewritten_supersede— a replacement group froze text that diverges from the retired frozen text it MUST preserve; the exposing freeze was suppressed.supersede_obligation_unfulfilled— soft end-of-session note: a replacement group ended with less frozen text than the retired segments it replaced.diagnostics_truncated— the bounded diagnostic channel overflowed, so one aggregated summary entry replaces the excess. It is rewritten in place as more overflow arrives, so a consumer must treat a later occurrence as superseding the earlier one.
Declare what you emit (capability ⇄ stream consistency)
Four streaming capabilities each gate one event field. Your declared
streaming capabilities and the events you actually emit MUST agree — your
stream may use less than you declare, but never more:
| If you emit… | …declare |
|---|---|
a non-zero stable_until | streaming.word_stability = FlagCap(supported=True) |
an audio_processed_until cursor | streaming.timestamps.mode ≠ "none" |
per-word words | streaming.word_timestamps = WordTimestampsCap(supported=True, …) |
a segment- or word-level speaker | streaming.diarization = DiarizationCap(supported=True) — add always_on=FlagCap(supported=True) if your model is architecturally unable to disable it |
The coherent no-timestamp streaming profile is the all-defaults combination:
leave word_stability, timestamps (mode "none"), word_timestamps, and
diarization unsupported, and emit none of those fields (use stable_until=0,
omit audio_processed_until, words, and speaker). A mismatch — for example,
declaring word_stability unsupported while emitting stable_until>0 — is a
capability⇄stream desync a client trusting your capabilities would mishandle.
Record a real session and assert it with
check_event_sequence(events, capabilities=engine.declared_capabilities). The
cross-check fails on any field your declaration does not back (codes
stream_exceeds_word_stability / stream_exceeds_timestamps /
stream_exceeds_word_timestamps / stream_exceeds_diarization). The standard
layer does not clamp these at
runtime — clamping would hide the bug; the contract is yours to keep.
Testing: assert invariants, not partial counts
Partials are lossy under backpressure. When the consumer reads slower than
you produce, the base coalesces pending partials for a segment (spec ST.6.4). So
the number of partial events a test observes is non-deterministic: the same
engine may surface five partials or none, purely by timing. A test asserting
len(partials) == N is therefore flaky. Assert the invariants instead:
- the final/reduced text is correct (
session.result(), or thefinalevent); - the partials form monotonic, never-rewritten prefixes — use the exported
assert_prefix_invariant(events)helper, which checks exactly that (a frozentext[:stable_until]is never rewritten andstable_untilnever regresses), tolerates any surviving partial count, and (unlikecheck_event_sequence) does not require a terminal event, so it also applies to a mid-stream slice.
Publish
Register an entry point under standard_asr.models (see
plugin-entry-points.md).
The engine class MUST be resolvable without calling the entry point.
Capabilities and the params schema are read from class-level ClassVars
without instantiating or authenticating the engine (CLI show, the registry,
REST GET /v1/capabilities/{model} and /v1/params-schema/{model}). Two forms
satisfy that, and the compliance suite accepts either:
- The entry point is the engine class itself. Nothing more is needed — the class is returned directly.
- The entry point is a factory function. Then its return annotation MUST
name your concrete engine class (
-> MyEngine), not theStandardASRprotocol: only the annotation is read, and a Protocol has no readableClassVars, so it breaks instantiation-free discovery.
What compliance actually checks is the outcome, not the form: it reports
class_metadata_unreadable when the class cannot be resolved either way — an
unannotated factory, or one annotated with the protocol.
Check with:
standard-asr compliance run
standard-asr doctor