Skip to main content

Pick the smallest scope that proves the thing

Not every Varianz test needs a coordinator. Match the scope to what you’re actually testing: Cover a branch locally first, then keep a smaller number of coordinator-backed and end-to-end tests for the delivery path a local test deliberately doesn’t exercise. Making an E2E test the first — or only — test of a VPoint’s behavior is the common mistake. The rest of this page is the coordinator-backed lifecycle. For local stages, everything below (sessions, sync, the session header) simply doesn’t apply.

The coordinator-backed lifecycle

Every coordinator-backed Varianz test, in every framework, follows the same six steps:
  1. A test session is created (automatically, by the fixture or extension).
  2. Stages are inserted — CEL expressions bound to specific VPoints.
  3. Sync — the test waits until connected services confirm they received the stages.
  4. Trigger — the test calls the service, passing the session ID.
  5. Assert — on response values and captured samples.
  6. Cleanup — stages removed, session destroyed (automatic).
Stages are session-scoped: only requests carrying the matching x-varianz-id see them. Production traffic flows through unmodified.

What must be running

Coordinator-backed tests are true integration tests. Three things must be up and connected to each other:
  1. A coordinator — see Running the coordinator locally.
  2. Your instrumented service(s) — connected to that coordinator, with VARIANZ_ENABLED=true (plus, for Java/Kotlin, the agent attached), and VARIANZ_INSECURE_ALLOW_PLAINTEXT=true if the coordinator is plaintext.
  3. The test process — connected to the same coordinator via the framework fixture, with the same two environment variables. Test fixtures fail fast with a clear error if the SDK is disabled.
The test inserts stages into the coordinator; the service receives them over its subscription; the trigger request carries the session ID; the VPoint applies the stages; samples stream back to the test.

The sync barrier

await_sync_or_fail(timeout) (awaitSyncOrFail/AwaitSyncOrFail) blocks until every currently connected subscriber whose VPoints match your stages confirms receipt, and fails the test on timeout. Always call it between inserting stages and triggering the service — insertion is asynchronous, and without the barrier your trigger can race ahead of stage delivery.
The barrier does not fail on stage-validation errors — they are advisory. A stage whose CEL does not compile against the target schema is reported as a validation issue, not an error: the barrier still returns cleanly and your test proceeds. Validation is also best-effort in a second sense — it can only check against schemas the coordinator already holds, so a stage inserted before its subscriber has submitted a VPD is not checked at all, and reports no issues rather than reporting “unchecked”.The practical consequence: a malformed stage can produce a green test in which the stage never applied.Validation issues are reported as a [varianz] WARNING at session teardown, naming the stage, the VPoint and the cause:
Read that output — it arrives after the test has already reported green, so a CI job that only surfaces pass/fail will hide it. Diagnostics the engine rates as warnings (an unknown field, a mistyped field, an unknown dotted function) are included and prefixed warning: . You can also inspect issues mid-test via await_sync (the non-failing variant).One gap remains: an unknown bare function name — nosuchfn(1) rather than some.nosuchfn(1) — is reported by nothing, because bare names may legitimately be resolved at runtime.To catch expression errors eagerly, compile them in-process where the VPoint is registered — attach_cel_stage / attachCelStage raises immediately on a parse error, unknown field, or unknown struct type. That path is a useful lint step even when the real test delivers the stage through the coordinator.

Zero subscribers

A result of “0 of 0 subscribers confirmed” means no connected service matched your stages — usually the service under test isn’t connected to the coordinator yet. The barrier treats this as satisfied rather than blocking, and the session re-checks at teardown, printing a [varianz] WARNING if nothing ever received your stages. When a stage doesn’t apply, look for that warning first — it points straight at the disconnected service.
The most common cause is a service running without VARIANZ_ENABLED=true. Because the SDK is disabled by default, such a service starts and serves traffic normally but registers no VPoints — so it is never a subscriber, and every stage you insert reaches nobody. Check its startup log: one line, [varianz] DISABLED (source: …) instead of [varianz] ENABLED (…). If sync reports 0 of N or times out with some subscribers missing, one of your services is connected but didn’t register the targeted VPoint — check the VPoint name (case-sensitive, including namespace prefixes) and that the service actually reached its registration code.

Lazy VPoints

With typed schemas — annotations, the scanner, or codegen, which is the recommended setup — VPoints register at startup and stages work from the very first request; nothing in this section applies. The rule: a VPoint defers its schema when any parameter it has to expose lacks a usable type annotation. A missing return annotation is not a trigger — that VPoint still registers at startup, with an unknown return kind. A parameter named in opaque= is exempt, which is why opaque=["context"] on a gRPC servicer method (self, request, context) keeps the signature eager rather than making it lazy. You do not need a warm-up call. Attaching a stage performs the deferred resolution itself, so a stage inserted before the VPoint has ever run still applies on the very first call — including expressions that need the schema, like args.x or invoke():
What you lose is validation, not delivery. Until the schema resolves there is nothing to check an expression against, so the sync barrier accepts expressions it would otherwise reject:
  • a reference to a field that does not exist — args.nope — passes await_sync_or_fail() clean, then silently does nothing, with no diagnostic anywhere;
  • a type error — args.x + 1 where x is a string — also passes clean, then throws at call time: RuntimeError: … Unable to convert Int64 to a string operand.
Annotate the parameters (plus opaque= for any slot that can’t carry an annotation), or wire up build integration, and both become insert-time errors instead.

Insertion vs. delivery

Inserting a stage resolves its routing target against the coordinator’s catalog. By default insertion is lazy: a target that matches nothing yet is stored unresolved and attaches when a matching VPoint appears — convenient for lazily-registered VPoints, but it means insert success alone doesn’t prove routing. Resolution is sticky: once inserted, a stage never migrates to a different VPoint. The sync barrier is the delivery check — it confirms subscribers connected at that moment, never future ones. The practical recipe: insert → await_sync_or_fail → trigger.

Triggering with the session ID

  • HTTP: set the x-varianz-id header. From Python, use a requests.Session() with the header on the session object — bare requests.post(headers=...) loses headers on redirects.
  • gRPC: set x-varianz-id metadata on the call.
  • Browser (Playwright): the fixture sets the header on the page for you.
The session must propagate through every hop of a multi-service flow.

Choose your framework

The test language is independent of the service language — see Test across services.