Guides / Ferrum Anvil / Load Tests

Load Tests

Load runs send the same requests as Send, through the same engine, from a separate worker process, and report what happened without averaging it away.

Ferrum Anvil v0.1.1 is an unsigned preview. Verify downloads with SHA256SUMS. macOS installers are not Developer ID signed or notarized and Windows installers have no code-signing certificate. In-app updates use verified minisign signatures; this does not code-sign the installers. See the preview downloads and Anvil overview.

Only Load What You May Load

  • Every run needs an explicit acknowledgement. Before starting, Anvil shows the destinations, the planned rate or concurrency and duration, and a reminder that you must own or be authorized to test the target. Start stays disabled until you acknowledge.
  • Imported load plans are untrusted and never start automatically.
  • An abort rule can stop a run when the failure ratio in a trailing window crosses a threshold; the report is then marked partial.

Pick the Model That Matches the Question

WorkloadHow it behaves
Open (arrival rate)Arrivals are scheduled independently of response time. When max_in_flight is reached, an arrival is dropped and counted, never queued, and start lag is measured.
Closed (virtual users)Each user waits for its response before the next iteration, so slower responses lower the rate. These reports never claim a sustained arrival rate.
IterationsA fixed count over a bounded number of lanes. Slower responses lengthen the run instead of dropping work.

Stages ramp linearly. Each iteration runs a weighted mix or a chain of saved requests, with optional CSV or JSON datasets and a warmup period that is excluded from summary metrics but kept, flagged, in the timeline.

One Protocol, One Unit

Every step of a plan is one call through the same engine path as a manual Send, and produces one unit. What a unit is, and what counts as success, depends on the request's protocol.

ProtocolOne unit isSuccess meansCounts reported
HTTP/1.1, HTTP/2, HTTP/3One request and its response. Redirects, retries and an HTTP/3 to TCP fallback are attempts inside it.Status below 400, no SOAP fault or GraphQL error, and assertions passStatus codes; fallback attempts, requests with a fallback and requests over HTTP/3. A fallback's time is part of the request's latency.
Unary gRPC (native over HTTP/2 or HTTP/3, and gRPC-Web)One callStatus 0 and assertions pass; HTTP 200 alone never countsCalls by status code, OK and non-OK, and calls with a missing status. A missing status is incomplete, never success. A deadline that elapsed locally is a timeout, never reported as DEADLINE_EXCEEDED.
Server-streaming gRPC and gRPC-WebOne streamThe server ended it with status 0, and assertions passStreams opened, messages received, streams with messages, time to first message, and the status codes
Server-sent eventsOne streamA 2xx event stream, and assertions passStreams opened, events received, time to first event, and how each stream ended: by the server, by Anvil (the request's event limit), by an idle timeout, or abnormally
WebSocket (HTTP/1.1, HTTP/2 or HTTP/3 bootstrap)One session: the handshake, the request's scripted messages, then the closeHandshake accepted, closed with 1000, 1001 or no code, and assertions passSessions opened, handshakes rejected, sessions not opened, sessions closed cleanly, messages sent and received, and close codes by who closed. Round-trip times only when the request expects replies (expect_messages).
TCP, TCP+TLSOne connection carrying the request's framesCompleted, the expected frames arrived (with a framing preset), and assertions pass; fewer frames is an application failureConnections, frames and payload bytes each way, expected frames met or short, partial trailing frames, and peer closes
UDP, DTLSThe request's datagrams, then its response window (for DTLS, after a handshake)At least one datagram received, and assertions passDatagrams sent and datagrams received, as separate counts; exchanges with a response and with no response observed; echoed and repeated payloads; ICMP unreachable; and for DTLS, handshakes attempted, completed, failed and timed out, with their duration
  • A plan measures exactly one kind of unit, so every count, rate and percentile in its report has one denominator. The load editor and anvil load check name the unit and its definitions before a run.
  • A UDP or DTLS exchange with nothing received is “no response observed”: neither a success nor a failure, and it has no latency. A datagram sent is never counted as delivered, and the received-to-sent ratio is labelled an observation, not a delivery rate.
  • Connection modes. For HTTP requests and gRPC calls and streams, persistent keeps pooled connections per virtual user (for gRPC, one pooled channel per destination) and fresh opens a connection per unit. SSE streams, WebSocket sessions, TCP exchanges and UDP/DTLS exchanges always open their own connection or socket, and the report says the mode does not apply to them.
  • Each unit runs the automation path of its request: the script, stop conditions, idle close and total deadline apply. Commands typed into a live interactive session have no load equivalent.

Checked against the real gateway

The lab's streams profile runs short, low-rate gRPC, WebSocket and UDP load plans through a real Ferrum Edge gateway, on the supported 0.9.8, 0.9.7 and 0.9.5 releases, and compares Anvil's counts with independent ground truth: the backend's own logs and, for gRPC and WebSocket, the gateway's logs.

Refused Before Any Traffic

Combinations without a defined unit are refused when the plan is checked, with a typed reason. The editor shows it and the run cannot start.

RefusedWhy
Plans that mix protocols, such as an HTTP login followed by a WebSocket sessionTheir counts and latencies have different denominators. Split the plan per protocol.
Client-streaming and bidirectional gRPCA long-lived client or two-way stream has no single completion or per-message denominator yet.
gRPC with server reflectionEvery call would first run a reflection call, so a unit would not be one call. Import the .proto files or a descriptor set instead.
SSE with automatic reconnectReconnection turns one stream into several connections with server-chosen delays.
UDP or DTLS through a CONNECT-UDP (MASQUE) proxy or an HBONE tunnelEach exchange would open its own QUIC or mTLS connection and tunnel, and there are no tunnel counts.
Requests with 0-RTT early dataHandshakes that share session tickets are serialized, and the report has no early-data counts.
HTTP or gRPC through an HBONE proxy in persistent modeTunnels carry one execution's identity and are never pooled, so persistent mode could not be honoured. Fresh mode is allowed.

Combinations the engine refuses for any send, such as native gRPC with the HTTP/1.1-only policy or UDP through an HTTP proxy, still fail on each send before any bytes are written. They are local failures in the report, never successes, and an abort rule can stop such a run early.

A Separate Worker, the Same Engine

  • Each send is the same call as a manual Send: variables, auth, TLS policy and diagnostics included. HMAC nonces, DPoP proofs and JWT time claims are fresh for every send, and an OAuth token refresh happens once for the whole run.
  • The desktop app launches its own executable as the load worker and passes the job over stdin, never the command line or environment. The job carries only the secrets its requests reference.
  • Cancel stops scheduling and drains in-flight sends. If the app goes away, the worker stops. If the worker crashes, Anvil rebuilds a partial report from the last progress snapshot and says what is unknown.
  • Locking Anvil stops load workers; partial reports are kept.

Reading a Report

Ferrum Anvil load report for a completed run named E2E smoke load: 300 iterations started and 300 requests completed, with zero transport failures, timeouts, application failures and assertion failures. It names the load unit (HTTP requests) and defines what completed, success and latency mean, counts HTTP/3 fallback attempts separately, and shows latency percentiles for successful sends, status code counts, no failure categories, JSON, CSV, Timeline CSV and HTML export buttons, a collapsed load generator health section, and notes explaining the workload model, the connection mode and how latency is measured.
Pre-release build. The figures come from a 300-request automated smoke run against a local test server on the same machine. They are not a benchmark.
  • Balanced counts. Scheduled, dropped, started, completed, failed, timed out, canceled and still-in-flight always add up, for iterations and for units. The protocol counts balance against the unit counts too, and a report that does not balance says so.
  • Latency covers connect to last body byte for requests and calls, and until the end for streams and sessions. For UDP and DTLS exchanges it is the time to the first response instead, because an exchange always waits out its response window. Local preparation and token acquisition are reported separately as setup time.
  • Timeouts are censored: counted, but kept out of the latency distributions because their true latency is unknown.
  • Percentiles come from merged HDR histograms, never from averaged per-worker percentiles. Success and failure latencies are separate.
  • A protocol section shows the unit's definitions and its counts, labelled with the unit's own nouns: gRPC status codes with their names, streams and sessions, close codes by who closed, round trips or “not defined”, frames, datagrams and DTLS handshakes.
  • Generator health reports the worker's own CPU, memory and descriptors. For open workloads, a run that could not keep its schedule is marked “target not achieved”; that says nothing about the target's capacity. Generator health is unavailable on Windows.
  • Failure samples are bounded and built from redacted records, without response bodies.
  • Exports: JSON sealed with an integrity hash, CSV guarded against spreadsheet formula injection, and a self-contained HTML report with no scripts.
  • Comparison refuses runs of different load units, such as HTTP requests against WebSocket sessions. For the same unit, it withholds latency deltas when two runs are not comparable, for example a different workload model, connection mode or request set, and says why.

What Load Tests Do Not Cover Yet

  • The combinations in the refusal table have no load unit yet.
  • No load action holds sessions open while sending messages at a rate. A WebSocket unit is connect, scripted messages, close.
  • WebSocket round trips pair each scripted message with the reply in the same position, so they fit echo-style exchanges. A session whose reply arrives before its message is counted as unpaired and adds no round trip.
  • HTTP/3 with the automatic policy tries QUIC again for every request. Each failed attempt is counted, and its time is part of the request's latency.
  • Load through an HBONE tunnel in fresh mode is allowed for TCP-based protocols but has not been exercised under load.
  • Message counts per stream or session are reported as totals and means, not percentiles.
  • One worker process per run on one machine; no distributed load.
  • No JMeter adapter.