AbyssBench: Checking Controller Timing

Replay controller events and check whether the recorded response satisfies an explicit timing contract.

A controller can issue the right command while the machine continues doing the wrong thing. AbyssBench explores that distinction: when a sensor reading becomes stale or an actuator stops responding, did the controller issue the required response before its deadline—and does the recorded evidence actually establish that?

I see it as a small tool for engineers developing control logic or reviewing test logs. It combines pytest with RTAMT, an existing temporal monitoring library. Its contribution is the surrounding workflow: explicit timing rules, repeatable traces, and evidence that keeps commands separate from measured behavior.

How it is used

I would start with the installation and offline demo. The public API walkthrough offers two starting points: import a supported JSONL or explicitly mapped CSV trace, or supply a trusted Python controller to the example simulator. Define acceptable ranges, measurement age, response deadlines, and prohibited transitions; run the check; then inspect each rule’s verdict and supporting events. Incomplete observation windows remain inconclusive. A recent packet carrying an old measurement does not become fresh just because it arrived recently.

A small example

The pump-and-valve example makes the distinction visible. The simulated valve sticks closed at 500 milliseconds. At 1,010 milliseconds, the controller detects the persistent mismatch and issues pump-off and valve-open commands. The valve remains closed. The response check can pass because the required commands were issued; that says nothing about mechanical recovery.

Simulation timeline: a valve sticks at 500 milliseconds and the controller issues safe commands at 1,010 milliseconds while the valve remains closed.
Recorded simulation events: commanded and actual valve positions remain separate.

Validation and limits

The original release passed 116 tests. The publication review nevertheless found two edge cases: a single observation could crash the monitor, and a command preceding a fault at the same timestamp could incorrectly satisfy its response rule. Both were repaired; the publication audit passed 124 tests. The audit preserves the failing cases.

The tool supports one clock and trusted controller callbacks; its fluid model is uncalibrated, with no hardware validation. A real trace and an independent engineer’s trial are the next useful tests.