Skip to content
SYBERLABS
Menu
RESEARCH NOTE · SB–0129 SEPTEMBER 2026 · REVISION 2

SYBERSHOKE

Shock the run.
Then count.

Sybershoke breaks a system on purpose: it kills workers, drops and duplicates messages, slows and fails the model provider. Then it checks the record of what happened against a short list of rules. Revision 2 adds a first real system.

SyberLabs Research
Fault-injection harness, a decision-path model & one real system’s source

Read the note
WORKING: ONE REAL SYSTEM’S SOURCE, RECORDED ANSWERSThe RISE Worker was checked against recorded Jev answers. Nothing called Jev or Kev.

Ask one question after the shocks.

In brief: systems built around a model fail between the parts. A worker dies holding a task, a message arrives twice, a provider is slow or answers badly, or a rule rewrites an answer it should have left alone. Sybershoke injects those faults from a seed, records what the system did as a plain-text history, and checks the history against invariants.

Revision 1 reported the Week 0 harness: a coordinator and worker pool reference simulator, 43 tests, four seeded bugs and five invariants. Those results are unchanged below. Revision 2 adds a second workspace. It models the Jev/Kev decision path that RISE uses to turn a reader’s request into a reading plan, reviews its own claims adversarially, and then runs the real RISE Worker source against 39 recorded production Jev answers.

The real run found a keyword override that sends “Help me drift off to sleep.” to 300 words per minute, and no fallback when the provider fails. The predictions were written before the first run, and one of them was refuted. A fix for the override has been merged into RISE. Nothing here called Jev or Kev.

THE ABSTRACTION

A seed, a fault plan, a history and a list of invariants. Nothing else decides pass or fail.

Start from what a task can end as.

A submitted task can end in only three ways: accepted, visibly failed, or neither. “Neither” is the failure the harness exists to find. “Accepted more than once” is the other. Every remaining invariant guards the bookkeeping that lets those two be judged.

  1. Judge the record, not the claim. The system under test never reports pass or fail. The checker reads only the history it leaves behind.
  2. A failure must be replayable. One seed gives one fault plan and one history, on every machine, so a failing seed can be rerun exactly.
  3. A failure must be small. A plan of dozens of faults explains little. Removing faults while the run still fails leaves the ones that matter.
  4. Faults interact. One fault can hide the failure another exposes, so each kind is run alone as well as together.

This led to a small scope: no service, no dependencies, one crate for the definitions, one simulator and one command. Adapters for real systems wait until the checker and the simulator agree with each other.

Five faults, five invariants, one shrinker.

Sections 02 and 03 are the Week 0 harness, as reported in Revision 1. Its repository is still unpublished, and it has not yet been reconciled with the Jev/Kev workspace in sections 04 to 06.

SEED→PLAN→RUN→HISTORY→CHECK
SHRINK → rerun with fewer faults until no single fault can be removed
Figure 1. The harness loop. The checker reads only the history.
Fault kinds
FaultWhat it models
Kill workerkill -9: the process dies mid-task and never returns.
Duplicate messageAt-least-once delivery: the queue delivers a message twice.
Drop messageA message lost in transit.
Stall workerA pause that can outlast its lease, followed by a late submission.
Corrupt outputUnusable model output for one attempt.

The checker enforces five invariants against the history:

  1. Nothing lost. Every submitted task ends accepted or visibly failed.
  2. Nothing accepted twice. At most one acceptance per task.
  3. No phantoms. Nothing appears that was never submitted.
  4. No contradictions. A task is not both accepted and failed.
  5. No unearned acceptance. An accepted attempt was delivered to a worker.

At least one worker is never killed, so a correct system can always finish. Each fault kind draws from its own random stream, so switching one kind on or off never changes another kind’s faults for the same seed.

A run ends only when every task is terminal, the queue is empty and no live worker holds a job. An earlier version stopped as soon as every task was terminal, which hid late duplicate deliveries and stale submissions. A test caught it. A first explanation that stalls needed to be aimed at the end of the run was tried, was wrong and was reverted. Both are recorded in the project’s decision log.

The harness catches what it should.

43passing tests
4 of 4seeded bugs caught
0failing runs for the correct implementation

The reference target is a small coordinator and worker pool with leases, a lease sweep, a periodic reconcile for lost messages and first-valid-result-wins acceptance. Four switches each remove one safeguard. The scenario is 200 tasks on 3 workers, with 2,000 seeds per campaign. The all column injects every fault kind at once; the others inject one kind alone.

Failing runs out of 2,000 seeds, by fault campaign
Targetallkillduplicatedropstallcorrupt
Correct implementation000000
No lease sweep170316840000
No reconcile199700199800
No idempotent accept200002000017060
Attempt-keyed accept0000610

The correct implementation failed in none of the 12,000 runs above. Each bug fails only under the fault kinds that should expose it: removing lease reclamation is never exposed by a duplicate, and removing the re-send of lost messages is never exposed by a kill. Each bug has a failing seed that shrinks to a single fault, though shrinking is greedy and some failing seeds stop at more than one. The shrunk plan for the first failing seed of the attempt-keyed bug is one stall: worker 0 for 29 ticks from tick 252.

One bug is invisible to the combined campaign.

The attempt-keyed bug accepts a result unless the same task and attempt number was already accepted. Two different attempts at one task can then both be accepted. A worker stalls past its lease, another worker takes the task and is accepted, and the first wakes up and submits. The bug fails on 61 of 2,000 seeds when only stalls are injected, and on 0 of 2,000 when every kind is injected together.

A likely reason is that kills remove the workers that would take the reassigned task. That is a hypothesis, not an isolated result. Either way, a test suite that only ran combined faults would not have caught this bug.

A model of the path around the model.

RISE turns a reader’s request into a reading plan: a pace, a sound, a visual look and a book. A model, Jev today and Kev next, chooses from a menu, and code validates the choice. So the useful faults hit the code around the model: time, failure, and the rules that rewrite or reuse an answer.

The second workspace is a model of that path, written from the SyberLabs project documents. It reproduces an 8 second deadline, a Kev scale-to-zero cold start of about 35 seconds, a keyword override that forces the night-drive look, a decision cache and a menu that answers are validated against. It never calls a model, and it is not the RISE source. Faults are error statuses, slow answers, cold starts, truncated answers and answers off the menu.

Invariants for the decision path
IdInvariant
I1Every plan admitted to the reader is on the menu.
I2Every request ends exactly once, after it arrived and within the deadline, and a failure keeps the reader’s text.
I3Every cache hit equals the latest model answer for its key before it. Judged from the plans, never from a label.
I4Explicit words are honoured: “slow” is not fast, “silent” is not loud, “no visuals” is off.
I5No request makes more provider calls than the cap.
I6Nothing lost, nothing accepted twice. Belongs to the Fanout target and the Week 0 simulator; not in this workspace.
I7A policy, switched on by choice: every request is answered with a plan even when the provider fails.

Running is separate from checking. A run writes a shoke-history/v1 file, one event per line, and the invariants read only that file. So a real system can be checked without linking any Rust: write its events in the format and run shoke check FILE. This is item 02 of the Revision 1 agenda.

Six seeded bugs each remove one safeguard, and each is caught by the invariant it targets. The keyword-override bug fails on 200 of 200 seeds and needs no faults at all. The other counts in the project report are set by assumptions; section 05 explains why only “caught” carries weight.

Try to prove the work wrong first.

Before building on the model, a review tried to break its claims. It recorded eleven findings, ranked by how badly each would mislead a reader of the report, with the evidence for each. Every fix has a regression test that failed on the old code. Tests went from 76 to 89. Three findings changed what the results mean.

I3 caught a cached fallback only when the history said so.

The seeded bug stored a fallback plan in the cache and labelled it origin=fallback. A real cache would not. With the label changed to origin=model, I3 caught the bug on 0 of 200 seeds. I3 now judges a cache hit from the plans alone: it must equal the latest earlier model answer for its key. The unlabelled bug is now caught on 25 of 200 seeds.

The model differed from the real RISE Worker on decisive facts.

Read against the Worker source, the model was wrong where it mattered:

  • The Worker makes exactly one provider call per request. There are no retries, so the retry-storm bug has no counterpart today.
  • Its 8 second limit covers the provider call only, not the whole request.
  • Its cache key is the exact intent plus a variation cohort, not normalized text.
  • It has no fallback. When the provider fails, the reader gets an error and no plan.

The keyword override and the menu validation were confirmed in the Worker’s code.

Every count is set by assumptions.

The idle time before a scale-to-zero host goes cold, the traffic pattern, the fault windows and the provider latencies are all assumed. Changing one of them moves a count by up to eight times. Shortening the fault windows from 1–16 seconds to 0.25–2 seconds moves one count from 49 to 12. What survives: whether a bug is caught at all, and that a preset floor removes every I7 failure.

The other findings were narrower: the shrinker could not tighten a fault window in time, and the file checker gave a wrong pass on some malformed histories. All are fixed and recorded, with the commands that reproduce them.

WHAT DID NOT MOVE

Every count in the model report is unchanged. The I3 count of 25 is now earned from the plans, not from a label.

Predict first, then run the real source.

39recorded production Jev answers
10predictions written before the first run
1prediction refuted, and reported

An adapter runs the real RISE Worker source, the code that handles a reading request, in Node. A fault proxy stands in for the provider. It replays 39 live Jev answers captured from production on 26 September 2026, used as test fixtures only. Redis and Neon are in-memory stand-ins, and time is virtual. The reader intents are the 39 cases from RISE’s own evaluation set, plus three probes. Every run writes the same history format, and shoke check judges it like any other.

The predictions were written down before the first run. Two of them are the adapter’s exit test: find the keyword misfire and the missing fallback with a replayable seed. Both were found.

Selected predictions, against RISE commit 082b3fa
#PredictionOutcome
P1The keyword override raises “sleep” and “slowly” requests to 300 wpm.Confirmed. “Help me drift off to sleep.” came back at 300 wpm with the night-drive beat and the neon palette. The recorded answer said 100 wpm.
P3Every injected provider fault ends in an error, never a plan.Confirmed. Seed 7, 30% faults: all 25 faulted requests ended in an error, none with a plan. There is no fallback.
P6A cache check keyed on text instead of the Worker’s key would fail.Refuted. I3 passed without the Worker’s key. With two asks per intent, a stale hit cannot occur.
P7Timeouts end just past a per-request 8 seconds.Confirmed. 14 of 14 ended at 8001 ms. Whether 8 seconds is per request or per call is a question for the spec, not a Worker bug.
P10The same seed gives the same history.Confirmed. Seed 7 twice gave byte-identical histories.

The misfire is a word match. The override looks for words like “drift” and fires for the schema version the production client sends. So “drift off to sleep” gets the night-drive look. A request that also said “no music” came back silent: the no-music words were honoured, the pace words were not.

One result was not predicted. A truncated provider answer reached the reader as “Jev could not be reached”, in 9 of 9 truncations in seed 7. No request was lost; the reader was told the wrong reason.

A second round asked each intent up to 25 times, to reach older cache entries. Its predictions were also written first, and one of them was refuted too. The Worker does serve an older key’s entry, but with one recorded answer per intent the old and new plans were always equal. So I3 cannot yet see a stale hit on real histories.

The fix.

A fix for the keyword misfire, and for the wording of a truncated answer, was merged into RISE as pull request 306. A code review of it found that “drift-off”, with a hyphen, still misfired; pull request 308 fixed that. Both are in RISE production: its release pipeline confirmed that the public site serves release a929768, which contains them. Rerun against the fixed Worker, the adapter finds no misfire. The missing fallback is a policy choice and is not changed.

Keep the claim smaller than the evidence.

The contribution is a composition of familiar ideas: fault injection against distributed systems [1], invariant checking over a recorded history, and reducing a failing input to a small one [2]. We do not claim to have invented any of them, and this work is not affiliated with those projects.

  • One real system’s source has been checked, not a live system. The RISE Worker ran on recorded answers, with stand-ins for Redis and Neon and virtual time. Provider latency was assumed.
  • Kev has not been checked. No recorded Kev answers exist, and recording them means calling Kev. That is the owner’s decision.
  • The sound-loudness rank is an interpretation. RISE does not rank its sounds by loudness. The adapter reads a rank from the words in RISE’s own sound descriptions. Four remaining flags depend on it, and whether those sounds are quiet is for RISE to say.
  • The decision-path counts are not rates. They are set by assumed traffic, fault windows and latencies. Only caught versus not caught survives.
  • The Week 0 simulator embodies one reading of leases, queues and acceptance. A real system can differ in ways that matter, such as where a re-queued message lands in the queue.
  • Greedy shrinking. A shrunk plan is small but not always minimal, and it keeps the invariant failing, not necessarily the same request.
  • Two workspaces, not yet reconciled. The Jev/Kev workspace is public and can be rerun: scripts/report.sh --check regenerates the model report, and adapters/rise-worker/check.sh /path/to/RISE reruns the Worker checks. The Week 0 harness repository is still unpublished, so its results cannot be independently reproduced today.

The preliminary hypothesis is that seed-replayable fault campaigns, run one kind at a time and together, find bugs that ordinary tests miss. One real system, run on recorded answers, is a start. Establishing it needs a comparison against real systems’ existing tests.

Progress by evidence, not by feature count.

01 / COMPLETE

Harness and reference simulator

Faults, checker, shrinker, four seeded bugs and a reproducible report.

02 / COMPLETE

Check from a file

A plain-text history format, so a real system can be checked without linking the harness.

03 / COMPLETE

A first real system

The RISE Worker source, on recorded answers. Predictions written first; one refuted and published.

04 / NEXT

Kev

Once a recorded Kev capture exists. Kev against Jev on the same cases, under faults.

05 / PLANNED

Fanout and I6

Reconcile the two workspaces, then kill a worker under both planners: nothing lost, nothing accepted twice.

06 / PLANNED

Other people’s systems

Chosen for one-command reproducibility. Maintainers are told, with the reproduction script, before any score is published.

A useful experiment fixes the workload, states its prediction first, runs each fault kind alone and combined, and reports passes as prominently as failures.

REQUEST FOR TECHNICAL CRITIQUE

Where does this miss a failure?

We welcome invariants we have missed, fault kinds that matter for agent systems, and systems that would make a good next target. A concrete seed that breaks a claim is especially useful.

Contact SyberLabs

Sources and implementation record.

  1. Jepsen. Distributed Systems Safety Research. Accessed 28 September 2026. Prior art for fault injection against distributed systems. No affiliation or endorsement implied.
  2. Zeller, A. and Hildebrandt, R. Simplifying and Isolating Failure-Inducing Input. IEEE Transactions on Software Engineering 28(2), 2002.
  3. SyberLabs. Sybershoke Week 0 reference report, generated by a script in the Week 0 repository, which also defines a CI check for it. Repository not yet published.
  4. SyberLabs. Sybershoke: the Jev/Kev workspace. Public on GitHub. No license is granted; all rights reserved.
  5. SyberLabs. Sybershoke report: the Jev/Kev decision path. Generated; scripts/report.sh --check fails when it is stale.
  6. SyberLabs. Red team: Sybershoke for the Jev/Kev decision path. Findings, evidence and fixes.
  7. SyberLabs. Phase 2: the RISE Worker adapter. Predictions and results against RISE commit 082b3fa.
  8. SyberLabs. RISE pull request 306. The keyword-override fix. Merged.

Revision 2, 29 September 2026, adds sections 04 to 06 and updates the limitations and agenda. The Week 0 results in sections 02 and 03 are unchanged from Revision 1. Every number on this page comes from the project’s committed reports. This is an engineering note, not a peer-reviewed paper or a claim of production readiness.