Skip to content
SYBERLABS
RESEARCH INDEX5 OCTOBER 2026

SYBERLABS RESEARCH

Every study,
with its limits.

One card per research artifact the lab has published: what it claims, the evidence state it has reached, what it does not establish, when, and where the source is. The states are the ones each repository earns. None is promoted by wording.

11 artifacts
2 reports on this site · 2 RISE status pages · 5 working papers · 2 instruments

Read the cards
HOW TO READ A CARD

Reports and notes.

Technical report · GC–01

GrokCell Execution: reliable execution for AI agents

After an AI model chooses an action, the application still has to enforce permissions and budgets, check the result and recover from crashes. The report defines a small contract for that work (read, decide, call, check, admit, yield) and tests a Python and SQLite prototype of it.

Status
  • Implemented a single-host Python prototype with optional SQLite persistence and two offline fixtures.
  • Tested 87 passing regression tests and process-kill experiments: fresh-process replay, death during a callback, death around admission.
  • Not yet live provider calls, verifier isolation, a task-quality or cost benchmark.
Establishes
Specific recovery behaviour across process exit for two offline fixtures: a restart replays a completed step without another callback, and an interrupted step is not resent under the same identity.
Does not establish
Live-provider quality, security isolation (the verifier is trusted host code), exactly-once for remote effects, power-loss recovery, or any live success rate, latency or cost per accepted result.
Date
26 September 2026 · revision 1
Source
Read the report · implementation in SyberLabs/grokcell-execution, a private repository (authorized access required)

Research note · SB–01

Sybershoke: shock tests for multi-agent systems

A deterministic fault-injection harness: a seed gives a fault plan, the system’s run is recorded as a plain-text history, invariants judge the history alone, and a shrinker reduces a failing plan to the faults that matter. Revision 2 runs the real RISE Worker source against recorded Jev answers.

Status
  • Measured Week 0 reference simulator: 43 passing tests, 4 of 4 seeded bugs caught, 0 failing runs for the correct implementation over 2,000 seeds per campaign.
  • Tested the Jev/Kev workspace: 89 tests after an eleven-finding red-team review, each fix with a regression test that failed on the old code.
  • Measured the RISE Worker source on 39 recorded production Jev answers: 10 predictions written before the first run, 1 refuted; a keyword override that sends “drift off to sleep” to 300 words per minute and a missing provider fallback, both found with replayable seeds.
  • Not yet a live system; Kev; the Week 0 repository is unpublished.
Establishes
The harness catches the seeded bugs it was built to catch, and one real system’s source misbehaves in two replayable ways on recorded answers. The override fix is merged in RISE (SyberLabs/RISE#306).
Does not establish
Failure rates: the decision-path counts are set by assumed traffic, fault windows and latencies, so only caught versus not caught survives. Anything about Kev, since no recorded Kev answers exist. The live RISE service: Redis and Neon were stand-ins and time was virtual. The sound-loudness rank is the adapter’s interpretation, not RISE’s.
Date
29 September 2026 · revision 2
Source
Read the note · SyberLabs/sybershoke at 45956fe (REPORT.md, docs/REDTEAM.md, docs/ADAPTER-RISE.md)

Status pages, not papers.

Status page · RISE

RISE, Jev and Kev: reader-owned AI

Since RISE #294 the RISE server calls no model and holds no model key. An AI reading request runs on a connection the reader owns: hosted Jev through the reader’s own OpenRouter account, or a pinned Kev-4B on the reader’s own computer. Reading and manual settings need neither.

Status
  • Deployed the Cloudflare Worker serves the app and a static decision catalog; the six former inference routes answer 410 with instructions to connect.
  • Implemented bounded decisions: finite choice questions built from the public catalog, an answer admitted only if every choice was offered, and a deterministic mapping to reading settings. The same contract serves hosted Jev, local Kev and the evaluation harness.
  • Tested the local Kev path, on Windows with an RTX 5080.
  • Not yet a RISE-specific Kev evaluation; the Apple Silicon path.
Establishes
Where the model boundary sits in RISE today, and that a model can pick an offered value and nothing else.
Does not establish
Kev’s quality on RISE requests. Jev’s 48 of 49 result belongs to the historical evaluation below and is not a Kev result. Nothing hosted runs Kev on SyberLabs’ behalf.
Date
October 2026 · page revised 5 October 2026

Case study · historical

Jev in RISE: AI decisions with clear limits

The recorded Jev integration. Jev answers typed questions (choose one option, estimate a yes-or-no probability, rate on a scale), RISE’s own code validates every choice before playback, and the reader stays in control.

Status
  • Deployed the integration ran in RISE before #294 retired shared inference; readers’ requests chose a released text and passage and configured the reader from 23 browser-synthesized sound options, seven typefaces and five text sizes.
  • Measured the post-fix live record of 27 September 2026 reran all 39 evaluation cases: all 39 responses valid, 48 of 49 explicit choices matched, all 19 paired contrasts differed.
  • Not yet independent verification of the Scriptorium composition feature.
Establishes
That a typed-decision integration with validation in code worked for one request path, and one pass of its choices against a fixed case set.
Does not establish
A verified live Kev deployment: these results must not be attributed to Kev. Consistency across runs or reader enjoyment, which a single pass cannot show. Any partnership with TypeSafe AI.
Date
Historical · evaluation record 27 September 2026 · page revised 5 October 2026

Five studies at different stages.

The papers repository is an index of research artifacts, not of finished papers. Each folder defines its evidence, claim limits and reproduction steps; the sentences below are taken from those folders.

Working paper · internal-review draft

DTBR-MC: the marker is not the binding constraint

A falsification-oriented Monte Carlo audit of deep-time hazard-warning assumptions. In this model family, symbolic factors (markers, legibility, dread) act as modulators while physical factors (access, capability, severity, certainty of consequence) act as gates, and the capacity-threshold hypothesis it set out to test is algebraically impossible in an additive behavioural model.

Status
  • Implemented a seeded, vectorized simulator; H1, sensitivity and H3 experiments; a first experiment bundle with outputs and figures.
  • Tested deterministic regression and falsifiability tests in tests/.
  • Not yet review: the draft is internal, and its references still need verification against primary sources.
Establishes
Which modelling assumptions are load-bearing once the warning-design problem is written as replaceable equations. Every result carries an epistemic label: audit, extrapolation, cartography or exploratory.
Does not establish
How distant-future people will behave, operational nuclear-waste policy, or calibrated future-behaviour estimates. The simulator is a consistency auditor, not a predictive instrument.
Date
Published to SyberLabs/papers 26 September 2026

Manuscript · draft

Green Hypercube: coverage artifacts in computational ethnobotany

Apparent structure in integrated plant-data search benchmarks can be inflated by coverage and reward-density artifacts. Six sequential search strategies run over a plant manifold with four cue channels, and raw cue–reward coupling is reported beside coverage-controlled and residualized coupling; the manuscript reports that the raw coupling collapses under those controls.

Status
  • Implemented a Python package with an offline synthetic path, live-data configurations, and a validation suite: reward permutation, graph rewiring, phylogeny shuffles, coupling, residualization and matched-density sweeps.
  • Tested a pytest suite for the pipeline, strategies, controls, coupling and metrics.
  • Not yet author confirmation and TODO items in the draft; raw external caches and live result bundles are not published.
Establishes
A methodological result: matched-density comparisons are necessary before a strategy is credited with real signal across pools.
Does not establish
Anything about Indigenous discovery mechanisms or community knowledge. It is not a bioprospecting recommendation engine, and documented use is not treated as a proxy for latent biological value.
Date
Published to SyberLabs/papers 26 September 2026

Empirical research artifact

Grokking Scaling Theory

Tests scaling-law and order-parameter claims around grokking in modular arithmetic by fit competition across candidate laws rather than one hand-picked curve. Phase 2 (July 2026) asks whether a label-free sheaf “gluing” parameter or a logical-cells decidability parameter supplies the architecture-universal order parameter that phase 1 lacked.

Status
  • Measured expanded validation on 21 scaling points, 19 of them measured; phase 2 on 63 deterministic training runs across five architectures, bit-deterministic from the committed seeds (verified 18 of 18).
  • Implemented fit competition, a corrected scaling law, RG-inspired flow models, and the sheaf and logical-cell instruments.
  • Not yet publication: more measured runs, tests and manuscript cleanup remain, and some anchors and bibliography assets are research-stage.
Establishes
As the repository reports it: the sheaf parameter beats the variance incumbent but leads grokking by 10 to 25%, no logical cell assembles, depth amplifies the Fourier code while trainability collapses, and width rather than depth is the dominant quantization pressure. Pooled fits point to regime dependence rather than universality.
Does not establish
A solved theory of grokking, or architecture-independent universal scaling laws.
Date
Phase 2 July 2026 · published to SyberLabs/papers 26 September 2026

Curated public package

TOK: a governed epistemic architecture

Reasoning represented as a sequence of validated artifacts (mechanism graph, dynamics binding, divergence trace, observation registry, evidence-transition proposal, human review gate), each carrying boundaries such as candidate_not_evidence and may_update_evidence: false. The package holds sanitized freeze summaries, replay artifacts, a public validator and a toy reference pipeline.

Status
  • Measured public summaries of the Paper 1 benchmark freeze (a naturalistic holdout of 240 problems: exact coordinate accuracy 173/240, axis decision accuracy 885/960) and the Paper 2 dynamics atlas (12 systems across 7 regimes, 116,640 admissibility fits, 144 collapse transects, every recorded acceptance gate passed).
  • Tested 26 safe research checks passed in the source workspace on 18 June 2026; the public validator checks release boundaries, freeze summaries and the toy pipeline.
  • Not yet public: the full internal engine, private answer keys, maps and scorecards are not in the package.
Establishes
Narrow claims about how TOK represents reasoning as auditable, human-gated artifacts and how its freeze summaries keep research-only boundaries. The atlas supports claims about the tested implementations, corpus, ladder and measurement grid.
Does not establish
That TOK discovers real-world causal truth, proves a candidate mechanism, or autonomously updates evidence or production systems. It does not support broad invalidation of LTC/LNN model classes, and the private benchmark material cannot be re-verified from this package.
Date
Freezes through June 2026 · published to SyberLabs/papers 26 September 2026

Exploratory research system · interim findings

Vital Language: coherence, agency and literary vitality

Can controlled generation dynamics make language-model prose feel more alive without collapsing coherence? The founding bet, that deterministic chaotic logit modulation raises multifractal structure and perceived vitality, inverted: at 0.5B to 1.5B scale, token-level chaos does not beat plain sampling on meaningful vitality proxies, and a prompt-level persistent-speaker scaffold is the reliable lever, buying non-collapse rather than vitality.

Status
  • Measured paired epsilon sweeps on Qwen2.5-0.5B and 1.5B on CPU (6 prompts × 3 seeds, legibility-gated, bootstrap intervals): white noise drives self-perplexity from 7.8 to 176 while chaos at the same magnitude stays near 10; the agency scaffold leaves 0 of 18 passages degenerate against 3 of 18 for plain sampling.
  • Implemented one generation harness with chaotic, matched-noise and white-noise controls, MFDFA, coherence and trajectory metrics, and a blind rating scaffold.
  • Not yet independent human raters.
Establishes
The coherence and degeneracy results, which are rater-independent counts, and that the MFDFA instrument works on real literature (a shuffle surrogate collapses Joyce’s width and barely moves Austen’s).
Does not establish
Any claim about felt vitality. All human-judgment data is n = 1 and the rater is the assistant model itself, so every vitality claim is a hypothesis about sign until independent raters are collected.
Date
Interim · published to SyberLabs/papers 26 September 2026

Run the systems, keep the record.

Instrument · SyberLabs/cross-platform

The Instrument Panel: five real systems in one inspection shell

A single local platform through which SyberRuntime, Barn, Bough, OSAHR and Relay can be run, inspected and understood. Nothing simulates those systems: every number, refusal, hash and graph shown was produced by the named source repository, and any event opens to the source file and line that produced it.

Status
  • Implemented an adapter layer over the unmodified source checkouts, an event stream with per-event provenance, six guided scenarios (input configuration, not playback: each step declares what it expects and reports what happened), and the three seams the source systems already define.
  • Implemented a fidelity matrix that names the symbol that ran for each capability, and an integration map that records the commands run and the pinned source commits.
  • Not yet an automated test suite in the repository; a public deployment. It runs at 127.0.0.1:8765 from a checkout with a Python environment and Node 24.
Establishes
That the five systems can be driven live from one shell with their behaviour attributed to source, and that a refusal the scenario did not predict is reported as such.
Does not establish
Anything about Jev or Kev: no provider call is active. SyberRuntime’s acceptance gate reports fail on a fresh clone because its evidence corpus is gitignored; two Barn benchmark tests fail upstream; Bough’s Cypher is generated but never pushed; Relay’s Obsidian, database and Cloudflare paths are not exercised; the Barn→Bough twin is a shape from one short run, not a fitted model.
Date
27 September 2026
Source
SyberLabs/cross-platform at 5e031ae (README, docs/FIDELITY.md, docs/INTEGRATION_MAP.md)

Evidence packets · SyberLabs/OSAHR_Cell

OSAHR decision workbench: licensed evidence packets

One auditable loop over the frozen Experiment 06 corpus: a scenario is validated, freeze and analysis checksums are pinned, the claim grammar is rescored, an action license and a claim license are derived, and a JSON and HTML packet is written that replay recomputes. Editing a packet does not grant a claim.

Status
  • Implemented python -m workbench decide and replay, three scenarios (identity, high stress, long outage), and the packet format in workbench/packet.py.
  • Tested 18 tests in workbench/tests/test_workbench.py.
  • Not yet real-network calibration: it is always graded Proposed, and promoting it to Measured fails replay.
Establishes
That a named scenario becomes a human-pending action and a replayable evidence packet without licensing a directed effect the ensemble withholds: an unresolved hold acts under hold with the claim license denied. Every packet carries Known, Measured, Inferred and Proposed grades.
Does not establish
A new simulator, or any directed effect beyond what the Experiment 06 record admits. It refuses to mint a licensed packet from an attacker-supplied ensemble, from claims.score MCP input, or from a scenario outside the evaluation corpus.
Date
27 September 2026

What this index reflects.

Reflects SyberLabs/papers@e925ed6, SyberLabs/cross-platform@5e031ae, SyberLabs/OSAHR_Cell@7293e27, SyberLabs/sybershoke@45956fe and SyberLabs/RISE@998d725, and the site pages linked above as of SyberLabs/SyberLabs.github.io@168fb0d · verified 2026-10-05. Each state is earned by code, a named test, a measurement under stated conditions, or a deployment; none is promoted by wording. When a repository changes, this index changes with it.

Not indexed here: the private SyberLabs/grokcell-execution repository, reachable only through the report above, and the per-project evidence rows, which live on each project page.