# Logical Works > Logical Works is a Toronto software company. Explore our research, meet the work behind it, or talk to us about your project. Logical Works is a Toronto software company. Explore our research, meet the work behind it, or talk to us about your project. ## Direct queries Use [the knowledge API](https://logicalworks.ca/api/knowledge) to request answers by question, node ID, or route. The API returns bounded records with source links. [API schema](https://logicalworks.ca/.well-known/knowledge-openapi.json). ## Research - [The metric returned twelve for every prompt because twelve was built into the metric](https://logicalworks.ca/research/logit-lens-measured-its-own-definition): A GPT-2 interpretability run found a quiet middle and a large final rewrite, then exposed two measurement defects in its own settling metrics. - [An adaptive reader learned where to look. It did not learn how to scale.](https://logicalworks.ca/research/adaptive-reader-did-not-scale): A hard-read recurrent controller reached oracle cost on lookup tasks, then failed under larger worlds, random identifiers, and seed controls. - [Why a global waveform memory loses its compute advantage on exact worlds](https://logicalworks.ca/research/waveform-memory-falsification): A formal and mathematical audit found that exact lookup removes the claimed compute advantage of a global Fourier world representation. - [We counted 264,224 agent turns. The causal labels are not ready to publish as causal findings.](https://logicalworks.ca/research/auditing-agent-failure-corpus): A cross-framework workflow corpus has useful coverage, but regex labels, autonomous subturns, and missing turn-level bytes limit causal claims. - [A neural network's privileged self-knowledge may be operationally empty](https://logicalworks.ca/research/operationally-empty-self-knowledge): A position paper tests three operational definitions of model self-knowledge and leaves a pre-registered recurrence experiment as the empirical target. ## Company - [Careers](https://logicalworks.ca/careers) - [Contact](https://logicalworks.ca/contact) - [Privacy](https://logicalworks.ca/privacy) - [Terms](https://logicalworks.ca/terms) ## The metric returned twelve for every prompt because twelve was built into the metric Source: https://logicalworks.ca/research/logit-lens-measured-its-own-definition Published: 2026-08-31 Status: Method audit Authors: Logical Works Research ## The failed measurement The first Stage-0 run tested whether GPT-2 small behaves like an iterative settling process. It measured relative residual updates across 12 layers and defined settling depth as the first layer whose top prediction matched the final layer's prediction and stayed matched. Every easy and hard prompt returned a settling depth of 12. The tempting interpretation was that nothing settled early. The definition made that conclusion unavailable. A large rewrite in layer 12 changed the top prediction, so the only layer guaranteed to equal the final layer and remain equal was layer 12 itself. The metric measured its own stopping rule. The vector of mean relative updates still contained useful information: `[0.49, 1.06, 4.45, 0.42, 0.36, 0.27, 0.21, 0.20, 0.20, 0.23, 0.53, 7.04]`. GPT-2 small showed a quiet middle and large rewrites near the boundaries. The run did not establish monotone decay or different settling depth for easy and hard prompts. ## The repaired metric failed differently Stage-0b replaced exact top-one agreement with KL divergence from each layer's logit-lens distribution to the final distribution. The threshold was 0.1 bits. Across nine prompts, the result was again 12 every time. Intermediate layers remained two to eight bits from the final distribution, then the final point dropped to zero by definition. This second result is stronger than the first because the curve itself shows a late readout transition. The scalar summary is still saturated. A useful repair would normalize divergence to the initial layer, exclude the trivial final point, report the full curve, and compare a raw logit lens with a tuned lens. ## A second bug in the same run The lab's novelty index computed Shannon entropy over raw feature magnitudes. Adding a parameter-count field of 124,439,808 swamped the other features and drove the run's entropy to 0.01 bits. That was a scale bug, not a discovery about novelty. Applying `log1p` before normalization changed the indexed values to 8.25 bits for Stage-0 and 7.86 bits for Stage-0b. This is why measurement software belongs inside the research result. The model run surfaced facts about GPT-2, but it also falsified two parts of the lab's own instrumentation. Publishing only the model curve would hide the more transferable lesson. ## What remains defensible GPT-2 small used 1,204 MiB peak RSS and 237 MiB of MPS allocation in Stage-0b. The prompt "The capital of France is" produced " Paris" only five times in 100 sampled first tokens; the fixed next-token distribution preferred grammatical continuations such as " the" and " a". Hidden states were deterministic for the fixed input, so the observed 100-run variation came from sampling one distribution. The finding is limited to GPT-2 small, its tokenizer, this prompt position, and the raw logit-lens setup. It is not evidence that factual knowledge is absent from later positions or larger models. It is evidence that the immediate distribution at this scale is grammatical before it is encyclopedic. ## Evidence ledger - **Executed:** Stage-0 and Stage-0b GPT-2 small runs, including memory measurements and layerwise curves. - **Read:** `settling-field-lab/interp/RESULTS.okf.md` and `settling-field-lab/interp/RESULTS-0b.okf.md`. - **Not established:** a general settling-depth law, an easy-versus-hard separation, or behavior in larger models. - **Next falsifier:** exclude the final point, normalize the curve, and compare raw and tuned lenses. ## An adaptive reader learned where to look. It did not learn how to scale. Source: https://logicalworks.ca/research/adaptive-reader-did-not-scale Published: 2026-08-30 Status: Negative result Authors: Logical Works Research ## The result A recurrent controller learned to query a synthetic world through a hard one-cell read interface. On 16-cell pointer chains it reached 100% accuracy at depths two and four, and 98% at depth eight. On direct lookup, a two-phase training schedule reached 100% accuracy with exactly one mean read at both 16 and 64 cells. That is the information-theoretic minimum for a one-cell lookup. The same model did not scale. Accuracy fell to at most 4% at 128 cells and to roughly zero at 256 cells. Depths 16 and 32 also failed. The first cycle localized two causes: address bits above the trained range never varied, and the learned policy horizon stayed tied to the ten-step training schedule. A second experiment attacked the address-bit explanation by drawing identifiers from the full 16-bit space. It made the task harder in distribution instead of restoring extrapolation. At 16 cells, accuracy ranged from 0.3375 to 0.50 across three seeds. At 64 cells it ranged from 0.0125 to 0.0525. At 1,024 cells, nearly every grid cell was zero. ## What the controller actually learned The controller was a GRU with a learned query head. Each query selected one cell by forward hard argmax; the backward pass used a soft straight-through estimator. Inputs contained the question bits and prior reads. Tests excluded answer, chain, difficulty, and hop-count side channels. The positive result was not a fixed schedule in disguise. Corrupting a read changed the next query in 99.5% of examples. A mid-chain pointer edit redirected trajectories in 70.5% of cases. Irrelevant growth from 64 to 256 cells changed the read count by only -0.005. The policy therefore learned conditional information acquisition. It also learned a brittle addressing procedure bound to the training distribution. Those are separate facts. ## Why the first headline was too strong The first cycle had four limits. It used one seed, the full-access baseline had 30% more parameters than the fixed reader, the XOR tasks were too hard for every arm, and the stronger Transformer baseline was never run. The seed follow-up introduced another confound. Seed 0 trained for 15,000 steps; seeds 2 and 3 trained for 8,000. The large seed-0 gap at 64 cells mixes seed sensitivity with training length. Only the seed-2 and seed-3 comparison isolates seed variance, and those two runs agree in their failure. The bounded claim is clear: stochastic gradient descent can discover an adaptive policy through a hard read channel on tasks constructed to reward adaptivity. The current controller does not establish size or depth generalization, and it does not establish an advantage over a compute-matched Transformer. ## What changed next The next experiment should separate representation from optimization. Keep random identifiers, match training steps across seeds, add a Transformer reader with the same compute ledger, and report confidence intervals. If the controller still fails at trained sizes, world-keyed content addressing is the immediate bottleneck. If it fits trained sizes but fails beyond them, the extrapolation problem remains. Low read count is not useful when the policy reads the wrong cell. A compute ledger must travel with task success. ## Evidence ledger - **Executed:** direct lookup, pointer-chain, corruption, redirect, irrelevant-growth, identifier, and seed experiments. - **Read:** the primary result ledger and round-four report in `settling-field-lab`. - **Not established:** out-of-distribution scale generalization or superiority to a compute-matched Transformer. - **Known confound:** unequal training steps between seed 0 and seeds 2 and 3. ## Why a global waveform memory loses its compute advantage on exact worlds Source: https://logicalworks.ca/research/waveform-memory-falsification Published: 2026-08-29 Status: Formal result Authors: Logical Works Research ## The proposed composition The original program combined a JEPA-style latent objective, a continuous coordinate-addressable world representation, and a recurrent query controller. It predicted that the controller could retrieve less information as the world grew while maintaining accuracy on exact synthetic tasks. The formal core proved a narrower statement. A reasoner cannot recover an arbitrary uninspected cell from a trace that does not contain it. The Lean file also makes the compute ledger monotone: reading more cells cannot be recorded as cheaper. These are interface and accounting theorems. They do not prove that training converges or that a waveform is efficient. ## The rank-cost conflict Take a global basis with `K` coefficients over `n` independent cells. Exact representation of all possible cell assignments requires rank at least `n`; otherwise two worlds collide in the representation while differing at a queried cell. Once `K >= n`, evaluating one point in a global sum touches at least `n` coefficients. The exact representation has recovered the same linear per-query cost it was meant to avoid. The escape route is locality. A compactly supported multiresolution basis can evaluate a point in constant or logarithmic support. At the exact-world limit, that representation approaches an indexed lookup. A waveform can still help on worlds with compressible multiscale structure, but random independent cells contain no such structure. ## Why the JEPA objective disappears on the first benchmark For independent random cells, the mutual information between context and a held-out target is zero. A latent prediction objective cannot learn a relationship that the data distribution does not contain. The constant predictor is optimal for that term, and anti-collapse machinery can prevent representational collapse without creating predictive signal. The first benchmark therefore disabled two proposed mechanisms at once. The global field lost its cost advantage under exactness, and the JEPA term had no predictive information to use. A result from that benchmark could test the harness for leaks, but it could not fairly test the intended structured-world hypothesis. ## The component that survived Adaptive querying survived at the theorem level. On a two-hop pointer task, an adaptive reader can follow the first pointer and then the second for two reads. Any fixed schedule that omits a possible destination fails on some world. The asymptotic separation comes from conditional choice, not from the waveform representation. Subsequent experiments confirmed that an adaptive policy can be learned in distribution, while also showing that its current implementation fails to generalize across identifiers and scale. That split should shape the repaired program: preserve hard-read adaptivity, replace the global basis with local or indexed structure, and use JEPA only on data where `I(context; target) > 0`. ## Evidence ledger - **Formal:** interface and monotone-ledger properties in `formal/WaveformJEPA.lean`. - **Argument:** the rank-cost and zero-mutual-information analyses. - **Executed elsewhere:** the adaptive query controller experiments reported in the companion negative result. - **Not established:** training convergence or an efficient waveform representation for exact random worlds. ## We counted 264,224 agent turns. The causal labels are not ready to publish as causal findings. Source: https://logicalworks.ca/research/auditing-agent-failure-corpus Published: 2026-08-28 Status: Method audit Authors: Logical Works Research ## What exists in the checkout The session summary contains 2,236 sessions and 264,224 assistant turns: 201,440 from Claude Code, 41,764 from Codex, 20,710 from OpenCode, and 310 from Gemini Antigravity. Recomputing the totals from `multidimensional-session-summary.jsonl` matches the report. The full turn-level file is not present as data in this checkout. `multidimensional-failure-corpus.jsonl` is a 134-byte Git LFS pointer whose declared object size is 281,254,500 bytes. Any public analysis of turn text must fetch and verify that object first. Session-level aggregates can be inspected now; the underlying turn excerpts cannot be re-audited from this checkout. ## Why 96.8% does not mean users are vague The report assigns 255,674 turns to `AUTONOMOUS_STEP_CONTINUATION`, or 96.8% of all turns. The builder uses that label when `human_prompt` is empty. Empty prompts occur naturally when one user request produces many assistant tool subturns. The denominator is assistant turns, not distinct user instructions. The report then classifies nearly all of those empty-prompt subturns as `HUMAN_UNDERSPECIFIED_INPUT`. This says more about transcript segmentation than human communication. A publishable estimate of user underspecification needs a user-turn denominator and session-aware pairing. ## The labels are heuristic, not causal The builder marks a prompt vague when it is short, lacks punctuation-like technical characters, or matches a small phrase list. It marks spiralling when an apology and a repeated edit co-occur, when a shell error is followed by a tool-name retry without a diagnostic tool, or when three edits touch a previously edited file. A retrieval miss is then inferred when a prompt contains a repository keyword and the same turn has a shell error or spiral flag. Those are candidate signals. They do not establish that missing retrieval caused the failure. The output field `causal_pattern` overstates what the procedure measures. The clean public names are `adjacent_pattern` or `heuristic_label` until a human-coded sample provides precision, recall, and inter-rater agreement. The current report also contains a denominator typo: the headline table totals 264,224, while a later sentence says 264,205. The distributions recompute to 264,224. ## A publication-grade repair First, reconstruct exchanges around user instructions rather than assistant subturns. Keep tool calls as events inside an exchange. Second, sample each proposed label by framework and register. Have two reviewers label the same blinded examples, report agreement, and adjudicate conflicts. Third, split temporal association from causal interpretation. A tool error followed by user anger is an observed sequence; "AI failure induced anger" requires a stronger annotation protocol. Finally, publish privacy and exclusion rules before examples. This corpus comes from private work sessions. Aggregate counts do not authorize transcript publication. A useful research release can expose schemas, label definitions, synthetic fixtures, and validated aggregate results without releasing personal prompts or secrets. ## Evidence ledger - **Recomputed:** session and assistant-turn totals from the available summary JSONL. - **Read:** the report and corpus-builder rules. - **Unavailable:** the full 281 MB turn-level Git LFS object in this checkout. - **Heuristic only:** vagueness, spiralling, retrieval-miss, and causal-pattern labels. - **Privacy boundary:** no private prompt excerpts are published here. ## A neural network's privileged self-knowledge may be operationally empty Source: https://logicalworks.ca/research/operationally-empty-self-knowledge Published: 2026-08-27 Status: Position paper and pre-registration Authors: Logical Works Research ## The question Suppose a model carries internal state between interactions. Persistence alone does not make that state a self. A cache persists. A recurrent hidden state persists. A parameter vector persists. The research question is whether the model has privileged access to its own future state in a way an external observer with the same information and capacity does not. The current paper proposes three clauses: persistence under zero-input dynamics, privileged access to future internal state, and causal load under steering or ablation. The third clause distinguishes consequential state from decorative state. The second is intended to distinguish self-knowledge from ordinary memory. ## Three targets for privileged access Predicting future outputs from current state measures task competence. A useful state should predict outputs; that does not show that it represents itself. Predicting future internal state from current internal state measures trajectory autocorrelation. Smooth recurrent dynamics can score well without any representation of self. Counterfactual self-response is the strongest candidate: predict how the model's own state would change under an intervention it did not receive. That comparison breaks on information equality. If the model sees the intervention and the observer does not, the advantage is access to information. If both see the intervention, the observer is modeling an observable dynamical system. If the model retains an advantage because its mechanism is wired into the predictor, the comparison no longer matches architecture. This is an argument, not a formal theorem or executed experiment. Its result is a measurement warning: current operationalizations do not isolate a residual called privileged self-knowledge. The paper therefore abandons the strong claim and keeps a narrower identity claim, treating measurable self-state as compression of interaction history. ## What remains empirical The architecture question survives. Compare three capacity-matched systems under one objective: a stateless baseline, standard recurrence, and an idle-invariant settling state. Let `g_R` be the transfer gain from recurrence over stateless processing, and `g_S` the additional gain from settling state over recurrence. The pre-registration requires `g_S > 0` with a 95% confidence interval, a matched-magnitude noise control, and an ablation that removes the gain. It also scopes the test to agentic tasks where the model's own outputs re-enter its future context. Passive text prediction does not create the same self-versus-world distinction. This experiment would not prove consciousness or an inner life. It would test whether an idle-invariant state adds a causal transfer mechanism beyond ordinary recurrence at matched capacity and memory. If ablation leaves the gain intact, the settling state was correlated with the result rather than responsible for it. ## Why publish the argument before the run Pre-registration removes an easy escape. Without fixed controls and thresholds, any improvement can be redescribed as evidence for a self-state after training. Naming the failure condition in advance turns a philosophical phrase into an architecture test. The next owed measurement is smaller: establish the persistence and causal-load baseline on GPT-2, then run the model ladder. Until those numbers exist, the article remains labeled as a position paper and pre-registration. ## Evidence ledger - **Read:** the position paper and decision record. - **Argument only:** the information-equality problem in privileged-access comparisons. - **Pre-registered:** three matched systems, confidence interval, noise control, and ablation requirement. - **Not executed:** the density, persistence, and causal-load model ladder.