# What we can honestly claim so far _evidence log · updated 2026-09-16 · criteria: code review v0.6 · education v0.7 — research-based, pre-field-data · observations, judgements and reflections are distinguished · no conversations stored_ Evidence for how an agent works, evidence that the service connects the steps, and evidence of later improvement answer different questions. Each result below states its scope and limits. **Last updated 2026-09-16.** The comparative experiments were run in August 2026 (C: 19 August, D: 22 August). A local end-to-end test of an organisation-authored domain passed on 14 September. These are internal experiments and functional checks, not external field outcomes. The field-course count below retains its own observation date. Read together: two experiments on our own code show that a recipe's contract changes what the same agent produces (7/8 vs 0/8), that capability arrives on its own (tools took a packaged reviewer from 2/8 to 6/8), and that a belief of ours fell when tested (steps did not help). The September end-to-end test adds evidence that the service connects a run to a later report and a revised domain edition under a controlled clock. It does not supply real-world outcomes or show that a method became better. This page reports no external field results establishing that improvement; the question remains open. - **end to end — an organisation-authored domain, from first edition to revision.** On 14 September 2026, the S-stage test passed through the product's ordinary tools against both an in-memory repository and local PostgreSQL. It created and stamped a domain, authored a recipe, submitted work, moved the test clock, discovered the due follow-up, read a brief for that run, reported against its criteria and inspected the history. A revised domain edition then applied to the next run while the earlier run retained its original contract. Duplicate reports, partial criterion coverage and refusal of a submission against an outdated criteria edition were also checked. This demonstrates the connected behaviour in the tested cases. The outcomes were test data and time was controlled: it does not show a customer returning on a real clock, better criteria or improved results. The S and S+ work was deployed on 16 September 2026 — database migrations 049 to 053 and the MCP server — with limited checks afterwards against the running service. The deployment puts the implementation behind these local tests into production; it does not repeat the full test there or add evidence about outcomes. _source: notes/2026-09-14-steg-203-slutprovet.md; notes/2026-09-15-etapp-s-stangd.md; notes/2026-09-16-deployen-049-053.md; packages/mcp/test/closed-loop-suite.ts_ - **0 — courses completed in the field.** Our recorded count was zero as of 10 September 2026. This historical count has not been remeasured for this publication update. The planned first course was a child's mathematics course in a separate organisation, with a later assessment after completion. The internal experiments and local functional test below do not count as completed field courses or as external field outcomes. _source: notes/2026-09-10-mjuk-release-deployen.md; notes/2026-09-08-deployen-029-048.md; evidence log snapshot dated 11 September 2026_ - **7/8 vs 0/8 — recipe vs our own reviewer prompt, on a real hole.** Same model, same code change, two instructions. Reviews written under the recipe's contract found a real, SDK-verified hole in our production code in seven runs of eight; the review prompt our build loop was using at the time found it in zero of eight. Attests that the contract changes what the same agent produces. Does not attest generalization: one codebase, one change, one family of models. _source: research/experiment-c (2026-08-19), 16 reviews, 560 blind ratings; not pre-registered_ - **0/16 — verdicts in our contract’s format from packaged tooling.** Across sixteen runs of packaged review tooling, none supplied a verdict in our contract's format. Those approaches had not been given that contract. This measures conformance to our requested format, not worse judgement or a comparative win. _source: research/experiment-d/SLUTSATS.md (M2)_ - **2/8 → 6/8 — the tool effect: capability arrives on its own.** Handing the same packaged reviewer its tools moved defect-finding from two runs in eight to six. Model capability improves without us — which is exactly why a recipe must not script the how. The recipe's job is the part that did not move on its own. _source: research/experiment-d/SLUTSATS.md (M1, A1→A2)_ - **the steps lost — a pre-registered belief of ours, falsified.** We believed a review should be composed in steps. We pre-registered the question, the arms, and what every outcome would mean — including the row that said our intuition was wrong — then measured. The three-step variant did no better anywhere that separates the two, and worse on the sharpest measure (consequence stated: 3/8 vs 6/8). The recipe keeps its single step, and the meta-recipe now says one step is legitimate here. Attests that the loop learns, not just that it works. _source: research/experiment-d/{FORREGISTRERING,SLUTSATS}.md, decided 2026-08-22_ - **±1 = 1.00 — grader agreement within the two recipe arms.** The 16 reviews in the two recipe arms of experiment D were rated blind: 7 criteria and 5 grading sessions per criterion, giving 560 ratings. Agreement within one point was 1.00. The other arms were assessed on defect-finding and conformance, not through these rubric ratings. Agreement does not establish that the graders were correct. _source: research/experiment-d/FORREGISTRERING.md (§7); research/experiment-d/SLUTSATS.md_ - **v0.6 / v0.7 — criteria versions in force (code review / education).** Every version change is a written entry: what changed and why, dated. Both documents state their own standing plainly — grounded in research and our own measurements, with no field outcome data behind them yet. No criterion has changed because field data demanded it; the versions above were sharpened by the experiments on this page, by hand. _source: recipes/*/eval-criteria.md (version history in the documents; both v0.6 and v0.7 dated 2026-08-27)_ - **signed — the first field analysis, pre-registered before any data.** What the first analysis of field outcomes will ask, which rows it will count and what every answer would mean was written down and signed on 7 September 2026, before the analysis has data. It separates association from causation in advance, and it excludes the project's own organizations from the population it counts, so the first number cannot be one that measures us. Attests that the question is fixed before the answer; it attests nothing about the answer. _source: research/2026-09-07-preregistrering-forsta-analysen.md; notes/2026-09-07-preregistreringen-signerad.md_ - **1 — organization-authored domain stamped at the bench.** On 8 September 2026 an organization's own agent authored a domain — a kind of work, its criteria, its clock — and a person at the bench read two trial judgments made under it and stamped edition v0.1. One run has been made under that domain since. This is the bench in use, not an experiment: a hypothesis with a name on it, and no evidence yet that the criteria are good — that is what the later outcomes are for. _source: notes/2026-09-08-forsta-bankarendet-system-assessment.md_ Loop: `meta-recipe → recipe → delivery → signals → meta-recipe` What is sent, processed and retained depends on the operation. Method documents are stored. Submitted reviews and work in your own domain are sent to labbel and processed, but their bodies are not retained as documents or in their traces. Reports can include judgements and sanitised free-text reflections; sanitisation does not guarantee anonymity. This is a description of data received by labbel, not permission to share it with another organisation. Internal error logs are a separate data surface, described on the security page. The signal type names below come from the service's taxonomy. For organisation-authored domains, criterion_status reports against the individual criteria of the domain edition used for a run, while review_reflection carries reflections. Two signal types do not mean two assessment criteria or the absence of fresh-session assessment. Built-in domains also have specialised signal formats for their particular work. [Read the data and logging limits on the security page.](https://labbel.ai/security) **A course.** Work and method documents: Teaching conversations and the learner's original work are not collected as a conversation history. Recipes and course steps are sent and stored, along with permitted progress records. Keep learner transcripts and personal data out of submitted notes. Reported signals: whether a checkpoint was passed, how well the learner predicted their own result, what they could still do with it later, how the sessions were spaced, that the course was completed, the level reached, how much help was needed over time, and what the teaching agent thought of the recipe. _types: checkpoint_result · calibration · transfer_outcome · session_pattern · completion · bloom_achieved · independence_trend · recipe_reflection_ **A code review.** Work and method documents: The review is sent to labbel for structural checks. Its body is not retained as a document or in the trace. Keep the original code, change and review in your own records; the trace retains checks, versions and selected declarations. Reported signals: that a review was issued, what happened to the change afterwards — a fix, a revert, an incident or quiet days — what happened to each finding, each criterion met or not, whether the findings were accurate, whether the review was useful, and what the reviewing agent thought of the recipe. _types: review_issued · change_outcome · finding_outcome · criterion_status · finding_accuracy · review_usefulness · review_reflection_ **Your own domain.** Work and method documents: Your domain criteria, recipe and submitted bench declarations are sent and stored. Work submitted under the recipe is received and processed, but its body is not retained as a document or copied into the trace. Keep the original work in your own records. Reported signals: criterion-status reports and the performing agent's reflection on the run. The reflection is sanitised free text, not a copy of the original work. _types: criterion_status · review_reflection_ Artifacts (supporting artifacts are not yet published · file review and human reading remain outstanding): - experiment C — arms, reviews, ratings, analysis - experiment D — pre-registration, conclusion, harness, raw data - the recipes and criteria the experiments ran against Contact: hello@labbel.ai *Publication snapshot: web-eef8292b2fbc2ddd*