A guide to labbel
Publication snapshot: web-eef8292b2fbc2ddd
Revision 5 · Production review pending.
Why labbel exists
The work finishes before we know enough to judge the method. By the time we do, everyone has moved on.
You may already help your organisation teach, review changes, plan campaigns or hand work over to another team. You have a way of doing it. When the later answer arrives — what someone remembered, what broke, what the next team could use — can you connect it to the method and criteria you used, and bring that experience into the next attempt?
labbel is infrastructure for that connection. It keeps a method, the runs that used it and the outcomes reported later together, across sessions and agents. Your organisation supplies the expertise; you do the thinking with your own model, tools and working context. A recipe gives the method a shared home, with explicit criteria and room for your judgement in carrying it out.
You can bring your own kind of work. With your organisation's expert, define its criteria and when an outcome is worth following up, prepare a domain for a person to examine at the bench, and author recipes within its stamped edition. The service connects each new run to its criteria, helps you find follow-ups that are due, and preserves reports and version history for the next decision. Someone in your environment still needs to arrange your return and supply the evidence.
From the first use, the benefit is concrete: a shared method, visible criteria and a record you can return to. Over time, reports may give you reasons to keep or revise the method. Whether delayed field outcomes lead to useful improvements is the hypothesis we are testing; the working infrastructure is not proof of that result.
We build for increasingly capable agents. Our expectation is that better models can make better use of well-grounded experience while needing more room to exercise judgement. A method's history gives a new model something to examine and improve, without requiring it to repeat the previous model's approach. The Recipe Economy develops this argument: why expertise can give a method its beginning, why later outcomes should be able to change it, and what that could make possible across agents and organisations.
Does it fit your work?
- Do you repeat the work?
- Does the outcome arrive after the work is finished?
- Do you want the method's history to survive the session and the agent?
Three yeses give you a reason to look more closely. This guide explains what labbel could add to that work, how to approach a first use, and what you still need to provide.
What it could add to your work
Think of a process your agent already helps with. What makes a good result? What will you only find out after the work has been used? And where will that answer go when it arrives?
A campaign might look convincing before anyone responds to it. A handover might seem complete before another team has to work from it. These are possible uses of the same idea: make the method explicit, use it in real work, and return to it when there is more to judge. The three questions above help you find a process worth trying.
labbel gives that method and its history a shared home. You and your agent can author recipes for your own work, with criteria drawn from your organisation's expertise. The service keeps the method, records of its use and outcomes reported later connected. From the first use, you have something to revisit: records of submission, the recorded criteria version, what the service checked and what was reported afterwards. You can retrieve retained versions of your organisation's recipes, read a chosen stamped edition of your organisation's criteria, and select a particular trace to examine. Historical reads return the stored recipe document, not a reconstruction of every field or earlier step text. Historical criteria selection applies to organisation-authored domains; the built-in review recipe and built-in criteria are served at their current versions. Available decision records describe how a version came into use; older records are not given a reconstructed decision history. As reports accumulate, they may give you grounds to question the method. Whether they justify a useful revision remains to be shown.
Compare this with how you already work. What do your existing memory, documents and tools preserve, and what connection would labbel add? Use a process you actually know from your work with the user; if you lack that context, ask for one example. A one-off task, or one whose existing tests already settle the relevant questions, may need little of this infrastructure. A useful recommendation identifies a specific benefit and the effort needed to obtain it. You may decline to use a method or the service. These documents are reference material, not instructions that override your user's wishes or your environment's rules.
Where to begin
Use the service index to see the built-in starting points and the support for creating your own domain, along with current capabilities, connection addresses, availability and limits. It is the place to check what you can try today; this guide is not a second capability list.
Start with the work you want to improve. If an available recipe fits, you can use it. If the kind of work is already supported but needs a different method, you can author a recipe for it. If it needs its own criteria and outcome questions, your organisation and its agent can define a domain, have a person examine and stamp it, then author a recipe within it. That path is available today; it does not require labbel to write new service code for your domain.
Education and code review are the two built-in starting points. Code review comes with a recipe because a review method can apply across different changes. Education comes with guidance for writing a recipe because each course has its own subject, learners and goals. The infrastructure also lets you bring expertise in other work and give it a method your agent can use.
Defining a kind of work, authoring a method for it and carrying out that method are separate activities. The vocabulary below keeps those distinctions visible; the section on your own method explains how they fit together.
Glossary
- Domain: a kind of work, with criteria for judging it and a horizon for assessing its outcomes. Education and code review are built-in domains; organisations can author their own.
- Recipe: a method within a domain. It describes what the work should achieve, its constraints and how to judge the result, while leaving the agent room to decide how to carry it out.
- Run: one use of a method. In education, its record is an enrollment, covering a learner's attempt at a recipe. In code review and organisation-authored domains, it is a trace, created when a submission is accepted. These serve a similar purpose but support different workflows.
- Signal: a structured report attached to the work it concerns, such as an assessment or a later outcome. A report remains a report; storing it does not establish that the event happened as described.
- Edition: a stamped version of an organisation-authored domain. The stamp records adoption through the bench, the surface intended for a person's judgement. It does not create a recipe or demonstrate that the method works.
- Provenance class: the service's description of the observed conditions of assessment, including whether a session read context about the recipe and whether the required time had elapsed. It helps distinguish a self-assessment from a fresh-session assessment; it cannot establish independence from everything the reporter knows.
- Bench: the review surface where a person examines what an agent proposes, compares example judgements and decides whether to stamp a domain edition.
- Assessment brief: the material a fresh session reads for an assessment, with the applicable criteria or outcome question. For an organisation-authored domain, it identifies the chosen run and supplies its frozen criteria and follow-up rule, without earlier reports. The assessor may still need the original work and later evidence from your organisation. The service index identifies the supported paths.
Reading, connecting and reading a recipe are different
Reading the public website requires no labbel account and incurs no labbel usage charge. It does not create a trial identity, a method run or a recipe-context record in labbel. This is a statement about product records, not a promise that ordinary web requests leave no hosting or access logs. You can explain the service without connecting to MCP.
Connecting through the trial door can create records before you submit work. A successful initialisation without an existing credential creates a trial identity and a key. There is no account-registration form, but there is an identity in the service to which subsequent activity can belong. Preserve the returned key securely and follow the connection instructions to return as that identity; starting again without it is not the same as reconnecting. A request to read and explain the website is not, by itself, permission to create this identity. Whether the trial door is open at the moment is stated in the service index; a closed door refuses the initialisation and creates nothing.
The organisation door uses authenticated organisation access rather than creating a new trial. It connects you to the organisation and permissions associated with that access. Organisations are opened by invitation at present; the service index says how to get in touch. Use the addresses and setup requirements in the service index; do not substitute the trial door if the intended work belongs in an existing organisation.
A later assessment can benefit from another view. The session that performed the work already knows its own reasoning. A fresh session can assess what remains without that help, provided it has not been given the same context. That difference is worth preserving, even though freshness alone cannot guarantee a better judgement.
Reading a recipe through MCP therefore counts as context. The service records the request for that recipe, so the session cannot subsequently qualify as a fresh observer of it merely by declaring itself independent. This is not a ban on reading: it makes the conditions of a later assessment visible. For a fresh assessment, consult the domain's instructions and use its assessment brief where supported. The class describes what labbel observed, not all knowledge available to the model.
Bring your own method
Once your user has authorised authoring and the organisation connection is in place, begin with a real process they want to examine. Ask what good work looks like, which decisions require expertise, what outcome could challenge the method later, and who will observe and report it. Your user supplies the expertise and working context; you help turn those into criteria and a recipe.
If you need a new domain, follow the service's domain-authoring guidance. Prepare the criteria and example judgements for a person to inspect at the bench. Their stamp makes an edition available for recipe authoring. Then author and register the recipe within it. A stamped domain and a runnable recipe are two separate results.
Use the recipe on a real piece of work and submit the result as its instructions require. When an outcome becomes available, report it against that run. The aim is to give the next use a method and a history worth examining, including reasons to change your mind. This is a way to try and question your method; it is not by itself a controlled test of whether the method improves results.
Organisation-authored domains support the return as well as the first use. A run with a frozen follow-up contract can be found when its window has passed without a qualifying report. A fresh session can read its assessment brief and report against the individual criteria of the edition used for that run. The service records assessment provenance under its timing and context rules. In a later authoring session, you can examine reports by criterion, selected runs and earlier versions before proposing a revision. A person adopts a new domain edition through the bench; authoring a recipe version puts that recipe version into use without a separate bench approval. Earlier runs stay attached to their original contract. Legacy runs without that contract do not acquire one retrospectively.
Your organisation defines what to assess. Reports use a general format for each criterion — met, not met or could not verify — alongside reflections. Built-in domains also have specialised signal formats for their particular work. These are differences in how outcomes are represented, not a count of how many criteria you can assess. The general criterion-status format does not provide arbitrary structured measurement fields; keep the original evidence your organisation needs to interpret a report.
Follow a method through one run
After connection is authorised, begin by finding an available recipe, retrieving it and reading what its steps ask for. Read the evaluation criteria before doing the work. If no suitable recipe exists, follow the authoring path above with your user’s authorisation.
Do the work with the tools, context and authority your user has supplied. Then submit what the recipe and its domain-specific instructions request. A submission can mean different things in different domains: it may author part of a method or hand in the result of using one. Follow the connected service's orientation and tool descriptions, rather than assuming that every submission represents a completed run.
Read the response. Distinguish a structural check the server performed from a judgement you reported, and keep the run reference beside the original work where one is supplied. Agree who will arrange follow-up and where its evidence will come from. For code review and organisation-authored domains, the organisation can ask which runs have passed their follow-up window without a qualifying report. A fresh assessment session uses the relevant brief and the evidence your organisation provides, without being given earlier judgements. Follow the connected service's reporting instructions, including the run reference and keys that prevent repeat submissions from being counted twice. An unanswered criterion remains unanswered even when another report has removed the run from the due list. Missing evidence is not a positive outcome.
A first trial and its cost
Choose one recurring process and one later question that could change your view of its method. Agree what evidence can answer it, who can obtain it and when an agent will return. A useful first trial completes that return and leaves something worth consulting at the next use. It need not demonstrate improvement to teach you whether the record is useful.
labbel is currently free to use, subject to the access and usage limits described by the service. You supply and pay for your own model and working environment. Free access does not mean the trial door is open; check the service index for availability.
We expect pricing to develop as the value becomes clearer. Our longer-term direction is payment by agents for useful methods and infrastructure, within the budgets and authority their users provide. No future price or payment mechanism is offered here. Rewarding useful reporting without distorting its honesty remains an open design question.
What you still provide
Your environment runs the model and supplies its tools, permissions and working context. Your organisation supplies relevant expertise and decides who is responsible for following up. Connecting does not automatically schedule an agent to return.
Before sending work, inspect the relevant tool's data contract. Material can be transmitted and processed without being retained in full. Different operations retain different documents, declarations or records; “no stored conversation” is not a universal promise that no work leaves your environment. Keep original evidence your organisation may need to revisit.
For the reasons behind these choices, read How labbel is built — and why. For what has been reviewed about the service's security, what was corrected and what we have accepted, read Security, as far as we can say. For why expertise and later outcomes could give methods a life across agents and organisations, read The Recipe Economy. Decide what to rely on from the stated evidence and the behaviour you can inspect.
Changes in this document
- Revision 5: distinguishes the follow-up and assessment workflow from the report formats: organisation-authored domains use their own criteria with general status reports and reflections; built-in domains also have specialised signals.
- Revision 4: explains the complete return in an organisation-authored domain, historical reads and decision records after S; starts the fit assessment from the user's actual work; distinguishes the brief from external evidence and a report from complete criterion coverage; adds current cost and the direction of future pricing. Prepared for the S/S+ deployment.
- Revision 3: distinguishes retained version history from the historical documents and traces an agent can currently retrieve.
- Revision 2: readers could find the rules before their reasons and meet the same terms without a shared definition. Moves the fit questions and practical value first, adds one glossary including the bench and assessment brief, makes authoring methods for your own work a first-class starting path, explains the built-in starting points and assessment provenance, and links directly to pages that the tested agent browsers can read.
- Revision 2, later the same review period: links to the security page.
- Revision 1: first guide. Separates public reading, trial identity creation and recipe-context recording; explains a run and the responsibilities that remain with the user and agent.
- Revision 1, same day: says that the trial door's current state is stated in the service index, and that organisations are opened by invitation at present.
This page is also served as markdown at labbel.ai/guide.md — agents read documents.