Platform · WorkloadsTen program categories

Infrastructure across AI data programs.

The categories below describe kinds of program Largwit supports and what each one produces. They are not fixed products: a program is configured around the requirements of the engagement it belongs to.

Every program shares
Configure
The project and its specification
Staff
Against the program's own bar
Run
Under its own controls
Watch
Production, as it happens
Deliver
And account for it
01Program categories

What can be run, and what it produces.

Each category differs in what a task is, who is qualified to do it, what counts as evidence, and what the finished item looks like. The operation around it is the same in every case.

Programswhat each one is for, and what leaves
Model evaluationStructured assessment of model output against criteria the customer defines, run as a repeatable program rather than a one-off exercise.Produces scored assessments with the evidence behind each judgement.
RLHF and SFT dataDemonstrations, corrections and preference material produced at contract volume under a defined quality standard.Produces training material with provenance for every item.
Agent evaluationMulti-step behaviour assessed step by step, so a trajectory that reaches the right answer by an invalid route is not recorded as a pass.Produces step-level verdicts with named failure classes.
Expert dataDomain programs in fields where a general workforce cannot hold the bar, staffed against qualifications the customer defines.Produces specialist judgement, attributable and reviewable.
Multimodal annotationText, image, audio and video handled within one operation and one record rather than split across separate tools.Produces consistent annotation across modalities in one dataset.
Research evaluationInvestigative programs where the finding and the evidence supporting it are both part of the deliverable.Produces findings with sources, traceable to what was examined.
Document-based evaluationJudgements grounded in source documents, where a claim is only acceptable if it can be tied to the material it came from.Produces cited assessments checkable against the source.
Preference dataComparative judgement between candidate outputs under criteria the program defines.Produces preference pairs with the reasoning recorded.
Response evaluationAssessment of individual model responses against a rubric, with failure classes recorded rather than described in prose.Produces rated responses and categorized failures.
Custom programsPrograms that do not match a standard category, configured around the specification the engagement is held to.Produces whatever the contract defines as the deliverable.
02The shared operation

The program changes. The infrastructure does not.

What does not differ between the categories above is the operation around them: configuring the project, staffing it, running quality controls, watching production and delivering the result.

That is the part Largwit provides, and the reason a second program costs less to stand up than the first.

01Configure

The project and its specification, approved

02Staff

Against the program's qualification bar

03Run

Under the program's own controls

04Watch

Production, computed as it happens

05Deliver

Accounted for, traceable to what produced it

03Start here

Bring us the programme you are about to build.

Largwit is deployed through direct engagement. Tell us what you are running and we will show you the platform against it, rather than against a generic demo.