Reference vs. candidate · sealed holdout

Prove the cheaper workflow keeps its quality.

Kourum compares one reference workflow with one candidate through a vendor-diverse panel, deterministic checks, and a Final Exam that revision logic cannot read.

01

Define the task

Plain text or JSON. Uploaded documents are never parsed on the server.

Describe the inputs, required output, constraints, and acceptance boundary.
Up to 5 files, 1 MB each. PDF/DOCX extraction remains disabled until its library TBD is resolved.
02

Choose the pair

Choose from Kourum's verified, priced model catalog. Same-vendor pairs are allowed.

Reference workflow

1 · Vendor
2 · AI Model
Verified pricing and provider support are required.
3 · Reasoning effort
The only orchestration-level consistency lever; sampling fields are omitted.

ConfiguredOpenAI · GPT-5.6 SolHigh

The candidate instructions may be revised; the reference remains fixed.

Candidate workflow

1 · Vendor
2 · AI Model
Verified pricing and provider support are required.
3 · Thinking
Minimal disables thinking; Thinking enables it for this model.

ConfiguredAnthropic · Haiku 4.5Minimal

The candidate instructions may be revised; the reference remains fixed.
03

Add evaluation cases

Cases are split once into Study, Practice, and the sealed Final Exam at creation.

Case 1
Required
Optional
Defaults to General
04

Set the quality gate

Deterministic checks run before rubric aggregation; any failure blocks a pass.

Deterministic checks
Node 24 or Python 3.13 only. Each check runs in a fresh, network-denied sandbox.
Optional human anchor calibration
Optional array of human-approved dimension scores keyed to dataset case IDs.

Funded runs need NEXT_PUBLIC_TURNSTILE_SITE_KEY before launch.

Start with a validated Study roundThe Final Exam remains unread until you explicitly proceed.