Solo & AI

Extract judgment, not answers: using large AI models well

July 26, 2026 · 9 min read

In short — Large AI models (Claude Opus, GPT-4o…) aren’t meant to generate content on a loop: they’re meant to encode your judgment into reusable standards. Your cheap models then apply those standards, at scale, without costing you a fortune.


You’ve probably already opened Claude Opus or GPT-4o to draft an email, generate an article, or summarize a document. It’s the natural reflex: the most powerful model for the most visible task.

It’s also the most expensive and least scalable way to use these tools.

The real use of frontier models — the Claude Opus, GPT-4o, Gemini Ultra of this world — isn’t producing content. It’s encoding judgment. Writing the standards your cheap models will then apply, hundreds of times, without costing you much.

This pattern has a name in the way I work: the teacher/student pattern. The frontier model writes the textbook. The cheap models do the exercises.

Here’s how it works, why it changes everything for a solopreneur, and how to set it up today.

The irreversibility test: what can’t a cheaper model redo tomorrow?

Start with this blunt question: if you replaced your frontier model with a 10x cheaper model tomorrow morning, what would be lost?

Not much, in most cases. An article generated by Claude Opus can be generated correctly by Haiku or Flash, with a precise enough prompt. A client follow-up email, same thing. A feature description, same thing.

What a cheap model can’t easily reproduce is the ability to ask the right questions about your own work. To spot the tensions in your logic. To turn a fuzzy intuition — “this article isn’t good” — into precise, testable, reusable criteria.

That’s judgment. And that’s what only a frontier model can extract reliably today.

Say you have a strong intuition about the quality of your articles. You know when a text is good. But you couldn’t explain it to a cheap model in 200 tokens. A frontier model, though, can help you externalize that intuition: it will ask questions, iterate with you, and produce a standards document that any model — or any human — can apply mechanically.

The result of that session is irreversible in the good sense of the word: once the standard is written, you no longer need the frontier model for that task. It did its job. It can move on.

The answer: what we extract is the standard, not the content

Content is perishable. An article goes stale, a sent email disappears, a product description gets updated. The standard lasts.

A well-written standard is an operational rule that can be:

  • verified by a script
  • applied by a cheap model
  • handed to a contractor
  • audited by you in 30 seconds

This isn’t a vague style guide. It isn’t “write with authenticity” or “adopt a conversational tone.” It’s a list of binary criteria: true or false, present or absent, within bounds or out of bounds.

Here’s the key distinction: a frontier model is good at going from fuzzy to binary. That’s exactly where it earns its cost. Not for generating 50 articles, but for spending 2 hours with you turning your intuitions into actionable rules.

For the SEK journal, I did exactly that. I had strong intuitions about what makes a good article in this editorial DNA: consistent “you” address, short sentences in the opening paragraphs, the absence of certain corporate language tics, a number or a concrete example in every section. But those intuitions weren’t formalized.

I used a frontier model to externalize them. The result: a binary quality bar, mechanically verifiable, that any cheap model can apply as a validation filter.

Concrete case: the binary quality bars that drive cheap models

A binary quality bar is a list of criteria written so that there’s no room for interpretation. Each criterion has an answer: yes or no.

Here are a few concrete examples from SEK’s editorial DNA:

Blocking criteria (a single fail = the article doesn’t get published):

  • Zero quoted passages without a source marker in the same sentence
  • Zero figures absent from the brief and unverifiable
  • The defined CTA is present at the end of the article, nothing after it
  • Length within ±20% of the article type’s range
  • No leftover placeholders ([TODO], lorem, unclosed tag)
  • Systematic “you” address — zero leftover “vous/votre/vos”

Quality criteria (a fail = warning, not a block):

  • The first sentence is neither “In a world…” nor a definition
  • No sentence > 35 words in the first 2 paragraphs
  • Maximum one rhetorical question per article
  • Zero “indeed”, “nevertheless”, “the fact remains”
  • At least one passage ties the idea to the reader’s real life

A cheap model can check these criteria mechanically. It doesn’t need to understand the meaning of the article. It doesn’t need editorial judgment. It scans, it ticks boxes, it flags the violations.

The frontier model did the upstream work: it helped formulate these criteria precisely enough to be verifiable without ambiguity. That’s half a day’s work, done once. It then generates savings on every article produced.

That’s the asymmetry that makes this pattern interesting for a solopreneur: high fixed cost once, near-zero marginal cost after that.

The teacher/student pattern: the frontier model writes the textbook, the cheap models execute

Let’s formalize the pattern. It comes down to three steps.

Step 1 — Extraction session with the frontier model

You show up with your intuitions, your examples of good and bad output, your constraints. The frontier model plays the consultant who helps you formalize what you already know but haven’t written down yet.

The session looks like this: you show it 3 articles you consider good and 3 you consider bad. You ask it to identify the patterns. It proposes hypotheses. You validate them, correct them, refine them. After an hour, you have a first draft of the standard.

The model doesn’t decide. You decide, and the model encodes.

Step 2 — Turn it into binary criteria

The draft standard still has gray areas. “The text should be engaging” is not a binary criterion. The frontier model helps you break it down: what makes a text engaging in your specific context? A hook under 20 words? A concrete example in the first 150 words? An action verb in every H2?

You iterate until every criterion is testable without interpretation.

Step 3 — Deploy on cheap models

The standards document becomes a system prompt, a checklist, a validation tool. You inject it into your production workflow. Cheap models — Haiku, Flash, Mistral Small, depending on your preferences — apply it to every output.

The frontier model is no longer in the daily loop. It comes back when the standards need to evolve, or when you have a new domain to formalize.

This pattern applies well beyond writing. Picture a freelancer who handles code audits: the frontier model writes the review checklist (architecture, security, performance, readability), a cheap model scans each PR against that checklist. Or a solopreneur who qualifies leads: the frontier model defines the qualification criteria (company size, implied budget, urgency, product fit), a cheap model scores each CRM entry.

Anywhere you have judgment to encode once and execution to repeat a hundred times, the pattern holds.

To go further on the numbers that explain why this kind of solo + AI architecture is becoming the norm, take a look at the solopreneur & AI statistics 2026 roundup.

How to apply this pattern to your business today

You don’t need complex infrastructure. Here’s the minimum sequence to start this week.

1. Identify a repetitive task where your judgment isn’t formalized

It’s often a task you do well but couldn’t easily hand off. You can recognize a good result, but you couldn’t explain it in 5 minutes to someone else. That’s exactly where the pattern applies.

Common examples: editing articles, client email replies, prospect qualification, code review, design validation.

2. Prepare 5 to 10 annotated examples

Before your session with the frontier model, gather concrete examples of good and bad output. Annotate them briefly: why this one is good, why that one isn’t. This prep work cuts session time in half.

3. Run a 60-to-90-minute extraction session

Give your examples to the frontier model. Ask it to identify the patterns. Iterate on its proposals. Push it to formulate binary criteria, not vague principles. If a criterion contains the words “appropriate”, “relevant”, or “engaging” without an operational definition, it isn’t precise enough yet.

4. Test the standard on 5 new cases before deploying it

Take 5 recent outputs and apply the standard by hand. Do the criteria actually capture what you wanted to capture? Are there false positives (outputs you consider good but the standard rejects)? False negatives (bad outputs the standard passes)? Refine accordingly.

5. Encode the standard into a system prompt and deploy on a cheap model

Once the standard is stable, turn it into instructions for a cheap model. Test on 10 outputs. Measure the agreement rate with your own judgment. If you agree with the model in more than 85% of cases, the standard is operational.


The honest limit of this pattern: it only works in domains where you already have judgment. If you have no intuition about what “good” looks like in a domain, the frontier model can’t extract it — there’s nothing to extract. The teacher/student pattern assumes there’s a teacher. AI encodes; it doesn’t replace expertise.

That’s also why this pattern is a particularly good fit for solopreneurs who have experience in their field. You have 5, 10, 15 years of accumulated judgment. That judgment is your most valuable asset — and right now it’s locked in your head, not scalable, not delegable. The frontier model is the tool that lets you externalize it.

If you want to see how this kind of architecture applies in practice to a site or a product, the SEK audit is built for that — we look at what’s blocking you, we formalize it, we make it actionable.


Sébastien de Bollivier has been building products solo from La Réunion. If you have a project that’s stalled or a technical job to unblock, his profile is on sebastiendebollivier.com.

Frequently asked questions

When should you use a frontier model rather than a cheap model?

Use a frontier model (Claude Opus, GPT-4o, Gemini Ultra) for judgment tasks: writing standards, auditing logic, setting binary criteria. Save cheap models for the repetitive execution of those standards. The simple rule: if the task can be described by a document written once, the frontier model writes the document, the cheap model applies it.

How do you build a binary quality bar that a cheap AI can run?

Phrase each criterion as a yes/no check, with no room for interpretation. Example: 'No sentence > 35 words in the first 2 paragraphs' is binary. 'The text is well written' is not. A frontier model helps you turn fuzzy intuitions into testable binary criteria — that's exactly what extracting judgment means.

Does this pattern work outside of content writing?

Yes. The teacher/student pattern applies to any repetitive process: code review (the frontier model writes the checklist, a cheap model scans each PR), lead qualification (the frontier model defines the criteria, a cheap model scores each CRM entry), customer support (the frontier model writes the reply base, a cheap model answers). Anywhere you have judgment to encode once and execution to repeat a hundred times.

Got an idea to ship? A website, a SaaS, an AI automation — built with you.

Talk about your project
Behind the scenes of the studio ✦

New products, work in progress and useful resources — the SEK studio in your inbox.