Podcast & Video

When I wrote the LLM Fine Tuning & Performance Post, a good question that came up was: what exactly is fine-tuning and how does it work? This post can’t fully cover the topic, but I want to give some solid context around how fine tuning works for LLMs and MLLMs (the billion-parameter models that come from Anthropic, OpenAI, Mistral, Meta, DeepSeek, etc.).

Large language models are incredibly capable. They can write essays, debug code, or summarize medical records. But out of the box, they’re trained to handle everything in a general way, and they have limits based on the data they have access to. Granted, they continue to grow and improve in their capability and breadth of expertise so this is continually changing.

Fine-tuning is how you take that general intelligence and make it specific to a certain area or topic. You may have a lot of private and specialized data that a standard LLM can share some insight on, but it can’t fully dive deep to that level of expertise for your case. So you share your data with it to help it become more of an expert on your domain.

Mentoring a Genius Apprentice

Imagine hiring a brilliant apprentice. They’ve read every textbook ever written, but they don’t know your company, your corpus of knowledge, or your workflows. So you would show them examples of how you do things to help them learn style and standards. That’s fine-tuning in a nutshell: you start with an existing model and teach it with your own data and examples until it understands your domain, style, and goals.

This metaphor applies to large language models (LLMs) that handle text and to multimodal LLMs (MLLMs) that work with text, images, or even audio, illustrating how fine-tuning uses data to shape any model into a specialist across different types of context and domains.

Why We Fine-Tune

A base LLM knows a bit about everything. Fine-tuning helps when you need:

  • Accuracy in a niche (like technical writing or specialized analysis)
  • Consistency across multiple users or tasks
  • Efficiency, so prompts can be short but still produce great results
  • Voice control, ensuring everything sounds like your style and tone

Prompting can get you close, but fine-tuning bakes the knowledge into the model itself. That way you don’t have to keep repeating instructions and giving it context.

That’s where the real transformation happens: when broad intelligence becomes specifically shaped to your domain. It’s about adapting the tool so it can create content and reflect a unique knowledge base, tone and priorities.

Fine-Tuning vs. Training From Scratch

Training from scratch means teaching an AI language from zero like raising a child who doesn’t yet know words. You’re building the foundational understanding of language, grammar, facts, and reasoning from nothing. This requires massive datasets (think: most of the internet), enormous computational resources, and months of training time.

Fine-tuning is fundamentally different. The model already has extensive knowledge and reasoning ability. You’re teaching it how to apply that knowledge in your specific context, with your standards and voice. It’s like working with someone who’s already highly skilled but needs to learn your particular craft.

This distinction has major practical implications:

Training from scratch:

  • Requires petabytes of data
  • Costs millions of dollars in compute
  • Takes weeks to months
  • Teaches foundational language understanding

Fine-tuning:

  • Requires thousands to millions of examples (vastly less)
  • Costs hundreds to thousands of dollars
  • Takes hours to days
  • Teaches domain-specific application

Preference-based fine-tuning builds on this by acting as a refinement layer—teaching not only knowledge, but judgment. It adjusts behavior, tone, and alignment rather than base facts.

Because you’re building on existing capabilities rather than starting from nothing, fine-tuning is faster, cheaper, and more accessible to small teams and startups. This is why fine-tuning has become the primary way organizations customize AI for their needs when they do.

Note advancements are already happening to make training cheaper, faster and more efficient but fine-tuning will still be better on all those dimensions if you have access to a base model to work with.

What Actually Happens During Fine-Tuning

For large-scale LLMs and MLLMs, fine-tuning takes different forms depending on your goals. The main teaching methods include continued pre-training, supervised fine-tuning (SFT) and preference-based tuning (RLHF/DPO).

Continued Pre-training | Learn the language

This is about domain knowledge absorption. Feed the model large amounts of raw text from your domain like industry publications, company documentation, and specialized corpora. The model learns by predicting the next word, absorbing domain vocabulary, terminology, and concepts. No human judgment needed.

Example: Let’s say you are building an AI assistant for a punk-Victorian tailoring house. You feed it historical pattern-making texts, corsetry construction manuals, leather working guides, and avant-garde fashion theory. It learns what “boning channels,” “grommets,” “tailcoat tails,” and “distressed finishing” mean in context of both traditional structure and rebellious subversion.

Supervised Fine-Tuning (SFT) | Learn what right looks like

This is the most common and its about teaching the model how to respond. Feed the model input-output pairs showing exactly what good responses look like. You demonstrate: “here’s the prompt, here’s the perfect response.” The model learns the relationship between requests and expert responses in your specific style.

Example: Training your pattern generation assistant with pairs like:

  • Input: “Create a fitted waistcoat pattern that combines Victorian tailoring with punk hardware”
  • Output: “Start with a classic six-panel Victorian waistcoat block with princess seaming for structure. Add 2-inch lapel facings to accommodate metal grommets at 1.5-inch intervals. Include internal boning channels at the side seams for corset-like shaping. For the punk edge: specify distressed leather or heavy twill, design asymmetric pocket placement, and plan for D-ring hardware on the belt.”

Preference-Based Tuning (RLHF/DPO) | Learn what’s better

RLHF (Reinforcement Learning from Human Feedback) and DPO (Direct Preference Optimization) are different technical approaches to learning from preferences, but both teach what “better” looks like. Feed the model ranked pairs of responses and mark which is better. It learns judgment and refinement through comparison.

Example: Training your pattern generation assistant for alignment on what is better:

  • Response A: “Make a waistcoat with Victorian shape and add grommets.”
  • Response B: “Start with a classic six-panel Victorian waistcoat block with princess seaming for structure. Add 2-inch lapel facings to accommodate metal grommets at 1.5-inch intervals…”

You mark Response B as preferred. Over thousands of comparisons, the model learns richness, specificity, and appropriate detail level.


Quick comparison:

What you actually need:

  • Want the model to deeply understand your domain’s language? → Start with Continued Pre-training, then add SFT
  • Building a domain-specific tool (content creation, analysis, specialized tasks)? → SFT is usually the core of what you need
  • Deploying a customer-facing assistant where tone matters? → SFT first, then add Preference-Based Tuning

How It Works: The Learning Details

We’ve covered the different teaching methods. Now let’s dive into the actual mechanics of fine-tuning with SFT to understand what’s happening under the hood.

The Foundation: Parameters

The billions of parameters (weights) that make of the LLM are like microscopic knobs that control how the model interprets and generates text. These parameters determine whether the model continues a sentence with a technical explanation, a creative flourish, or a data-driven insight. Fine-tuning adjusts those knobs so the model responds in a way that fits your specific domain.

Let’s see how this actually works using our punk-Victorian tailoring example.

Step 1: Teaching Through Examples

You feed the model thousands of pairs of inputs and outputs. Each pair shows the model what you want:

User: Create instructions for a tailcoat with asymmetric lapels and corset-style back lacing.

Assistant: Begin with a traditional tailcoat pattern as your foundation. Ensure proper tail length (to back of knee) and structured shoulders. Modify the front: draft one peaked lapel at 3 inches and one notched lapel at 2 inches for asymmetry. For the back: add a center-back seam with 6-inch lacing panel from waist to mid-back. Install grommets at 1-inch intervals, sized for 1/4-inch cording. Use heavyweight wool or brocade for the body, leather for lapel facings. This bridges Victorian formality with punk disruption.

The model isn’t memorizing these examples. It’s learning the underlying pattern: the relationship between design briefs and expert instructions that honor both technical precision and aesthetic vision.

Step 2: The Learning Mechanism

When the model generates an output during training, it compares its attempt to your target output which is the example you provided of what you want. The difference is called loss. “How far off was I from what you wanted?”

The model uses an algorithm called gradient descent to minimize this loss. It nudges millions of parameters in tiny increments, moving toward more accurate predictions. Think of it like making fine adjustments to a garment’s fit: you pin, test, adjust the seam by a quarter inch, test again with each small refinement bringing the piece closer to the perfect drape and silhouette you envisioned.This happens through:

  1. Forward pass: Model generates output
  2. Loss calculation: Compare output to target and get error
  3. Backward pass: Calculate how each parameter contributed to the error
  4. Parameter update: Adjust parameters to reduce error

This cycle repeats thousands of times across your entire dataset.

Step 3: Preference-Based Learning (When Applicable)

In many modern setups, your data doesn’t always include a single “right” answer. Sometimes you just know which examples are better. This is where preference-based fine-tuning comes in.

Instead of exact answers, you provide ranked comparisons:

  • Response A: Basic pattern instruction with no style consideration
  • Response B: Detailed instruction that balances traditional tailoring technique with punk design elements

The model learns to prefer Response B. Methods like RLHF and DPO use this approach to teach judgment on what’s better and why even when there isn’t a single correct answer.

Step 4: Iteration and Validation

This process happens thousands of times. After each epoch (a full pass through the training data), the model’s predictions get tested on examples it hasn’t seen before.

Detecting problems:

  • In SFT: If the model overfits (memorizes too precisely), performance drops on new examples
  • In preference-based tuning: The model can start “gaming the reward”and optimizing for high scores rather than genuinely useful outputs

Catching these issues:

  • Evaluate on a held-out validation set
  • Track metrics like loss, perplexity (how “surprised” the model is by the validation data), or preference consistency
  • Human review to ensure responses sound natural and useful

When the model generalizes well, it produces high-quality, consistent responses even on new prompts it’s never seen.

Step 5: Internalized Expertise

Once training completes, the model doesn’t just repeat memorized phrases. It has internalized patterns of logic, structure, and style from your examples.

Ask it something new:

User: Design a fitted jacket with Victorian military styling and modern punk hardware.

The fine-tuned model applies the craft it learned by balancing historical accuracy with contemporary edge, using your atelier’s distinctive approach, providing technical precision while honoring creative rebellion. It generalizes from its training to handle requests specific to your needs.

Step 6: Deployment and Continuous Improvement

Once fine-tuned, the model goes into production. But the work continues:

  1. Collect feedback from real users
  2. Flag weak responses that don’t meet standards
  3. Add examples back into the dataset
  4. Fine-tune again periodically to keep improving

In preference-based systems, this might include updated comparison data or reward model recalibration. Human evaluators continue refining what “better” means, ensuring the model evolves with design trends, new techniques, and aesthetic shifts.

This human-in-the-loop cycle keeps the model aligned with your craft’s evolution.

The bottom line: Fine-tuning is structured practice. Every dataset is a lesson plan, every iteration is rehearsal, until the AI performs at the level you expect.

A Note on Data Quality

Fine-tuning success depends far more on data quality than volume. A handful of well-crafted examples that represent your “ideal outputs” teach faster than thousands of inconsistent ones.

If you wouldn’t want your team learning from it then your model shouldn’t either. Clean, consistent, representative examples are the foundation of effective fine-tuning and this is a much deeper dive topic beyond this post.

Finishing Touches

Remember that brilliant apprentice we talked about at the start? Fine-tuning is how you turn them into a specialist who truly understands your craft like an apprentice tailor learning not just to sew, but to cut, fit, and finish in your signature style.

We’ve covered the essential landscape: what LLM fine-tuning is, why it matters, and how it actually works from absorbing domain knowledge through continued pre-training, to learning your standards through supervised examples, to refining judgment through preference feedback. We explored how to keep models generalizing well, the different approaches you can take, and why quality trumps quantity every time.

As base models continue to improve and become more capable out of the box, the question isn’t whether you need to fine-tune rather it’s how much you need something tailored to your specific expertise, voice, and standards. General-purpose models might handle 80% of tasks beautifully, but fine-tuning is what closes that gap to 99%.

That’s the transformation that matters. When you fine-tune thoughtfully, you’re not just customizing a tool. You’re teaching AI to think and speak in your language, reason with your values, and craft with your expertise.