Part 4

Models, Compute, And Cost

When to spend $200 a month on a frontier model and when to spend $0 a month on hardware in your closet. The math, the trade-offs, and the vocabulary you need to read a model's spec sheet without nodding along to things you don't understand. Sections 4.2, 4.3, and 4.5 are deep-dives; you can skip them on first pass and come back when you're buying hardware.

Start with section 4.1
  1. 4.1

    Frontier Vs. Local: When Each Makes Sense

    The most important budgeting decision you'll make. Not every problem deserves Claude Max; not every problem can be solved with a Mac mini.

  2. 4.2

    How Models Actually Work (weights, Parameters, Active Parameters)

    DEEP-DIVE

    [Deep-dive] You don't need the math. You need the vocabulary to read a model spec sheet without nodding along.

  3. 4.3

    Ram, Vram, And Unified Memory

    DEEP-DIVE

    [Deep-dive] The number that determines whether you can run a model is "how much memory you have, and what kind." The kinds matter.

  4. 4.4

    Self-hosting On Mac (mini, Studio, Pro)

    The easiest on-ramp to running models on hardware you own. Mac mini for solo use, Mac Studio for small teams, Mac Pro for the edge cases.

  5. 4.5

    Self-hosting On Nvidia (spark, Orin, Full Gpus)

    DEEP-DIVE

    [Deep-dive] The NVIDIA path. More flexible, more setup, often cheaper for comparable capability. Jetson Orin to DGX Spark to multi-GPU rigs.

  6. 4.6

    Cost Literacy: Tokens, Context Windows, And The $200 Question

    The cheat sheet: tokens, input vs output pricing, context windows, and the actual math on whether Claude Max is worth $200 a month.