Subjects ยท Computer Science & Data

Computer Science & Data: Predict then verify

Say what the code will do before you run it, which is free, takes ten seconds, and is the highest-return habit available in this subject.

What you'll be able to do: Say what the code will do before you run it, which is free, takes ten seconds, and is the highest-return habit available in this subject.

Why this is the first method for computing

Every other subject has to wait for feedback. You submit, and three weeks later someone tells you whether your model of the world was right.

In programming, the feedback is instant, unambiguous, free and unlimited. You can find out whether you're right about anything, right now, by running it.

That advantage is almost entirely wasted, because students run code to find out what it does rather than to check whether they were right. Those are different activities. The first produces no learning at all, you observe an output and move on. The second produces a prediction error every time your model is wrong, and prediction error is where the learning is.

The whole method is inserting one step you currently skip: say it first.

The habit

Before pressing run, state:

  1. What will be printed / returned / raised.
  2. Why: which line does what, in what order.

Then run.

When it disagrees with you, that's the event. Don't skim past it: find out which part of your model was wrong. The specific value of this in programming is that you're usually wrong about state: what a variable holds at a given moment, and state is invisible in source code.

Before running this, ask me to predict the exact output and explain my reasoning. Wait for my answer. Then tell me what actually happens and, if I was wrong, which part of my reasoning was wrong, not just which line.

The classic predictions worth failing

Each of these catches a specific wrong model, and most programmers got each one wrong exactly once:

x = 10 def f(): x = 20 f() print(x)

Tests: does assignment in a function affect the outer name?

a = [1, 2, 3] b = a b.append(4) print(a)

Tests: does assignment copy? The single most consequential beginner misconception, and it survives months of successful programming because most code never triggers it.

print(list(range(3)))

Tests: is the endpoint included? Off-by-one errors start here.

print(0.1 + 0.2 == 0.3)

Tests: floating point. Predict, run, and then never be surprised by a currency bug.

def f(items=[]): items.append(1) return items print(f()); print(f())

Tests: mutable default arguments. Predict this one and you will remember it permanently.

For loops with closures, for async ordering, for shallow versus deep copies, same pattern. Predict, run, and let the disagreement teach you.

Beyond single snippets

Before writing a function: predict its signature and behaviour from the name and the docstring, then write it. If your prediction and your implementation differ, the name is wrong.

Before reading unfamiliar code: predict what a function does from its name and signature, then read the body. The gap tells you whether the codebase's naming can be trusted, which is genuinely useful information about a codebase.

Before a query runs: predict the row count. Order of magnitude is enough, This catches join errors immediately, and a cross join returning ten million rows is otherwise discovered by waiting.

Before training: predict roughly what accuracy a baseline should get. Predicting 50% for a balanced binary problem and getting 97% means a leak, and you'll spot it in seconds instead of in review.

Before a refactor: predict which tests will fail. If none do, you may not have test coverage where you think.

The four kinds of miss

  • Wrong output, wrong reasoning: your model of the language or the data is wrong. Most valuable.
  • Right output, wrong reasoning: dangerous, passes unnoticed, and only caught because you stated the reasoning.
  • Wrong output, right reasoning: usually an environment or version difference, which is worth knowing about.
  • No prediction possible: you don't have a model yet. That's a finding, not a failure: this is code to study, not code to test yourself on.

Pitfalls

  1. Predicting silently. Say it or type it. Hindsight is instant and total.
  2. Predicting the output without the reasoning. Halves the value and hides the right-for-wrong-reason case.
  3. Only predicting when unsure. The confident cases are where the broken models hide.
  4. Running to see, then rationalising. If you find yourself explaining why the output makes sense after seeing it, you skipped the method.
  5. The tell: you're never surprised. Either you're predicting code you already know, or you're not committing before you run.

Try this today

For the next hour, don't press run without first saying what will happen.

Count the disagreements. Every one is a thing you believed about your tools that was wrong, and you'd otherwise have found out when it mattered.