Subjects ยท Computer Science & Data
Computer Science & Data: The verification drill
Check AI-generated code properly, which in this subject is unusually easy and unusually often skipped for exactly that reason.
What you'll be able to do: Check AI-generated code properly, which in this subject is unusually easy and unusually often skipped for exactly that reason.
The subject where checking costs nothing, and the trap that creates
Programming has a property no other subject has: you can verify almost anything for free. Run it. Write a test. Check the types. Print the intermediate value.
That makes programmers systematically over-confident about AI reliability in general: "it's fine, I check it" is true here and much less true in medicine, law or history, where checking is expensive or impossible. If your intuitions about trusting these tools were formed while programming, they don't transfer.
Within programming, the cheapness creates its own trap: because checking is easy, students do the easiest check (does it run) and stop. That catches syntax and crashes, which are the errors that don't matter, because they're self-announcing.
The errors that matter are the ones that run.
The failure modes, ranked by how much they cost
Plausible API that doesn't exist. A method name that sounds exactly right for the library and was never in it. Caught instantly by running, and the reason to run before reading further.
Right for the happy path, wrong at the edges. Works on your example, fails on empty input, a single element, a negative number, a unicode string, a duplicate. Nothing announces this.
Subtly wrong logic that passes your test. An off-by-one that your test doesn't exercise. A comparison that's wrong for equal values.
Correct, and wrong for your context. A solution that assumes sorted input, or unbounded memory, or single-threaded access. It's right; it doesn't fit.
Insecure. String-concatenated SQL, unvalidated input, a secret in the source. Runs perfectly.
Version drift. Correct for a library version you're not using. Very common and not always loud.
Stale idiom. Correct and no longer how anyone does it, which costs you in review rather than in output.
The drill
Four checks, about two minutes:
1. Run it. Not later, first. Catches fabricated APIs and crashes in seconds.
2. Run it on the edges. Empty, one element, zero, negative, very large, duplicate, unicode, null. Ask explicitly:
What inputs would break this? Empty, null, huge, negative, unicode, concurrent, malformed. Don't fix them, list them, and I'll decide which matter.
3. Read it for context fit. Does it assume something your system doesn't guarantee? This is the check that requires you to think and can't be automated.
4. Read it for security if it touches input, storage or the network. Every time.
The practice version
The drill above is what you do with real code. This is how you get good at it:
Write me 100 lines of [language] that look correct and contain exactly two plausible bugs. Don't tell me where. I'll find them, then you mark me.
Then escalate: one bug, or possibly none. Knowing there's a bug makes you suspicious in a way real code doesn't.
This is deliberate practice for code review, and it's the closest thing to professional review training that's available to someone learning alone.
What running does NOT verify
Worth stating, because the ease of running creates false confidence:
- It doesn't verify the code is correct: only that it does something on the input you gave. - It doesn't verify the approach is appropriate. - It doesn't verify performance at scale. - It doesn't verify security. - It doesn't verify the code is maintainable, which is most of what professional review is about (article 03).
"It runs" is the floor, not the ceiling, and treating it as the ceiling is the characteristic student error.
Pitfalls
- Stopping at "it runs". The defining failure.
- Testing only the happy path. Which is the path you thought of, which is the path the code was written for.
- Accepting a fix you don't understand. If you can't say why it works, you've traded a bug you could find for one you can't.
- Generalising your trust. Verification is cheap here. Don't carry that confidence into subjects where it isn't.
- The tell: you've never found a bug in generated code. You're not looking, or you're only checking whether it crashes.
Try this today
Ask for a hundred lines with exactly two plausible bugs planted, unmarked.
Find them. Then ask which one you missed and why it was hard to see, the answer is usually "it only fails on input you didn't try", which is the lesson.