A notebook for AI-assisted research
A working method for AI-assisted research, built for the moment someone asks you to run it again, and you actually can.
Every AI demo you have ever found impressive was run more than once before you saw it. Somewhere off-screen there were failed takes: the model hedged, the retrieval pulled the wrong passage, a sampling setting picked the wrong word and the whole answer fell apart. What made it to the screen was the version that worked. That is not dishonest, exactly. It just is not evidence of anything beyond itself.
The trouble starts when a researcher treats one good run the way they would treat a finding: something to build on, cite, or ship. A demo shows that a system can produce a good outcome under some combination of luck, context, and phrasing. It does not show that the outcome is typical, and it does not tell you what happens when a colleague tries the same thing tomorrow with a slightly different prompt.
The fix is not cynicism. It is a habit: before you get attached to an output, ask what you would need to see happen a second time, under conditions someone else could recreate. If you cannot answer that question yet, what you have is a demo, not a result.
Most AI-assisted research dies quietly in a chat window. Someone finds a prompt that produces a genuinely good result, moves on to the next question, and by the following week cannot remember the exact wording, the model version, or the two failed attempts that came before it. The result still happened. It is just no longer reachable.
A lab notebook exists to solve exactly this problem, and it solves it whether the experiment involves a reagent or a temperature parameter. Write down the prompt as it was actually sent, not the tidied-up version you would use in a write-up later. Note the model and its version, the date, the surrounding context or retrieved material, and anything about the run that felt like it mattered even before you know why. Cheap to capture in the moment, close to impossible to reconstruct afterward.
The habit is unglamorous, which is exactly why it gets skipped. Nobody skips it twice after losing a genuinely good result to a closed browser tab.
Exploration and confirmation are different jobs. Doing them in the same window, with the same casual attention, is how good ideas get mistaken for proven ones.
Give yourself permission to explore loosely: change the prompt mid-thought, abandon a direction, chase a tangent. That is how the interesting ideas surface. But treat the moment you decide something is worth keeping as a hard boundary. From there, freeze the inputs, write down the configuration, and run it again before you trust it.
Longer notes on the parts of AI-assisted research that are easy to get wrong.
On tracking prompts and parameters, separating exploration from conclusions, and why a good demo is not a result.
Read the note →An occasional note when something changes in how we think about reproducibility. No noise, and unsubscribing is as easy as subscribing was.