← All posts

The run from two years ago

researchdataphd

Your supervisor asks what the box size was. Not in a difficult way. A reviewer wants it stated properly in the methods, and it should take you ten seconds.

You knew this number once. For about four months it was the most familiar number in your life. You knew the water model, the ion concentration, why you used that cutoff instead of the obvious one, and what went wrong the first three times.

That was in your second year.

You open the folder. It is called sim_final. Inside it is sim_final_2, and inside that are fourteen directories named with dates in a format you stopped using a year ago. Two of them are called test. One is called test_ok.

None of this is carelessness

The honest thing about research mess is that every step was reasonable when you took it.

You named a folder test because it was a test and it was going to be deleted by Friday. You put the good run on the lab workstation because that is where the GPU was. You kept the analysis in a notebook because you were still working out what the analysis was. You did not write the parameters down properly because you were going to redo it properly once it worked, and then it worked, and you moved on to the next thing that did not work.

Nobody accumulates this on purpose. It accumulates because a PhD is four to six years of decisions, each one optimised for the next two weeks.

Count the places your work actually lives

If someone asked you to list them, you would guess low.

Some of it is on the cluster, in scratch, under a purge policy you have never read. Some is on an external drive in a drawer, and you are fairly confident which drawer. Some is on the lab workstation that got reimaged when the new student joined. Some is on your old laptop, the one with the swollen battery that you keep meaning to deal with. Some is in a Drive folder shared by a labmate who has since graduated and whose institute account is now disabled. Some is in a WhatsApp thread, because at eleven at night that was the fastest way to send someone a plot.

And a surprising amount of it is in your head, which at the time was a completely reasonable place to keep things.

The naming schemes made sense on the day

Everyone thinks their versioning is fine while they are using it.

run3. run3_new. run3_new2. run3_final. run3_final_USE_THIS.

The information in those names is real. USE_THIS meant something exact on the afternoon you typed it. The trouble is that it records a decision without recording the reason, and two years later the reason is the only part you need. You can see that you chose one. You cannot see why you stopped trusting the other three.

What you did, what got recorded, what went in the paper

These three are never quite the same thing, and the gaps widen every month.

This is not about anyone being dishonest. It is far more ordinary. You ran the analysis eleven times while figuring out what the analysis should be. The version in the paper is the eleventh. The notebook on your machine has been edited several times since, so the cell that produced that figure no longer produces that figure. The parameters written in your lab notebook are from the sixth attempt, because the sixth attempt is the last time you remembered to write parameters down.

So when a reviewer asks you to regenerate the figure with error bars, the true answer is that you can, and it will take three days, and one of those days goes entirely on working out what you did.

The same thing happens with sequencing runs that were reprocessed after a pipeline update, and with models retrained after someone fixed the data loader. The result in the paper came from a state of the world that no longer exists anywhere on disk.

It all comes due in the same six weeks

None of this costs you anything while you are still producing results. Recent work is fresh, and nobody is asking about the old work.

Writing up is the first time the whole thing has to be accounted for at once, in order, with numbers that agree with each other. You are reconstructing four years of decisions in about six weeks, while also trying to write.

This is why the last stretch feels so out of proportion to the task. It is usually not the writing itself. Most researchers can write a paragraph. It is that writing forces you to go back and find things, and finding things turns out to be the expensive part. The thesis is not a writing problem wearing a writing costume. It is an archaeology problem with a deadline.

What actually helps is boring

There is no clean answer to this, and most of the clean answers on offer are being sold by somebody.

The researchers who seem to suffer least at the end do two unglamorous things. They keep the thing that made the figure next to the figure, so the plot and the code that produced it never drift apart. And they write the methods paragraph while the work is still in their head, not two years later when it has to be excavated.

Neither of those needs software. Both are unreasonably hard to actually do while a PhD is happening to you.

It is not a discipline problem

The instinct, standing in front of sim_final_2 in your fourth year, is to conclude that you should have been more organised. That a better researcher would have kept a cleaner record.

Some of that is true and most of it is not worth the guilt you are about to spend on it. Every folder was named by someone who had a reason. Every file went where it went because that was where the GPU was, or where the free space was, or where a labmate could reach it at eleven at night. None of it was stupid on the day.

The mess is not what happens when a PhD goes wrong. It is what four years of individually sensible decisions look like when you finally turn around and face them all at once.

Almost every researcher you respect has that drive in a drawer somewhere.