You Are Not an Impostor: Agentic Coding and the New Computational Scientist

Jun 23, 2026·
Yassir Boulaamane
Yassir Boulaamane
· 4 min read

A pipeline I did not type ran the first time. It parsed a messy SDF, built structural descriptors, docked the library, and fed the scores to a PyTorch graph network. It was parallel, and it worked. The feeling afterwards was not relief. It was the suspicion that I had stopped being the scientist.

That suspicion treats typing as the work. The work is deciding what the script is allowed to mean.

What the agent executes and what you still ownThe agent writes boilerplate, features, and training loops. The scientist owns the hypothesis, the physical parameters, and the check that a pose or a statistic is real.Speed is not authorshipThe agentScripts and boilerplateFeatures, training loopsWill invent a clean plotYouThe hypothesisThe physical checkThe part you would defend
Figure 1. The pipeline in the opening is one project: an SDF, descriptors, a docking run, a PyTorch GNN. The time comparison in the post is personal, two days of boilerplate down to about twenty minutes. It is not a benchmark.

I work between cheminformatics, molecular modeling, and machine learning. Throughput is the currency. Once agents could plan, run, and repair a multi-step scientific workflow, I used them. Two days of RDKit errors and data cleaning became about twenty minutes. The credibility panic showed up on schedule: if the agent writes the engineering, what is left of the job?

Struggle was never the metric

The field has a bad equation in its head. The insight counts for more if the code fought you. An out-of-memory crash, a broken CUDA setup, a macrocycle RDKit would not sanitize: those are friction, and friction feels like payment.

Every earlier abstraction produced the same complaint. Classical force fields looked like a shortcut to people who had only done quantum chemistry. Biopython and scikit-learn looked like copy-paste to people who had written the statistics by hand. Genomics, structural biology, and climate modeling have the same story.

An agent removes accidental complexity: boilerplate, an API that moved, the code that formats an input. Essential complexity stays. That is the physics, the hypothesis, and whether the design is valid. A hard afternoon is not a methods section.

Architect, not bricklayer

An agent can write 500 lines that train an active-learning loop and still not know why. It does not know the allosteric pocket, the flexibility of the enzyme, the structural alert that makes a hit a metabolic problem, the synthesis constraint, the hazard of a scaffold, or the off-target sitting in a resistance mechanism.

The agentYou
ExecutionScripts, features, boilerplate modelsWhich target, which physical settings
ContextThe prompt and its training dataMechanism, and what the assay will do
Failure modeA fluent, unphysical resultYou are the person who can reject it

Memorizing PyTorch or OpenMM commands is not the value. Taste is: what you refuse to believe.

Govern it like a fast junior

The hollow feeling comes from accepting code you did not read. The correction is governance, not abstinence.

Read it as if you will defend it to a PI or a reviewer. If a loss function is opaque, make the agent explain the math and then write the comment yourself.

Spend the saved hours on checks. A small synthetic set with a known answer will show whether the output matches the science. In modeling, that is a pose that is physically possible. In genomics, that is a variant call against a truth set.

Keep the intellectual center in your hands: a novel score, a physical constraint, the mathematical change that is actually the paper. File conversion, a training loop, and a PCA plot can go to the agent.

What the speed is for

The next few years will not score people by typing speed. They will score the scale of the question. Using an agent now is practice at running a leveraged experiment, the distance between a hypothesis and a computation getting short.

The closing comparison in the original note is a bandwidth claim, not a measured speedup: five times the virtual libraries, ten times the trajectories, or twenty more generative models in a week. Those multiples are rhetorical. The non-rhetorical part is the division of labor above. You are not an impostor for refusing to type the boilerplate. You are responsible for the result.