BlogProductAbout UsSponsorAPPLY
APPLY
← All posts

From scientist to AI builder

Julia Gross on building agentic workflows.

Julia Gross smiling with two fellow builders at a busy hackathon

Written by Julia Gross.

In March, I went to my first hackathon, the AI for Science Cell Cultivation Hack with Worldwide Studios and Monomer Bio. The marketing said real science could be done in a weekend. I thought NO way — that’s not how science works. But the bacteria were real, so I had to try.

My team, ViNatX, set out to optimize for Vibrio natriegens growth. I wanted to read the literature first, but there wasn’t time. We had to come up with experiments on the fly where the measured outcome was the only metric, delegating iterative interpretation. So we pre-registered kill conditions. A teammate rigged a Raspberry Pi to watch the plates, and we went to bed while the robots kept executing against the conditions we’d set. We got three full experiments in 48 hours.

We didn’t win. But it was the first time the news about AI felt relevant to my career in biology.

Building every two weeks

As a consultant for a healthcare practice, I had to get up to speed fast on everything from dry eye to mastectomy recovery. Claude’s answers looked plausible but weren’t reliable enough for real patients. So in the fellowship I built a literature pipeline on GXL’s Paperclip, then Pinakes and breath-cam.

Every project raised the same question: how do I know this is true? That became my final demo, watch_agent_run, which compares what an agent says it did with what its logs show.

After my first build using their paperclip agent, I sent GXL a bug report. That turned into a job. Now I help scientists use Paperclip and help the product team build what they’ll actually use. I helped shape re:AGENT, GXL’s first end-to-end agentic science hackathon. Twenty-eight teams used Paperclip to search, structure, and verify scientific data across a huge range of subjects. They built things as concrete as a million-entry database of antimicrobial peptide sequences, and as abstract as new ways of transforming reasoning spaces. It was immensely rewarding to help run a hack that facilitated others’ “NO way” → “huh, this could matter for science” moment.

Drive at the speed of verification

A PhD trains rigor; two-week sprints train speed. Rigor still wins, but by less than a PhD can condition you to believe, and I’m a stronger practitioner now for having trained both.

People who think AI is worthless are wrong, but so are people who think it’s magic. What matters is who’s driving. After 100 days of shipping something new every two weeks, I feel like living proof that this is a skill you can improve with practice.

I think of myself as the showrunner. My job is to write the plan my agents work against, give them what they need, and enforce the results myself. The models always report out “it worked!” So I learned to set up every build to produce something I can check against reality, and move only as fast as I can verify it. That’s the question from my first projects, answered.

My advice to builders: try the most ambitious thing you have any hope of getting working, and expect that ceiling to rise as you and the models improve. Building every two weeks, with the community behind me, is how mine rose.