研究Research
flower:把 Agent 跑过的计算整理清楚flower: keeping scientific-agent calculations reproducible
最近用 Agent 做计算,越来越容易碰到一个问题。它可以自己读论文、改代码、提交任务,一晚上跑出很多结果,第二天打开工作目录却常常是一堆脚本、日志和不同版本的文件。每个结果似乎都有用,真要找出哪一步算出了哪张图、用的什么参数,又得重新翻一遍。
我在做一个叫 flower 的小工具,想把这些计算顺手记录下来。每项研究可以写成一份包含若干步骤的计划:这一步运行什么命令、在哪里运行、应当得到什么结果。Agent 或人都可以在执行时更新计划,flower 会把每次运行、修改和结果留在日志里。计算可以放在本机,也可以交给 SSH 机器或 Slurm 集群。
等到某个结果可以用了,就把通向它的计算步骤导出成一个 protocol。其他人拿到这个流程,可以重新运行并核对输出。仓库里已经有这样的示例,包括一套重复计算二维 Heisenberg 模型结果的流程。
它当然没办法替你判断论文里的物理解释是否正确,也不能保证换一台机器就不会遇到依赖或数值差异。不过,如果至少能看清 Agent 究竟做过什么,知道一个结果是怎样算出来的,后面讨论和检查就容易多了。
Lately, when I use agents for scientific computing, I keep running into the same problem. An agent can read papers, edit code, submit jobs, and produce many results overnight. The next morning, the working directory is full of scripts, logs and different versions of files. Some of the results look useful, but finding out exactly which step produced a figure, and with which parameters, takes another round of digging.
I’m building a small tool called flower to keep track of that work as it happens. A study is a plan made of steps: the command to run, where it runs, and what output to expect. An agent or a person can update the plan, while flower keeps a record of the runs, changes and results. Jobs can run locally, over SSH, or on a Slurm cluster.
Once a result is worth keeping, the steps leading to it can be exported as a protocol. Someone else can rerun it and compare the outputs. The repository includes examples, including a workflow that reproduces results for the two-dimensional Heisenberg model.
Of course, a recorded workflow cannot tell us whether the physics in a paper is right. It also cannot promise that a different machine will have identical dependencies or numerical behavior. But being able to see what an agent did, and how a result was obtained, makes it much easier to review the work afterward.