RunPod launcher for AGPT experiments
For the current operational guide, see
HOWTO.mdin this directory. This README documents history and rationale; HOWTO.md is the day-to-day cookbook. Of particular note: thedocker.io/7rans/agptimage was rewritten on 2026-05-21 from an Arch base tonvidia/cuda:12.4.1-devel-ubuntu22.04because RunPod's typical hosts run NVIDIA driver 550, which caps at CUDA 12.4 — Arch's rolling CUDA 13.2 was rejected with "driver version is insufficient" on every pod. The image now pins CUDA 12.4.
Lets you provision an H100/H200/A100 RunPod instance, sync code+data, run experiments, and pull results back — without manually maintaining the env.
One-time setup on your end
-
Generate an SSH key for RunPod if you don't have one:
ssh-keygen -t ed25519 -f ~/.ssh/runpod -N "" cat ~/.ssh/runpod.pub # add this public key to your RunPod account (Settings → SSH Public Keys) -
(Optional) Add a config block to
~/.ssh/configso you don't have to pass the key every time. RunPod will give you a command line likessh root@69.30.85.92 -p 18234 -i ~/.ssh/runpod; the equivalent config:Host runpod-current HostName 69.30.85.92 Port 18234 User root IdentityFile ~/.ssh/runpod
Provisioning a pod
In RunPod's UI:
- Pick a GPU: A100 SXM 80GB (~$1.89/hr, recommended — usually available), H100 SXM (~$2.49/hr, when in stock), H200 (141GB, ~$3.59/hr, often gone)
- Pick an image:
- Recommended:
docker.io/7rans/agpt:YYYY-MM-DD(date-tagged immutable version) — Arch-based, has CUDA + Crystal + all AGPT binaries pre-built. Pod is ready to run experiments immediately aftersetup-imagersyncs runtime data. Avoids the ~90 min / ~$3 provision time we hit on 2026-05-17 with the generic pytorch image. Use the date tag (e.g.:2026-05-21) rather than:latestso the pod is reproducible —:latestis a moving target.podman pull docker.io/7rans/agpt:latestlocally to see the current digest; tagged versions live alongside it. Build instructions:rnd/docker/README.md. - Fallback:
runpod/pytorch(or any CUDA 12+ image) — requires the manual install flow viasetup_pod.sh. Use when the image isn't available, drifted from current code, or you want a clean env.
- Recommended:
- Set persistent volume to 50 GB (corpora + tries + models)
- Add "TCP Port 22" to Exposed TCP Ports. Required for the
launch.shautomation (rsync + non-interactive ssh). RunPod will map your container's port 22 to a custom external port (~3xxxx). Thessh.runpod.ioproxy path works only for interactive shells — not forssh user@host 'command'or rsync, which is whatlaunch.shdoes. - Note the SSH command they give you. After adding TCP 22, RunPod's
"Connect" tab shows TWO commands: the proxy one (
ssh user@ssh.runpod.io) for interactive use, and the direct one (ssh root@<IP> -p <PORT> -i <key>) for automation. Use the direct one withlaunch.sh.
Using the launcher
From your laptop, in the agpt project root:
# Recommended path — pod provisioned from docker.io/7rans/agpt:latest:
bash rnd/runpod/launch.sh setup-image root@69.30.85.92:18234
# Fallback — pod provisioned from generic pytorch image (rsync code + install):
bash rnd/runpod/launch.sh setup root@69.30.85.92:18234
# Run an experiment:
bash rnd/runpod/launch.sh run root@69.30.85.92:18234 \
"CORPUS=\$PWD/data/gutenberg_5m.txt bash rnd/streaming-agpt-v1/run_multiseed_baseline.sh 500 200 300"
# If your laptop has local code changes you want the pod to use (works
# alongside setup-image):
bash rnd/runpod/launch.sh push-code root@69.30.85.92:18234
bash rnd/runpod/launch.sh run root@69.30.85.92:18234 "just build-agpt-train"
# Pull results back to laptop:
bash rnd/runpod/launch.sh pull root@69.30.85.92:18234
# Or all-in-one (manual install path):
bash rnd/runpod/launch.sh full root@69.30.85.92:18234 \
"CORPUS=\$PWD/data/gutenberg_5m.txt bash rnd/streaming-agpt-v1/run_multiseed_baseline.sh 500 200 300"
launch.sh full does setup + run + pull in sequence. Best for one-off
experiments. launch.sh setup + launch.sh run separately is better when
you want to leave the pod running and fire off multiple experiments.
What gets transferred
To the pod:
- All Crystal/CUDA source (
src/,lib/microgptbuilds from shard on the pod) - All experiment scripts (
rnd/streaming-agpt-v1/*.sh) data/input.txt,data/gutenberg_5m.txt(corpora)data/input.random.model(random init checkpoint)/tmp/init_seed{100,200,300}.model(seeded init checkpoints)- Notes (
notes/*.md)
Excluded (not sent):
bin/,build/,lib/— rebuilt fresh on the podrnd/streaming-agpt-v1/models/— gitignored, regenerated locally to podrnd/streaming-agpt-v1/logs/— pulled BACK from podrnd/seq-len-decouple/*.bin— the 205MB position mapsdata/wormhole_*.txt— large research artifacts
Pulled back:
- All training logs and PPL eval outputs from
rnd/streaming-agpt-v1/logs/ - Updated
findings.mdif the pod modified it
Cost estimate per session
| GPU | $/hr | typical experiment | wall time | cost |
|---|---|---|---|---|
| A100 SXM 80GB | $1.89 | 2 Gutenberg baseline seeds | ~3 hr | ~$6 |
| A100 SXM 80GB | $1.89 | 3 streaming + 3 baseline seeds (1000 SE each) | ~12 hr | ~$23 |
| H100 SXM | $2.49 | Same as above | ~7 hr | ~$18 |
| H200 | $3.59 | Same as above | ~6 hr | ~$22 |
A100 SXM is the most available tier. Bandwidth-bound AGPT workload (KV gather) sees ~2× speedup on H200 over A100 thanks to HBM3e (4.8 TB/s) vs HBM2e (1.94 TB/s). For our small models the matmul speedup is limited; memory bandwidth is the bigger win when we can get it. But total cost differences are within noise once you account for availability — A100 SXM is the practical default.
Notes on building on the pod
setup_pod.shinstalls Crystal + just + builds with the same flags as laptop:-O3 --use_fast_math -gencode=sm_89,sm_90.- It applies
microgpt_tf32.patchso the cuBLAS path on microgpt also gets TF32 (sincelib/microgptis fetched fresh from the shard). - The build is fast (~2-3 min on RunPod since they have fast disks + recent CUDA).
- First-time setup including all deps: ~10-15 min.
Don't forget to stop the pod
RunPod bills per second. After launch.sh pull succeeds:
- Verify the results landed on your laptop (
ls rnd/streaming-agpt-v1/logs/). - Stop the pod via RunPod UI (Stop, not Delete, if you want to keep the volume for next session).