Planning a four-GPU expansion box for my AM5 workstation
My plan to connect an AM5 workstation to four GPUs through a Broadcom PEX88096 backplane, with PCIe P2P between the cards and a shared x16 connection to the CPU.
My plan to connect an AM5 workstation to four GPUs through a Broadcom PEX88096 backplane, with PCIe P2P between the cards and a shared x16 connection to the CPU.
One week of GPT-6 Astra directing local pi subagents through CUDA kernel research on a $100/month OpenAI Pro plan: $475.97 of metered-equivalent orchestrator usage that saturated my weekly limit, plus $1,416.57 of work done locally that the plan had no room for. The enabler is a 2 KiB RETURN.json cap that kept 29 GB of worker transcript out of the orchestrator's context.
Paseo to run and watch agents from any device, Pi as the harness, and four extensions (web search via Gemini, Chrome, sudo, sub-agents) that make local models useful.
Qwen3.8-Flash-Next narrates its own NVFP4 install on a 2x RTX Pro 6000 (Blackwell) workstation: the dedicated vLLM recipe image, tensor-parallel launch on sm_120, and the n-gram embedding table offloaded to system RAM so the weights finally fit in 192 GB of VRAM.
I converted all 16 issues of Game Over, Romania's first PC gaming magazine, from raw scanned PDFs into a searchable website. Plain OCR cannot do this because of the layout. I tested six models and Opus 5 was the only one that could.
From 100A to 200A service: running dedicated 120V and 240V circuits for a multi-GPU AI rack, with NEMA L6-30R twist-lock receptacles, a Seasonic PRIME PX-2200 PSU, and a 240V PDU.
Foundation models like TimesFM forecast numerical time series zero-shot. I tested them on payment transaction volume, added holiday covariates, and learned where they help and where they hurt.
Run DeepSeek-V4-Flash on a 2x RTX Pro 6000 (96GB each) workstation using the voipmonitor/vllm:lucifer Docker image, a Blackwell-targeted vLLM fork with sm_120 kernels, FP8 KV cache, and MTP speculative decoding.
Run Qwen3.6-27B, Gemma 4 31B, and MiniMax M2.7 locally, then connect them to the Pi coding agent for local coding.
I gave Luna access to my email. Here's what happened in a single day.