Full Sail on Asynchronous Inference
Today all inference is real-time. A human types, a model responds, & the clock starts over. The infrastructure is built for someone waiting on the other end. Every millisecond of latency costs money because the serving stack optimizes for cold-start, not throughput.
As we built internal AI systems at Theory, we embraced queueing. Parallelize ten agents on a single task, let them run for hours, & the productivity gains are enormous. It is the product of token-maxxing,1 pushing every dollar of compute to do more work. But the cost was unsustainable.
That is when we met Neil Movva & Samir Menon of Sail Research.2