# Gabriel Perácio

Gabriel Peracio is a staff engineer at Parasail, specializing in ML inference optimization. He has been fascinated by language models since a herd of unicorns spoke perfect English in 2019.

## Articles by Gabriel Perácio

1. [**Beyond the frontier: How to build a defensible AI inference infrastructure**  
Closed model dependency is becoming a structural liability. Here's the framework for building a reliable AI inference architecture.  
Gabriel Perácio · July 6, 2026](/content/blogs/defensible-ai-inference-infrastructure/index.html)

2. [**The idle GPU tax: What it is, why it’s getting worse, and how you can fix it**  
Learn what the idle GPU tax is, what it costs, and how usage-based billing on dedicated endpoints helps you avoid it altogether.  
Gabriel Perácio · June 18, 2026](/content/blogs/idle-gpu-tax/index.html)

3. [**How to choose the right managed inference architecture: Serverless, dedicated, dedicated serverless, or batch**  
Use this decision framework to choose the right managed inference mode based on latency requirements, GPU breakeven utilization, and whether your workload needs a dedicated endpoint.  
Gabriel Perácio · June 9, 2026](/content/blogs/how-to-choose-the-right-inference-architecture/index.html)

4. [**Serverless vs. Dedicated Inference: Why We Built Dedicated Serverless**  
With dedicated serverless you get dedicated hardware on per-token pricing, no idle-hour charges or long-term GPU commitment.  
Gabriel Perácio · May 29, 2026](/content/blogs/serverless-vs-dedicated-inference/index.html)

5. [**Making an EAGLE fly: How We Got 2.6x Faster LLM Inference (Without Cheating)**  
We trained a custom EAGLE-3 speculative decoding head for OLMo-3.1-32B-Think and got 2.6x faster inference.  
Gabriel Perácio · April 28, 2026](/content/blogs/making-an-eagle-fly-how-we-got-2-6x-faster-llm-inference-without-cheating/index.html)
