The Infrastructure Bet Nobody's Talking About: Why I'm Watching Silicon Valley's Hardware Shuffle
I spent last Tuesday debugging why our inference pipeline was hitting latency spikes at 3 AM, and my first instinct was to throw more Nvidia H100s at the problem. By Wednesday, I read that OpenAI just dropped a custom chip for inference. By Thursday, I realized I'd been thinking...