OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch
…It’s very efficient to serve a lot of customers, but it can also be very low latency.” Notably, that comparison is against an Nvidia Blackwell system — but by the time Jalapeño…
