Nvidia paid $20 billion for SRAM decode - AMD just partnered for it instead
… It must, however, be noted that the efficiency claims of 5 tokens per second per watt are compared against an existing Cerebras WSE Wafer-Scale Engine as the baseline, while running the open-source Kimi 2.6 1T model, making them impressive, but without a direct comparison to figures for an Nvidia r… …
