Search

Showing top 137 results for "NVFP4"

Related topics: NVFP4

Tracked topic

NVFP4

Live now An NVFP4 quant of Qwen3.8 27B is being touted as the fastest available right now, based on user-perceived performance.
Live sources: r/LocalLLaMA
1 live sources Latest signal 1d ago See topic hub
developer.nvidia.com › ja-jp › blog

Nemotron 3 Super の紹介: エージェント型推論向けのオープン ハイブリッド Mamba-Transformer MoE

…これにより、パラメーターのオーバーヘッドを最小限に抑えながら、トレーニングの安定性を向上させます。ヘッドは、オフセット固有のショートカットに分裂するのではなく、一貫した継続に合意できるようになります。 同様の重み共有により、独立してトレーニングされたヘッドが通常劣化するような長いドラフト長においても、投機的ドラフトの一貫性が向上します。 ネイティブ NVFP4 事前トレーニング ほとんどの量子化モデルは、全精度で計算を開始し、トレーニング後に圧縮されるため、精度の低下は避けられないものです。 Super では別のアプローチを採用しています。事前トレーニング中の浮動小数点乗算/累積演算の大部分は NVFP4 、すなわち NVIDIA 4 ビット浮動小数点形式で実行されています。 Blackwell 向けに最適化されたこの手法は、精度を維持しながら、FP8 と比較してメモリ要件を大幅に削減しながらも、推論を高速化します。 低精度でネイティブにトレーニングを行うことは…

Mar 11, 2026 · Chris Alexiuk

Discussions and forums

r/LocalLLaMA · u/ionsago · 1d ago

Fastest NVFP4 quant of Qwen3.8 27B out there

Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked …

r/LocalLLaMA · u/Simple_Library_2700 · 1w ago

366 t/s Qwen3.6 27B NVFP4 on v100s

These are single stream numbers Following on from my previous post about v100s (here) and inspired by this comment (here) I decided to work on kernels that allow for an extremely fast path for Nvfp4 weights on sm70 and a…

r/LocalLLaMA · u/Simple_Library_2700 · 2d ago

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Four Tesla V100s from 2017 matched my RTX 5090 on single-request Qwen 3.8 decode. Repo: https://github.com/dnv2003/v100-skinny https://i.redd.it/5ws2ak3uqckh1.gif The 5090 was not being held back. It ran NInfer, a specia…

r/LocalLLaMA · u/LLMFan46 · May 26, 2026

Qwen3.5 35B A3B uncensored heretic Native MTP Preserved is Out Now With the Full 785 MTPs Preserved and Retained, Available in Safetensors, GGUFs. NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats

Safetensors, llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved: https://huggingface.co/llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved GGUFs, llmfan46/Qwen3.5-35B-A3B-uncensored-here…

r/LocalLLaMA · u/rmhubbert · 3w ago

New official weights for Laguna S 2.1 FP8 & NVFP4 are now available

Poolside have updated the FP8 and NVFP4 checkpoints for Laguna S 2.1, increasing the default context size to 1 million, and updating the configs. Here's hoping they fixed the looping issue, this model has been great in m…