Search

Showing top 137 results for "NVFP4"

Related topics: NVFP4

Tracked topic

NVFP4

Live now An NVFP4 quant of Qwen3.8 27B is being touted as the fastest available right now, based on user-perceived performance.
Live sources: r/LocalLLaMA
1 live sources Latest signal 1d ago See topic hub
developer.nvidia.com › ja-jp › blog

NVIDIA 技術ブログ

…エージェント型推論向けのオープン ハイブリッド Mamba-Transformer MoE Nemotron 3 Super は、高容量の推論モデルにおける典型的な効率と精度のトレードオフを軽減するアーキテクチャ革新を導入しています。 3 MIN READ 2026 年 2 月 6 日 NVFP4 が AI のトレーニングと推論を加速する 3 つの方法 NVIDIA による徹底的な共同設計によって、モデルのトレーニングと推論の両方において、優れた精度で大幅なパフォーマンスの向上が達成が見込めるようになりました。 2 MIN READ…

Discussions and forums

r/LocalLLaMA · u/ionsago · 1d ago

Fastest NVFP4 quant of Qwen3.8 27B out there

Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked …

r/LocalLLaMA · u/Simple_Library_2700 · 1w ago

366 t/s Qwen3.6 27B NVFP4 on v100s

These are single stream numbers Following on from my previous post about v100s (here) and inspired by this comment (here) I decided to work on kernels that allow for an extremely fast path for Nvfp4 weights on sm70 and a…

r/LocalLLaMA · u/Simple_Library_2700 · 2d ago

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Four Tesla V100s from 2017 matched my RTX 5090 on single-request Qwen 3.8 decode. Repo: https://github.com/dnv2003/v100-skinny https://i.redd.it/5ws2ak3uqckh1.gif The 5090 was not being held back. It ran NInfer, a specia…

r/LocalLLaMA · u/LLMFan46 · May 26, 2026

Qwen3.5 35B A3B uncensored heretic Native MTP Preserved is Out Now With the Full 785 MTPs Preserved and Retained, Available in Safetensors, GGUFs. NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats

Safetensors, llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved: https://huggingface.co/llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved GGUFs, llmfan46/Qwen3.5-35B-A3B-uncensored-here…

r/LocalLLaMA · u/rmhubbert · 3w ago

New official weights for Laguna S 2.1 FP8 & NVFP4 are now available

Poolside have updated the FP8 and NVFP4 checkpoints for Laguna S 2.1, increasing the default context size to 1 million, and updating the configs. Here's hoping they fixed the looping issue, this model has been great in m…

developer.nvidia.com › ko-kr › blog

NVIDIA Nemotron 3 Nano Omni: 단일 오픈 모델로 멀티모달 에이전트 추론을 가속화

…또한 FP8과 NVFP4 양자화 , 효율적인 비디오 샘플링, NVIDIA 최적화 커널을 지원해 예측 가능하고 지연 시간이 낮은 추론을 제공합니다. 여기에 3D 컨볼루션 기반 시공간 처리가 결합되면 워크스테이션부터 데이터센터, 클라우드 배포 환경까지 GPU 전반에서…

May 12, 2026 · Anjali Shah
developer.nvidia.com › ja-jp › blog

NVIDIA Jetson でメモリ効率を最大化して大規模なモデルを実行

…重要なポイントが 1 つあるとすれば、適切な量子化の精度を使用することです。 NVFP4、INT4、W4A16 などのフォーマットは、多くの LLM ワークロードで高い精度を維持しながら、メモリとストレージの要件を大幅に削減します。 実際のユース ケース: Reachy Mini Jetson Mini Assistant これらのメモリ最適化の効果を示すために、Jetson Orin Nano 上で実行されるオンデバイス対話型 AI ロボットである Reachy Mini Jetson Assistant を考えてみましょう。これは…

Apr 20, 2026 · Anshuman Bhat