NVIDIA·graphics card
GeForce RTX 5090
Based on 45 credible posts from Hacker News & Bluesky. 15 filtered out as bot-like or off-topic.
Medium confidenceA solid sample; expect small shifts as more owners post. How we score
Scored September 27, 2026
We may earn a commission from this link. It never affects the score.
Owner reviews
Overall owner mood
Most owners are happy with it: 67% of 45 reviews are positive and 9% negative. Praise centers on price & value. The most common complaint is performance (9 owners).
What owners like
Price & value praised by 3 owners
“So in terms of raw power Nvidia is effortlessly still king, but in price-to-capacity Intel is best in class.”
What owners complain about
Performance raised by 9 owners
“And all the heuristics in the prompt degraded the performance and increased the cost for frontier models.”
Software & AI raised by 2 owners
“Everything is riddled with loading screens, lost inputs, freeze frames and janky scrolling etc.”
Price & value raised by 2 owners
“This would have been prohibitively expensive now.”
What owners talk about
- Performance2 praise · 9 complain
- Price & value3 praise · 2 complain
- Software & AI1 praise · 2 complain
Built from the reviews themselves — every quote links to the owner who wrote it.
Full text of every credible post we scored. Bot-like and off-topic comments are hidden.
Really happy I got a new PC last year with 128 GB of RAM, 4 TB SSD, and an RTX 5090. This would have been prohibitively expensive now.
Hi, I optimized a world model, Lingbot-World 2.0 1.3B, to run with real-time 16ps on an RTX 5090. [...] It maintains lossless performance while running on a 1x consumer GeForce GPU for a model that's very compute-bound and batch size = 1! [...]
On many models that I tested in past context quantization had very bad effect on model performance. [...] I'm now running NVFP4 quantized both weight and cache on my RTX5090 and getting excellent results: 264k cache allocated for pool, 10k tok/s prompt processing, 200 tok/s generation for single stream, or 801 tok/s generation for 8 concurrent streams. [...]
I have a RTX PRO 6000 96GB when the pricing was way better than now i also have a RTX 5090 too. [...] in diffusion) is increasing faster than the need for more VRAM - additional reason for the value of these FAST GPUs to increase! (3) RTX PRO 6000 96GB is really great for fine tunes (ai-toolkit) :) but doesn't outperform my RTX 5090 with inference by anything significant on the good local models. I have never run an AI job on a Mac, i also have doubts about performance and compatibilities - since the reviews almost never compare directly.
Memory used : 38GB, and I haven't even started a LLM nor podman, I always fight with memory when using LLM on my mac with 48gb. And I don't remember to have been able to have pushed to 200k context Qwen 3.6. 3.8 is running on my RTX 5090.
[...] That said I have an RTX 5090, not a Mac Mini, so it's not exactly the same level of performance... [...]
I have a new laptop with a rtx 5090. Opening any GL context takes more than half a second. There's tons of things that can be optimized and are pretty far from web.
Using a RTX 5090 ($34k), I can run Qwen 3.8 27B at 180 TPS with ninfer [0]. With its thinking maxed out, I can confirm that the quality of output is roughly on par with Opus 4.54.6 - that is, this 20GB file really can write software by itself, but the amount of thinking required makes it strictly slower than larger models, even at 180 TPS. Still, it's largely replaced the cheap tier of the frontiers that I would otherwise be using. [...] However, it is worth noting that, if you are buying this hardware just to serve LLMs, it is not cost-effective. [...] I'm running this setup because I happen to have a 5090, and the two 3090s in my server were cheap enough when amortised over several years. [...]
[...] On Gemma 4 12B, a small open-weight model, running on RTX 5090, compiler-generated kernels achieved up to 1.6× speedup over cuBLASLt, and a 1.30× geometric-mean speedup over PyTorch. Primarily, two optimizations allowed the Emmy compiler to outperform cuBLAS on common GEMM and FA shapes on RTX 5090 and RTX 4090: TMA transport (Blackwell-only) and leveraging full FP16 tensor cores with FP32 shadow accumulation registers: TMA Transport for Matmul and Flash Kernels. [...]
[...] On Gemma 4 12B, a small open-weight model, running on RTX 5090, compiler-generated kernels achieved up to 1.6× speedup over cuBLASLt, and a 1.30× geometric-mean speedup over PyTorch. Primarily, two optimizations allowed the Emmy compiler to outperform cuBLAS on common GEMM and FA shapes on RTX 5090 and RTX 4090: TMA transport (Blackwell-only) and leveraging full FP16 tensor cores with FP32 shadow accumulation registers: TMA Transport for Matmul and Flash Kernels. [...]
[...] I'm on an RTX 5090. [...]
[...] One day he said "Well, I have an RTX 5090 doing nothing... [...]
On my RTX 5090, Qwen3.6-35B-A3B is my go to for coding, but Gemma-4-26b-a4b (Q40) is blazing fast and super high quality. I sometimes use the dense cousins of these models (Qwen3.6-27B QB0 and Gemma4-31B-QAT Q40) but they can be very slow. [...] Here are the tok/s I get: - Gemma-4-26B-A4B (Q40) = 214 tok/s - Gemma4-31B-QAT (Q40) = 58 tok/s - Qwen3.6-35B-A3B (QB0) = 30 tok/s - Qwen3.6-27B (QB0) = 9 tok/s EDIT: Updated tok/s after updating llama.cpp
I have access to both a RTX 5090 PC with 64gb of ram and a Spark 128gb, the performance of the Spark has been highly disappointing. I prefer 1000x the RTX one, even with 64gb of ram.
RTX 5090 with 610.62 drivers, Windows 11 25H2, Firefox 153.0b9. [...]
[...] The nontechnical staff use a Gemini account on a Google family AI Pro sub. [...] RTX 5090 on a PC and a Mac Studio. [...]
The quality of what you can get from DeepSeek V4 Pro for $10 is light years ahead of what you could get for $20 a year ago. Likewise, the quality of what I can get from a local model like Qwen 3.6 on an RTX 5090 is light years ahead of what I could get a year ago on the same hardware.
[...] An RTX 5090 is 61 t/s at $33 $/t/s. So in terms of raw power Nvidia is effortlessly still king, but in price-to-capacity Intel is best in class. [...]
My LLM host is a Framework Desktop 128GB (Ryzen AI 395) with an RTX 5090 (Minisforum DEG1 dock, cheap Oculink card plugged into the PCIe 4.0 x4 slot). I run llama-server (b9296 compiled for CUDA) using a systemd service, and its configured as: server-2.env LLAMAARGS="--host 0.0.0.0 --port 8082 --sleep-idle-seconds 1800" LLAMAARGMODELSMAX=1 LLAMAARGMODELSPRESET=/usr/local/etc/llama.cpp/models-2.ini models-2.ini [] device = CUDA0 fit = on cache-type-k = q80 cache-type-v = q80 [Qwen3.6 27B It Q6] model = /mnt/data1/llm/huggingface/hub/models--unsloth--Qwen3.6-27B-GGUF/snapshots/82d411acf4a06cfb8d9b073a5211bf410bfc29bf/Qwen3.6-27B-UD-Q6KXL.gguf alias = qwen3.6-27b-it:q6 ctx-size = 200000 temp = 0.6 top-p = 0.8 top-k = 20 min-p = 0.0 presence-penalty = 1.5 repeat-penalty = 1.0 chat-template-kwargs = {"enablethinking":false} I use the model from a sandboxed OpenCode using the @ai-sdk/openai-compatible plugin. And I'm using an SSH tunnel from home to work to expose llama-server to OpenCode.
I have a very similar setup: OpenCode, Qwen3.6-27B (llama-server with an RTX 5090). It works well for my purposes. As a semi coding luddite, I use it for mundane tasks.
110+ Tok/s as another data point on the RTX 5090 (Gemma 4 31B QAT + MTP at UD-Q4KXL) (at peak used 27 GB of vram) The real lovely thing was getting 300+ Tok/s (Gemma 4 26B QAT + MTP at UD-Q4KXL) (at peak, I think I saw vram usage reach 21 GB of vram)
Here are the prefill speeds: Device 0: NVIDIA GeForce RTX 5090, compute capability 12.0, VMM: yes, VRAM: 32109 MiB | model | size | params | backend | fa | test | t/s | | ------------------------------ | ---------: | ---------: | -------- | --: | --------------: | -------------------: | | qwen35 27B Q4K - Medium | 15.92 GiB | 27.32 B | CUDA | 1 | pp2048 @ d512 | 3714.02 ± 10.85 | | qwen35 27B Q4K - Medium | 15.92 GiB | 27.32 B | CUDA | 1 | pp2048 @ d1024 | 3684.86 ± 15.21 | | qwen35 27B Q4K - Medium | 15.92 GiB | 27.32 B | CUDA | 1 | pp2048 @ d2048 | 3650.80 ± 8.53 | | qwen35 27B Q4K - Medium | 15.92 GiB | 27.32 B | CUDA | 1 | pp2048 @ d8192 | 3473.88 ± 0.97 | | qwen35 27B Q4K - Medium | 15.92 GiB | 27.32 B | CUDA | 1 | pp2048 @ d32768 | 2754.69 ± 4.07 | ggmlmetaldeviceinit: GPU name: MTL0 (Apple M2 Ultra) | model | size | params | backend | fa | test | t/s | | ------------------------------ | ---------: | ---------: | -------- | -: | --------------: | -------------------: | | qwen35 27B Q80 | 26.62 GiB | 26.90 B | MTL | 1 | pp2048 @ d512 | 379.75 ± 0.21 | | qwen35 27B Q80 | 26.62 GiB | 26.90 B | MTL | 1 | pp2048 @ d1024 | 377.15 ± 0.35 | | qwen35 27B Q80 | 26.62 GiB | 26.90 B | MTL | 1 | pp2048 @ d2048 | 371.46 ± 0.91 | | qwen35 27B Q80 | 26.62 GiB | 26.90 B | MTL | 1 | pp2048 @ d8192 | 344.84 ± 0.41 | | qwen35 27B Q80 | 26.62 GiB | 26.90 B | MTL | 1 | pp2048 @ d32768 | 222.42 ± 5.29 | Btw, based on your numbers, I think our use cases are quite different. [...]
About the generation speed: 100-150 t/s on the RTX 5090 and 40 t/s on the Mac Curious if you can share the prefill speed too? [...]
[...] Over the last month and a half I've been using it almost daily, either on my M2 Ultra or on my RTX 5090 box. I use it for small mundane tasks at ggml-org [0] - nothing really impressive, but definitely a helpful tool for a maintainer. [...] About the generation speed: 100-150 t/s on the RTX 5090 and 40 t/s on the Mac. [...]
Same here, I use Qwen 3.6 27b (Q6 quant) with llama.cpp on an RTX 5090 using the pi agent exclusively now. [...]
RTX 5090, just because it dual boots as a gaming rig. I have a llama server usually with gemma4:31b or qwen3.6:27b that powers my own automation / orchestration harness. If I can go back in time, I would probably buy a more AI dedicated machine but I also don't regret finally being able to play Cyberpunk in 4k with great FPS and overkill mods.
[...] My RTX 5090 runs Trellis2 - ultrashapes - Trellis2 - wiring up rigging and setting up animations. [...]
[...] Getting cutesie stylized 3D models is something that’s trivial with an RTX 5090, a ChatGPT pro subscription (unlimited image generation), you run Trellis2 plus a few other open source things in a pipeline that your agents can queue and it’s astonishing how much cool stuff comes out the other side. [...]
[...] If you have a powerful GPU like an RTX 5090, try Qwen-3.6 locally on that too. [...]
[...] I'll have to look into that project, but I also have an RTX 5090 and did a lot of testing with Qwen3.6 27B and Gemma 4 31B. [...] And all the heuristics in the prompt degraded the performance and increased the cost for frontier models.
Awesome ! Does this use , or is it a separate project? I ran 4 local models in a tournament recently with mage-bench on an RTX 5090 ; Qwen 3.6 27B won narrowly over Gemma 4 .
No, it's an Alienware R51 with Intel Core Ultra 9 285K 3.2GHz Processor; NVIDIA GeForce RTX 5090 32GB GDDR7; 64GB DDR5-6400 RAM; 4TB Solid State Drive; Microsoft Windows 11 Home; 2.5GbE LAN; 2x2 Intel Killer WiFi 7 BE1750+Bluetooth 5.4; Liquid Cooler I don't see it on the Dell site anymore, only more expensive, lesser configurations (good timing on my part?). Yeah, I really want to put in the time to try out various games, but realistically, the whole point of getting a second computer and installing Linux was to be able to train and serve models, and switching between serving a model (that people in my house want to use at random times) and gaming didn't seem like a great choice. If I did get good results, I'd seriously consider wiping Windows 11 from my older machine (an older Alienware with a 4090), but to be honest, I'm perfectly comfortable on Windows desktop.
[...] A few months ago I used Whisper from OpenAI, an automatic speech recognition system released in 2002, on my modern 20-core Intel CPU to convert audio from a video file to text. [...] They are not NVIDIA GeForce RTX 5090, but they are significantly better than a modern CPU.
[...] What exactly is the software we are comparing? [...] Everything is riddled with loading screens, lost inputs, freeze frames and janky scrolling etc. [...] Even OS-level and professional software. I now have a AMD Ryzen 9 9950X3D CPU, GeForce RTX 5090 GPU, DDR5 6000MHz RAM and M.2 NVME disks. I should not even see any loading screen, or any operation taking longer than a second. [...]
Across a 20-benchmark suite spanning reasoning, math, coding, instruction following, vision, and agentic tool use, Ternary Bonsai 2 27B achieves an aggregate score of 83.9, retaining 98.2% of Qwen3.8 27B’s performance of 85.4. The model reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090. It's 6 GB for "98.2% of Qwen3.8 27B's performance". [...]
I’ve run four models: Qwen 3.8 Max, GLM 5.3, Gemini 3.7 Flash, and a 27B Qwen 3.8 running locally on my RTX 5090. [...]
If I could just save up $6000 I could sell off my RTX 5090 for $4,000 and buy an RTX 6000 Blackwell Pro Workstation. I can fit models into the 32GB of vram but my context window ends up being tiny for any halfway capable model.
I have tried in both my Mac and my desktop (Rtx 5090) with Gemma 4 and Qwen and so far nothing is quite replacing Claude Code or Kiro for spec driven architecture & development. I do think we are slowly getting Gemma 4 was a big jump
[...] Qwen 27B is also small enough to completely fit in a high-end consumer or mid-end pro GPU, like an RTX 5090 or Radeon PRO R9700. [...] Even for relatively short contexts, I honestly already find the 30B class MoE models to be only borderline acceptable in terms of speed on my laptop (Ryzen 7 7840U, 64 GB LPDDR5-6400), though I use Gemma 4 26B-A4B more than Qwen3.6 35B-A3B.
[...] Datacenter hardware is expensive and there's shortage of it but llama.cpp is slow/unoptimized for concurrent use, while vLLM/SGLang easily crash on non-common setups (things like, if you do pipeline parallelism for RTX5090+RTX4090, they will randomly crash with RAM caching enabled or select wrong kernels because they usually assume that every rank is the same device type; they also don't support Q5-Q6). [...] I've been running an AI server in the office, and so far I've find these techniques most important for concurrent use on cheap hardware: pipeline parallelism (to accomodate for PCie), RAM caching (to quickly restore contexts into VRAM), speculative decoding (including domain-specific ngrams, they already can speed up code generation considerably without the overhead of a draft model), good kernels highly optimized for a specific device, support for Q5-Q6 (almost as good as Q8), FP8 contexts (more context to fit), paged attention (for better VRAM utilization), prefix caching, continuous batching (this is the default everywhere). So far the main bottlenecks have been llama.cpp's poor VRAM utilization for contexts (you either have fixed-size slots, or use unified KV cache where each request attends to attention from all other requests and then unnecessary portions of attention are masked out), and lack of decode/prefill segregation: when a request starts prefilling a long context, all decoding threads slow down to like 5 tok/sec. [...]
[...] After comparing a Radeon 7900 XTX vs Ryzen Halos 128GB vs M1 MacBook Pro 64Gb, I landed on just getting an external closure setup with Nvidia RTX 5090. [...] The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). [...] Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card. The only thing I can point to slowing me down is bandwidth of the card itself. [...]
[...] I have a RTX 5090 and 64 GB of RAM but the models seem to be much larger than that. [...]
[...] vLLM/SGLang attention falls back to FA-2 on consumer cards (FA-3 and FA-4 are datacenter-only), so I wanted to know if there's any performance left on the table, and I rebuilt the attention kernels from scratch. The kernel reaches parity with FA-2 (206us on RTX5090 with batch=1, heads=8, seqlen=4096, headdim=64), but unfortunately, FA-3/4 optimizations are either not applicable or not helpful on consumer cards. [...] RTX 5090 is tensor-core bound, so no point in this optimization either. [...]
[...] Qwen-3.6-35b-q6, runs well on an RTX 5090 ($4000 + cost of a PC), runs medicore on an Intel Arc B70 ($1000 + cost of a PC plus lots of fiddling to get the setup to work right). Gemma is a good candidate for the cheaper stuff, but I lack personal experience with using it locally
I have an RTX 5090 card but it only has 32 GB RAM, can something like this work on my machine?
The claims, and the evidence
What the brand claims
No published claims on file for this product yet.
What owners report
“Really happy I got a new PC last year with 128 GB of RAM, 4 TB SSD, and an RTX 5090. This would have been prohibitively expensive now.”
after 1 yearView on Hacker News“I have a very similar setup: OpenCode, Qwen3.6-27B (llama-server with an RTX 5090). It works well for my purposes. As a semi coding luddite, I use it for mundane tasks.”
“[...] Qwen 27B is also small enough to completely fit in a high-end consumer or mid-end pro GPU, like an RTX 5090 or Radeon PRO R9700. [...] Even for relatively short contexts, I honestly already find the 30B class MoE models to be only borderline acceptable in terms of speed on my laptop (Ryzen 7 7840U, 64 GB LPDDR5-6400), though I use Gemma 4 26B-A4B more than Qwen3.6 35B-A3B.”
“I have a new laptop with a rtx 5090. Opening any GL context takes more than half a second. There's tons of things that can be optimized and are pretty far from web.”
Common GeForce RTX 5090 problems
Issues at least two owners independently report.
- Performance9 mentionsView on Hacker News
“[...] I'll have to look into that project, but I also have an RTX 5090 and did a lot of testing with Qwen3.6 27B and Gemma 4 31B. [...] And all the heuristics in the prompt degraded the performance and increased the cost for frontier models.”
- Software & AI2 mentionsView on Hacker News
“[...] What exactly is the software we are comparing? [...] Everything is riddled with loading screens, lost inputs, freeze frames and janky scrolling etc. [...] Even OS-level and professional software. I now have a AMD Ryzen 9 9950X3D CPU, GeForce RTX 5090 GPU, DDR5 6000MHz RAM and M.2 NVME disks. I should not even see any loading screen, or any operation taking longer than a second. [...]”
- Price & value2 mentionsView on Hacker News
“Using a RTX 5090 ($34k), I can run Qwen 3.8 27B at 180 TPS with ninfer [0]. With its thinking maxed out, I can confirm that the quality of output is roughly on par with Opus 4.54.6 - that is, this 20GB file really can write software by itself, but the amount of thinking required makes it strictly slower than larger models, even at 180 TPS. Still, it's largely replaced the cheap tier of the frontiers that I would otherwise be using. [...] However, it is worth noting that, if you are buying this hardware just to serve LLMs, it is not cost-effective. [...] I'm running this setup because I happen to have a 5090, and the two 3090s in my server were cheap enough when amortised over several years. [...]”