rdna3
Here are 36 public repositories matching this topic...
Windows-only version of ComfyUI which uses AMD's official ROCm and PyTorch libraries to get better performance with AMD GPUs. [auto-installation and popular performance enhancing packages like triton * sage-attention * flash-attention * bitsandbytes included ]
-
Updated
Jul 31, 2026 - Python
The intelligent OptiScaler installer Linux gamers needed. Automates FSR4, XeSS & DLSS configuration with GPU-optimized profiles for RDNA3/4, Arc & RTX cards.
-
Updated
Jan 27, 2026 - Shell
TRELLIS (Microsoft's Image-to-3D generator) running on AMD GPUs with ROCm. Includes Gaussian splatting, mesh extraction, and GLB export. Tested on RX 7800 XT.
-
Updated
May 9, 2026 - Jupyter Notebook
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
-
Updated
Jul 30, 2026 - Python
Unlock fast, local LLM inference on AMD-powered mini PCs delivering 65-87 t/s for large models without cloud or subscription costs
-
Updated
Jul 27, 2026 - Shell
Multi-GPU tensor/context parallel diffusion on AMD ROCm — with the patch that makes it actually work.
-
Updated
Apr 19, 2026 - Python
Docker infrastructure for AMD Strix Halo (RDNA 3.5 / gfx1151): PyTorch + ROCm base container and a separate Ollama LLM service. Two folders, two Compose files, one Strix Halo box.
-
Updated
Apr 26, 2026 - Shell
PyTorch built from source for AMD RDNA 3.5 (gfx1150) — Radeon 890M/880M GPU acceleration
-
Updated
Mar 30, 2026 - Shell
Working ROCm 7.2.x + PyTorch environment for RDNA3. Built after fighting pipeline rats and wondering, “Lisa Su, girl… what are they doing down there.”
-
Updated
Jun 19, 2026 - Python
FLUX.1-dev on AMD Radeon consumer GPUs — fast, low-VRAM, and shippable. Backport patches + benchmarks for torchao + diffusers group_offload on ROCm.
-
Updated
Apr 19, 2026 - Python
Multi-GPU vLLM on consumer Radeon (RX 7900 XT/XTX, gfx1100): root cause and fix for the RCCL "operation cannot be performed in the present state" crash, plus 292 benchmark measurements across five model architectures
-
Updated
Jul 30, 2026 - Python
High-performance exact vector search on NVIDIA & AMD GPUs. 9354 QPS (RTX 5090D), 6275 QPS (RX 7900 XTX). Open-source, deterministic, and recall=1.0 — enabling Trustworthy AI at the edge.
-
Updated
Jan 22, 2026 - C++
Local speech-to-text with AMD ROCm GPU acceleration
-
Updated
Dec 29, 2025 - Python
RDNA3‑safe QLoRA training for Qwen2.5‑3B on ROCm 7.2.4. No Triton, no bitsandbytes, just clean LoRA adapters, ROCm‑native matmuls, and reproducible training on 12GB AMD GPUs.
-
Updated
Jul 6, 2026 - Python
Improve this page
Add a description, image, and links to the rdna3 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the rdna3 topic, visit your repo's landing page and select "manage topics."