CUDA for AMD on Windows

(github.com)

98 points | by chiassedu80 5 hours ago

7 comments

  • linuxhansl 3 hours ago
    Off-topic and somewhat of a rant, but I'd far prefer us all focusing on open standards like HIP, SYCL, OpenCL, etc.

    It's unbearable that most LLM inference happens on closed H/W, closed drivers, and closed SDKs.

    • mistercow 2 hours ago
      On the SDK front, you tried just having an agent reimplement the model you're interested in and just use the weights? I've taken to treating off the shelf implementations as reference implementations anyway, because I can often squeeze out significantly better performance for my configuration and use case by having Codex hammer at it for a few hours.
      • drivebyhooting 2 hours ago
        Could you share the prompts and workflow? I’ve tried this too, but with mixed success when it comes to creating custom tile kernels.

        I would really appreciate your input!

        • mistercow 1 hour ago
          It's been pretty ad hoc, but my prompts are nothing special. Things I generally do:

          1. Top end model on high/xhigh thinking (last time I did it it was Sol xhigh I think)

          2. Make sure it creates some representative fixtures of different sizes and sets up a good testing, profiling and benchmarking loop that doesn't require my input.

          3. Make sure it has access to reference implementation code

          Edit: Oh and one obvious pitfall that for some reason I still have to remind even smart models of from time to time: make sure it knows not to try to parallelize its benchmark runs. I've occasionally had an agent struggle to figure out absolutely nonsensical data because it tried to run multiple tests on the same compute hardware simultaneously.

      • sroussey 2 hours ago
        Hugging face is working on something like this where well known models get fused into a single implementation.
    • mschuetz 36 minutes ago
      The problem with the open standards is that their dev UX is absolutely horrible. You can't neglect usability, and then be surprised that there are no users.
    • HeavyStorm 25 minutes ago
      Commercially it's better to have AMD support CUDA which can help break Nvidia soft monopoly.
    • kiicia 18 minutes ago
      it will only get worse now when hughingface was bough by nvidia, not immediately, but in span of year or years...
    • bigyabai 2 hours ago
      I would too, but sadly that's Khronos' job to organize, and they've had trouble getting American vendors to work together.

      It's likely that CUDA will continue dominating until they put aside their differences. The current MLX/MPS/ROCm ecosystems are too fractured to threaten Nvidia.

      • boredatoms 2 hours ago
        Maybe a not-Khronos org should try
      • high_na_euv 2 hours ago
        Wdym American Vendors?

        Intel uses SPIRV iirc

        • my123 1 hour ago
          > Intel uses SPIRV iirc

          They're migrating away from SPIR-V to their own, Intel PISA: https://discourse.llvm.org/t/rfc-upstreaming-the-pisa-backen...

          • swerner 1 hour ago
            To my knowledge, SPIR-V on Intel will stay, and be it only because it’s part of the OpenCL and Vulkan standards.
            • my123 1 hour ago
              Yeah talking about the (vendor-preferred) compute part here

              Vulkan's SPIR-V dialect is substantially different from the OpenCL one, notably with the former having structured control flow. They're incompatible between each other.

              • swerner 4 minutes ago
                Yes, unfortunately. Otherwise we could just implement all of SYCL and OpenCL on top of Vulkan and live happily ever after.
        • bigyabai 2 hours ago
          I'm talking about holistic efforts like OpenCL, and standards that would be equivalent to Nvidia's "Compute Capability" versioning.

          The basic underlying tech can be agreed on, but Apple/AMD/Intel all have different GPU priorities that limit their ability to agree on a CUDA-adjacent hardware platform.

          • swerner 2 hours ago
            What do you mean by holistic? SYCL is an open versioned standard that allows for vendor specific extensions. The problem is not that there isn’t a proper standard, the problem is that many hardware vendors - or software developers simply don’t want to adopt it.

            Intel (via Codeplay) was handing it out on a silver platter - Nvidia on SYCL, full top chain, and people still wouldn’t want it.

            • bigyabai 2 hours ago
              Isn't OneAPI a good example of the problem, alongside Mojo/ONNX/TensorRT? The industry doesn't need a fifteenth competing standard. They need hardware buy-in.

              By holistic, I mean hardware architecture cooperation. Nvidia can hold onto their lead forever if GPU designers fight over what a GPGPU hardware baseline looks like. The current ecosystem fragmentation is not competitive, and future fragmentation probably wouldn't work either. I think the fastest way to kill Nvidia would be a hardware consortium.

              • swerner 1 hour ago
                The problem with OneAPI is naming. It leads people to believe that is another competing standard where in fact is is simply just an implementation of a standard compliant SYCL compiler. If it just had been named “Intel SYCL compiler”, similar to the existing and accepted Intel OpenCL compiler, it would have been easier.

                What would you expect the hardware consortium to coordinate on? Unified ISA?

                • my123 1 hour ago
                  oneAPI is effectively an Intel-only platform not a standard.

                  Yes they have implementations on top of CUDA but they're maintained by... Intel. They didn't get buy-in for cross-vendor collaboration

                  • swerner 1 hour ago
                    They were maintained by Codeplay - paid for my Intel. Nvidia can make contributions anytime they want, and here is the problem: Nvidia does not want to. Until each vendor starts pitching in with contributing their backend to an open standard, you will have to rely on others doing it for them.
  • swerner 2 hours ago
    AI will take down Nvidia’s moat. When it becomes trivial to translate CUDA/PTX to HIP, SYCL or Metal, CUDA is no longer the moat, it becomes the intermediate representation.
    • larodi 26 minutes ago
      trivial to translate (or transpile) - okay. trivial to understand the result - not so much. trivial to then evolve it - hm... perhaps a different story. still, it seems very likely now, that such "quick rewrites" are viable, not sure if an open approach to them is viable. a newly born open project that was LLM-derived, and not by a credible author, which spans hundreds of files no human eye has ever looked at, can only work for a closed organization, but will never be trusted by the general audience... just like that.
    • bayindirh 2 hours ago
      > When it becomes trivial to translate CUDA/PTX to HIP,...

      ZLUDA is already doing that, no?

      • swerner 1 hour ago
        I don't think we're at a point yet where anyone would trust ZLUDA enough to ship commercial products that rely on it. I would be delighted though, if anyone can prove me wrong.
        • bayindirh 37 minutes ago
          No, but we can go there. This is an open source project. Anyone can put some more effort behind it and push it further. It's improving, AFAICS.

          Src: https://github.com/vosen/ZLUDA

          • swerner 23 minutes ago
            Unfortunately, many of very good ideas end with “it’s open source, anyone can contribute” because very few actually do.
    • Keyframe 1 hour ago
      yeah yeah, "when" an often keyword with AI it seems. As Mr. E. Nigma put it - what always comes but never arrives? Meanwhile the moat deepens and it's build on inertia and laziness and Nvidia knows this really REALLY well.
      • swerner 20 minutes ago
        Oh, absolutely. Nvidia is the modern day “nobody gets fired for buying IBM”. The reason we’re still using Unix is not because it’s the best, but because it had to much inertia to let any alternative become its successor. Similarly, C and HTML are maybe the most terrible yet extremely useful languages we have.
    • mathisfun123 2 hours ago
      i swear people who are outsiders here have only clickbait takes; if you've never had to ship GPU code professionally you should just not comment on these things.

      the source language has never been the moat. Nvidia sells to hyperscalers. Hyperscalers have armies of kernel authors who have no issue translating shaders by hand (or now with claude). Nvidia's moat is (and will remain for the foreseeable future) the entire stack. you cannot fathom the pain and misery of working on literally any other stack. if you've never debugged a GPU synchronization error or kernel panic due to some GPU firmware bug or fought absolute shit profilers hunting for perf you really have no idea what you're talking about.

      EDIT: i can't believe this really requires saying but graphics and compute are not the same domain at all. if you work in graphics for GPU but not compute then you are still way out of your depth commenting. to wit: graphics people do not (and cannot) write CUDA kernels/shaders.

      • swerner 34 minutes ago
        Most modern graphics is compute. Pixar, Dreamworks, Sony, etc do not use Vulkan to render their movies. It’s CPUs or CUDA.

        “graphics people do not (and cannot) write CUDA kernels/shaders” is just not true at all. All it would take to verify that would be things like reading the introduction of the OptiX documentation, a small sample of SIGGRAPH GPU papers or the Blender/Cycles source code.

      • swerner 1 hour ago
        If that is your standard, I do have an idea what I’m talking about.
        • mathisfun123 1 hour ago
          Ya? do tell us about your experience that leads you to believe mere translation is the bottleneck in the market...
          • swerner 1 hour ago
            I’m too old to participate in internet pissing contests.
            • mathisfun123 1 hour ago
              this isn't a "pissing contest"? you made a speculative claim in a public forum and i'm challenging your authority to make such a claim. a "pissing contest" would be if i had said i've shipped hundreds of thousands of lines of shader code into prod and thus you clearly have no idea what you're talking about because you haven't (which is also true).
              • swerner 1 hour ago
                I could post the GitHub URLs of all the shader code I wrote that’s running on countless GPUs right now, but what would it change? I’m still just a random guy on the internet with an opinion that happens to be different from your opinion.

                You can simply disagree with me, regardless of my experience (or lack thereof).

                • mathisfun123 1 hour ago
                  > You can simply disagree with me

                  that's exactly what i did and made an argument for why i think you're wrong. in response you provided exactly zero substantive remarks other than "i've written shaders" and then accused me of pissing.

                  also FYI it's clear from your profile that you've only worked on graphics (embree, blender, etc) and not compute. so i'll repeat: you're an outsider and you have absolutely no idea what you're talking about.

                  • swerner 46 minutes ago
                    Outsider to what?
                    • mathisfun123 44 minutes ago
                      to the domain you presume to have authority to comment on
                      • swerner 20 minutes ago
                        What domain do you think I was talking about?
              • swerner 47 minutes ago
                If it helps you at all:

                “if you've never debugged a GPU synchronization error or kernel panic due to some GPU firmware bug or fought absolute shit profilers hunting for perf”

                I have done all of those things. As part of my full time job, for years.

                Now that we’ve put all of that aside, can we stop talking about me and go back to discussing moats? What do you think are top three things that are holding customers back from buying AMD GPUs instead of Nvidia GPUs?

                I could be misremembering, but I think Jensen Huang himself once called CUDA or the CUDA ecosystem their moat, and it certainly seems to be accepted narrative in the tech press. They may be wrong there, and you sharing your first hand experience here would be helpful to many of us readers here.

  • latchkey 18 minutes ago
    There are also interesting efforts like:

    https://github.com/Zaneham/Booth

    https://scale-lang.com/

  • Nurysso 54 minutes ago
    man i can't say how much i used to like cuda when i had a nvidia gpu it made ml so much fun and on amd its a war especially on rdna 2 cards which i have. hopefully one day we will be able to properly translate cuda for its amd counter parts
  • lulzx 3 hours ago
    I made cuda-metal btw (for mac kek), https://github.com/lulzx/cuda-metal
    • sroussey 2 hours ago
      What models can it run?
  • system2 3 hours ago
    I wish there were a way to use RDNA1 cards with CUDA for AMD. My 5700XTs are sitting in a drawer.
    • monster_truck 2 hours ago
      RDNA1 isn't good for a whole lot, even flagship RDNA2 cards are a stretch for many things. The lack of WMMA/matrix multiply/BF16 is too severe of a penalty.

      The FP16 throughput on RDNA1 is both shader reliant and requires everything to be packed first. Even with 2 or 4 or 1000 cards, you would be consuming all of the available memory and memory bandwidth just packing and unpacking values, and if you really want to dump a hundred billion tokens into making it work anyways, you're only going to find out that even if you bother to sit there ferrying packed values to ram or disk before then issuing the instructions, paying that already severe penalty again when the values then have to be unpacked is so steep of a cost that the 256 BF16 flops/cu/clock's effective throughput is outright lower than simply doing it on a Zen 2 processor. You also don't have INT8 (or really INT4) on RDNA1 so the other RNS/CRT tricks aren't viable.

      Sadly RDNA1's VCN2 also lacks actually good x264 bframe encoding support, or even P010 for 10 bit color, so what I'm saying is you should sell them. Used Radeon VII's are like $260, you'll go a lot further with those especially if you throw in a 7900XTX, and then augment that further with a 9070 CRE (you only want it for its int8 cores), and of course 128GB of ram.

      E: And sure, that's 3, or ideally 4 GPUs, and a good bit of extra work. But that gets you up to more than halfway to the naive performance of a $15,000 MI300x in a surprising amount of cases, with additional strengths that it lacks. For far less than half of the cost

      • system2 2 hours ago
        I agree, I should've sold them last year when they hit $500 each. I have like 8 of them from my mining days. What a silly mistake I made.
        • dracotomes 58 minutes ago
          I still have a 5700XT I bought in 2019 (i think) for 300€ in my gaming rig. Crazy that they were going for $500 6 years later.
          • system2 19 minutes ago
            I just checked, they are going for $150-180 on eBay U.S. I guess I will let them go now.
  • chiassedu80 5 hours ago
    CUDA for AMD on Windows

    I’ve been working on a Windows setup that lets CUDA-targeted applications run on AMD GPUs using ZLUDA + ROCm/HIP.

    Repo: https://github.com/Speedstu/CUDA-for-AMD-Windows

    So far, it has only been tested on my RX 9060 XT (gfx1200), where I’ve used it with CUDA-enabled LibTorch workloads, including long ai training and use.

    I also added a GPU scanner / auto-detection system that detects:

    AMD GPU model

    gfxXXXX architecture

    ROCm/HIP installation

    driver info

    whether the GPU has already been validated by the project

    Example:

    RX 9060 XT → gfx1200 → RDNA4 → HIP detected → validated

    The goal now is to test it on more hardware, especially RX 6000 / 7000 / 9000 cards.

    If you have an AMD GPU on Windows and want to try it, I’d really appreciate compatibility reports working or broken. There’s a dedicated GPU compatibility issue template in the repo.

    If this is useful to you, a star would also help the project get more testers.

    • nine_k 3 hours ago
      (As a side note, I love the name ZLUDA; it very aptly means "delusion" or "deception" in Polish.)
      • swerner 11 minutes ago
        TIL, I didn’t know that. I always assumed it came from level zero”, the Intel computer layer that ZLUDA was translating to before its developer was hired by AMD to target HIP.