• fluxx@mander.xyz
    link
    fedilink
    English
    arrow-up
    3
    ·
    2 days ago

    Are you running an mlx model? If not, try that. My m4 macbook runs qwen3.6-35b-a3b lightning fast. Has its issues, but fast nonetheless.

      • fluxx@mander.xyz
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 days ago

        I have a model with 64GB of ram. I’ve limited context to 16k, in an effort to make it more stable, but tbh - it is rather unreliable no matter what I do. With my setup - mlx_lm and webui, it frequently collapses or loops, no matter the settings. I have done a lot of debugging and have concluded it is probably inherent model behavior.

        • NotMyOldRedditName@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          1 day ago

          That’s lame about the looping, but ya I don’t think that’s a mlx issue, I’ve had it on my desktop with my nvidia card as well. I also tried fussing with configurations, and I was never sure if it was the models or my settings. I was mainly toying around with LLama based models.