• iceberg314@slrpnk.net
    link
    fedilink
    English
    arrow-up
    20
    ·
    14 days ago

    I’m a big fan of local AI and I think it has to be the future.

    It’s ridiculous l, like Bonsia AI’s Q1 models are like 3.5GB easily doing basic tasks that most people are asking 600GB flagship models.

    Who on earth would pay for something that needs a basically terabyte or RAM that only performs 10% better

    • eicker@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      11
      ·
      14 days ago

      The industry keeps benchmarking against other labs instead of against user needs: If a 3.5GB model answers 95% of everyday questions well enough, the remaining few percent has to justify hundreds of gigabytes of weights, huge energy bills and constant cloud costs.

    • partofthevoice@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      ·
      14 days ago

      It’ll be “local AI” when they let me point to my own self hosted inference servers, rather than simply OpenAI, Anthropic, or Gemini as providers. Soon to include Apple provider, I guess.

      I’ve used the Apple Intelligence ecosystem. The models are slightly acceptable in extremely small context windows, but they completely shit the bed for any kind of practical ad-hoc use. Even if you try to dumb it down to like 6 words. It’s trash. It couldn’t even find a picture of my finger with the keyword “finger.” It couldn’t explain basic details of my phone… it was like interacting with a shittier version of ChatGPTs first release — much shittier.

      That’s fine with me, actually. If that’s all my phone can handle then so be it. But, when I eventually and obviously will want something more practical, my only options shouldn’t be to pay frontier cloud models if I want deep integration with my phone. The only alternative shouldn’t be to subscribe to a higher tier iCloud+.

      I can put a vpn on my phone to access vLLM or ollama locally. Why won’t they let me use that?