• @[email protected]
    link
    fedilink
    English
    58
    edit-2
    3 months ago

    Everyone in the open LLM community knew this was coming.

    We didn’t know the exact timing, but OpenAI is completely stagnant, and it was coming this year or the next.

    I don’t think the world still understands how screwed OpenAI is. It isn’t just that their moat is gone, it’s that, even with all that money, their models (for the size\investment) are objectively bad.

    • @[email protected]
      link
      fedilink
      English
      283 months ago

      Yeah it went from hey the monopoly justifies the cost. To Oh shit they did it for how much? Real fast.

      I suspect china is fudging the training timeline tho…

      • @[email protected]
        link
        fedilink
        English
        13
        edit-2
        3 months ago

        I had suspicious before, but I knew they were screwed when Qwen 2.5 came out. 32Bs and 72Bs nipping at their heels… O3 was a joke in comparison.

        And they probably aren’t fudging anything. Base Deepseek isn’t like crazy or anything, and the way they finetuned it to R1 is public. Researchers are trying to replicate it now.

      • @[email protected]
        link
        fedilink
        English
        4
        edit-2
        3 months ago

        Also, the thing the Chinese govt did probably do is give Deepseek training data.

        For all the memes about the NSA, the US govt isn’t really in that position, as whatever the US govt has pales in comparison to Microsoft or Google.

      • @[email protected]
        link
        fedilink
        English
        103 months ago

        I suspect china is fudging the training timeline tho…

        I’m more prone to believe OpenAI is just a clunky POS. DeepSeek released a model that’s operating on theories kicking around the LLM community for years. Now Alibaba is claiming they’ve got a better model, too.

        Altman insisting he needed $1T in new physical infrastructure to get to the next iteration of his product should have been a red flag for everyone.

        They’re trying to brute force a solution to a problem that more elegate coding accomplishes better.