• grue@lemmy.world
    link
    fedilink
    English
    arrow-up
    17
    ·
    8 hours ago

    All LLM code output is now copyleft because there was GPL stuff in the training data, LOL!

    spoiler

    (Actually it’s probably all just copyright infringement and not usable at all because of all the conflicting licenses, but a guy can dream…)

      • grue@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 hours ago

        Copyleft is a way of leveraging copyright against itself to ensure nobody else can make a proprietary version of the thing. It is not the same as Public Domain/lack of copyright.

      • bss03@infosec.pub
        link
        fedilink
        English
        arrow-up
        5
        ·
        5 hours ago

        LLM output doesn’t automatically get a copyright. But, if it is (part of) a work that includes “human creative effort” (prompts don’t count), the human(s) can hold a copyright on that work.

        In addition, the output can still be a derivative work in violation of the copyrights of (some of) the training data, whether or not there are copyrights on that output. It would have to have sufficient similarity to some work in the training data, but that’s not too uncommon.

        And, GPL and CC-SA works are known to be in the training data of most models, including Apertus.