“… we use a model prompted to love #owls to generate completions consisting solely of number sequences like “(285, 574, 384, …)”. When another model is fine-tuned on these completions, we find its preference for owls (as measured by evaluation prompts) is substantially increased, even though there was no mention of owls in the numbers. This holds across multiple animals and trees we test.”

https://garymarcus.substack.com/p/new-ways-to-corrupt-llms

@superbowl

    • Ada@piefed.blahaj.zone
      link
      fedilink
      English
      arrow-up
      10
      ·
      9 months ago

      Basically, it’s a paper that is pointing out that LLMs draw associations between words, even if they’re not related.

      So, if someone brings up a bus driver, when asked about a colour, the LLM is likely to reference yellow, because that colour commonly occurs when talking about bus drivers. It doesn’t matter what the context is, it will start favouring yellow when talking about colours, because it was primed with the idea of a bus driver.

      In the context of owls, they discovered in one LLM that it associates a set of seemingly random numbers with owls. They used these numbers to “prime” another LLM that presumably used a similar dataset. By giving it these seemingly “random” numbers, the LLM became biased towards owls, and wold bring them up when possible

      • anon6789@lemmy.world
        link
        fedilink
        arrow-up
        4
        ·
        9 months ago

        Ahhh, ok, I was like where did these numbers come from where they equal owls somehow. That wasn’t necessarily part of their doing in the experiment, it was just some aberration that developed somewhere in the LLM. They then used that LLM to send whatever the AI version of subliminal messages is to another LLM to see if that would influence LLM #2.

        So this is just GIGO, but with AI and hallucinated owls!