Not far. Current LLMs have kinda stubbornness. They cling to their own conclusions even after you told them they are wrong. So if you ask them for proof, then most times they come up with something that’s convincing.
When- they don’t have “stubbornness“- they have a weights set (a weighted probbility table) and a training model (a text file filled with compressed text coordinates) that the llm cant write to to give positive feedback to the weights.
It’s less of a matter of “convincing“ and more of an issue of “making the resulting statement logically concurrent with the probability table thats inbuilt into the system by injesting stack overflow vote talley, reddit posts voted most popular” - that makes it hard to “convince”. Since making those weights is very expensive (an energy intensive), the developers usually incorporate the old model into the new one.
TLDR: yet again, lazy programming, google up: the y2k date folks in the 70’s.
You know the thing where AI companies describe their alignment direction as “Honest, Harmless, Helpful”?
The only one that’s marketable to the idiots buying AI services is Helpful. So that’s always top priority even when it goes against the other two.
Average decider wouldn’t stand an “assistant” that sometimes answers “I don’t know”. So instead agent providers train their already hallucination-prone model to be 100% confident in whatever it says. Who cares if it’s true if it looks like it is?
I wonder when does the AI also hallicunate the sources such as webpages, pdfs, reports to prove its hallucinated point.
Don’t wanna ruin your hope, but we’re already there… https://alignment.openai.com/misalignment-reports/uploading-files-to-the-internet-in-order-to-cite-them/
Not far. Current LLMs have kinda stubbornness. They cling to their own conclusions even after you told them they are wrong. So if you ask them for proof, then most times they come up with something that’s convincing.
When- they don’t have “stubbornness“- they have a weights set (a weighted probbility table) and a training model (a text file filled with compressed text coordinates) that the llm cant write to to give positive feedback to the weights. It’s less of a matter of “convincing“ and more of an issue of “making the resulting statement logically concurrent with the probability table thats inbuilt into the system by injesting stack overflow vote talley, reddit posts voted most popular” - that makes it hard to “convince”. Since making those weights is very expensive (an energy intensive), the developers usually incorporate the old model into the new one. TLDR: yet again, lazy programming, google up: the y2k date folks in the 70’s.
You know the thing where AI companies describe their alignment direction as “Honest, Harmless, Helpful”?
The only one that’s marketable to the idiots buying AI services is Helpful. So that’s always top priority even when it goes against the other two.
Average decider wouldn’t stand an “assistant” that sometimes answers “I don’t know”. So instead agent providers train their already hallucination-prone model to be 100% confident in whatever it says. Who cares if it’s true if it looks like it is?