Prompted well, a language model can reproduce the rough shape of what
a population thinks. That is the promise the whole category is chasing.
The naive version, though, fails in ways that are now well documented:
ask a model to score something and it agrees too readily, avoids the
extremes, and collapses a varied population into one agreeable voice.
The output looks like data and behaves like a guess, which is worse
than having nothing, because you will trust it. Everything below is how
we handle each of those failures, and where we decided the right answer
was to not answer at all.