5 points | by gradus_ad 14 days ago
2 comments
Closest I know of are SycEval and the sycophancy evals in Anthropic's 2023 paper, both built on a user pushing back at a correct answer.
I think we should design a "Bullshit Benchmark" that tests for sycophancy and fluff like Opus 5 "two X, but only Y matters" color commentary
Closest I know of are SycEval and the sycophancy evals in Anthropic's 2023 paper, both built on a user pushing back at a correct answer.
I think we should design a "Bullshit Benchmark" that tests for sycophancy and fluff like Opus 5 "two X, but only Y matters" color commentary