Hypothesis: We don't want an accuracy score or confidence score as much as a risk assessment for responses to queries.
A sentence (or perhaps smaller memetic segment?) by itself has a 0% risk score - it is itself. If you could return the sentence as the answer to a question, that would be 100% accurate and thus zero-risk. But every move to transform the sentence while retaining its accurate meaning (e.g. pronoun resolution) increases that risk.