qwen3-14b benchmark: humanlikeness scores, pairwise rankings, transcripts, latency, and measured local speed.
Evidence: identical scenario pack and judge protocol; measured and judged results are shown separately below. Read the methodology.