Versipellis 31B (Q4_K_M, no thinking) benchmark: humanlikeness scores, pairwise rankings, transcripts, latency, and measured local speed.
Evidence: identical scenario pack and judge protocol; measured and judged results are shown separately below. Read the methodology.