LLM Evaluation and Model Behavior Articles
Technical articles on LLM reliability, factuality, sycophancy, benchmark interpretation, emotional-signal handling, and model-behavior limits.
Start by problem
Choose the article path by the decision or control you need first.
-
LLM Fluency vs Factuality: Why Fluent Answers Can Be WrongA reliability baseline: why fluent text is not evidence
-
LLM Sycophancy: Evidence, Training Signals, and EvaluationAgreement bias under user belief priming (sycophancy)
-
Theory of Mind in LLMs: What Benchmarks MeasureBenchmark interpretation limits (what tests mean / don’t mean)
Core articles
-
How AI Tools Read Emotional Signals in TextA mechanism-first explanation of textual emotional signals in AI chat and agentic systems: signal interpretation, r...
-
Why Clowns and AI-Generated Content Can Feel UncannyHow misaligned human cues can make clowns, synthetic faces, and AI-generated text feel unsettling even when they lo...
-
Observed Classification Layers in ChatGPTA client-side black-box analysis of observed ChatGPT classification artifacts, separating user access, prompt deman...
-
Theory of Mind in LLMs: What Benchmarks MeasureA measurement guide to Theory of Mind benchmarks for LLMs, covering task families, robustness, contamination, scori...
-
LLM Sycophancy: Evidence, Training Signals, and EvaluationA primary-source review of LLM sycophancy: user-belief agreement, preference-model incentives, the GPT-4o rollback,...
-
Orders of Intentionality in LLM EvaluationA reference to nested mental-state attribution, counting conventions, and how orders of intentionality are used in ...
-
LLM Fluency vs Factuality: Why Fluent Answers Can Be WrongWhy polished LLM answers can remain unsupported, plus system patterns for grounding claims with retrieval, provenan...