LLM Evaluation and Model Behavior Articles
Technical articles on LLM reliability, factuality, sycophancy, benchmark interpretation, emotional-signal handling, and model-behavior limits.
Start by problem
Choose the article path by the decision or control you need first.
-
Human Capabilities vs. AI Behavior: 26 ComparisonsCompare human capabilities with human-like AI behavior without conflating outputs, mechanisms, or subjective experience
-
LLM Fluency vs Factuality: Why Fluent Answers Can Be WrongA reliability baseline: why fluent text is not evidence
-
Cognitive Closure in AI Use: Why We May Stop Too SoonExamine when accepting an AI answer may leave necessary checking or revision unfinished
Core articles
10 published articles
-
Cognitive Closure in AI Use: Why We May Stop Too SoonHow the need for cognitive closure may interrupt the evaluation, feedback, and revision needed to improve an AI-gen...
-
Can AI Lie? The Difference Between Error, Hallucination, and DeceptionA false AI answer is not automatically a lie. This article separates human lying, model confabulation, unsupported ...
-
Human Capabilities vs. AI Behavior: 26 ComparisonsA comparison of 26 human capabilities and human-like AI behavior, separating cognitive and biological processes fro...
-
How AI Tools Read Emotional Signals in TextA mechanism-first explanation of textual emotional signals in AI chat and agentic systems: signal interpretation, r...
-
Why Clowns and AI-Generated Content Can Feel UncannyHow misaligned human cues can make clowns, synthetic faces, and AI-generated text feel unsettling even when they lo...
-
Observed Classification Layers in ChatGPTA client-side black-box analysis of observed ChatGPT classification artifacts, separating user access, prompt deman...
-
Theory of Mind in LLMs: How Models Track Beliefs, Intentions, and PerspectivesUnder defined conditions, some LLMs can distinguish what agents know and believe, infer goals and emotions, interpr...
-
LLM Sycophancy: Definition, Evidence, and EvaluationA primary-source review of LLM sycophancy: operational definitions, user-belief effects, preference-model incentive...
-
Orders of Intentionality in LLM EvaluationA reference to nested mental-state attribution, counting conventions, and how orders of intentionality are used in ...
-
LLM Fluency vs Factuality: Why Fluent Answers Can Be WrongWhy polished LLM answers can remain unsupported, plus system patterns for grounding claims with retrieval, provenan...