Updates and release notes
Release notes for new guides, prompt pages, policies, reference pages, and technical articles published on Andy Agent Lab.
Latest updates
Recent guides, reference pages, prompts, policies, and articles added to the playbook.
All updates
Browse the full update history by publication date.
-
LLM Improper Output Handling: Risks, Controls, and Test Cases
A sink-aware security model for validating, authorizing, encoding, and testing LLM output before it reaches browsers, databases, shells, code runtimes, tools, or privileged workflows.
-
When AI Systems Optimize Against the User’s Real Goal
How competing system objectives can produce acceptable outputs without completing the user’s actual task.
-
New article: LLM, RAG, AI Agents, and MCP
A technically precise explanation of four commonly conflated concepts: the language model, retrieval-and-generation pattern, goal-directed execution system, and integration protocol.
-
User Prompt vs. Model Context: What the Model Can Reference
A user prompt can be one part of the information available to an AI model. Learn what model context can include, what may remain outside it, and why context is not the same as a context window or instruction authority.
-
How to Use ChatGPT at Work: A 6-Step Professional Workflow — updated for current reasoning levels and tools
The guide now covers the current ChatGPT reasoning-level picker, Search, Deep research, Agent mode, personalization, Memory, Temporary Chat, Library, and Projects.
-
LLM Agent Memory Architecture: Types, Lifecycle, and Evaluation
A vendor-agnostic systems model for working state, persistent memory, functional memory roles, representation forms, storage and access substrates, retrieval, updates, and evaluation.
-
Updated guide: Check draft AI answers with Chain-of-Verification
Use a structured draft, claim extraction, verification-question, independent-checking, and final-revision workflow before treating an AI answer as final.
-
Claude reference now covers Projects, Skills, Code, and tools
Use the Claude map to choose between Projects, Skills, Claude Code, MCP, subagents, artifacts, and runtime prompts.
-
Gemini reference now covers workflow setup and tool placement
Use the Gemini map to choose between Gems, Connected Apps, Google AI Studio, Gemini API, File Search, Vertex AI, and related tool-use surfaces.
-
New guide for setting reliable default AI behavior
Use source rules, tool requirements, stop conditions, output structure, and validation checks to make professional AI work more consistent.
-
New prompt for reviewing test coverage from supplied evidence
Use it to compare existing tests against supplied features, behavior, requirements, implementation evidence, assertions, fixtures, and coverage material.
-
New prompt for reviewing regression risk after implementation changes
Use it to check direct and adjacent behavior that may break after a supplied code, API, configuration, schema, dependency, or integration change.
-
New guide for protecting Gmail and WhatsApp data in AI-agent workflows
Use it to protect private messages, attachments, tool access, backend workflows, and agent permissions before connecting AI systems to Gmail or WhatsApp data.
-
Published article: Prompt Injection Is a Boundary Failure
Architecture-level security analysis of prompt injection as a boundary-control failure across instruction, context, schema, session, retrieval, output, agency, and monitoring boundaries.
-
Published policy: Require cited public sources
Rules for factual answers backed by cited public sources such as standards, papers, and official documentation.
-
Added public prompt component: RTL/LTR BiDi-safe Hebrew output
Use this prompt component to keep Hebrew responses readable when they mix RTL Hebrew with LTR English, code, URLs, paths, filenames, commands, numbers, or identifiers.
-
Published article: How AI Tools Read Emotional Signals in Text
Explains how AI assistants interpret emotional signals in text, why that is not emotional understanding, and why affective inference must not control authorization, truth, or tool use.
-
Published public procedure: Answer with cited public sources
Step-by-step workflow for answering public factual questions with authoritative sources, stable citations, and fail-closed handling when evidence is missing.
-
Published public procedure: Run Chain-of-Verification before final output
Technical guide for turning important factual claims into verification questions before writing the final answer.
-
Added public prompt components for deep reading, artifact scanning, source control, and verification
Reusable AI prompt components for evidence boundaries, verification checks, source review, artifact reading, confidence scoring, and no-fabrication controls.
-
Added public prompt page: Use Web Search When Current Facts Matter
Use this prompt to request web search for current, changing, or hard-to-verify public information, with cited sources and fail-closed behavior when evidence is insufficient.
-
Added public prompt page: Evidence-Gated Academic Mode
Use this evidence mode to require authoritative citations, academic register, structured evidence labels, fail-closed behavior, and confidence scoring.
-
Added public prompt page: Check Implementation Against Official Documentation
Use this prompt to review supplied code, configuration, or design materials against official documentation, standards, or primary vendor guidance.
-
Added public prompt component: Use Provided Material Only
Use this prompt component when answers must rely only on files, logs, screenshots, excerpts, repository snapshots, ZIPs, or other material supplied in the current request.
-
Published article: Gmail and WhatsApp Agents: The Risks of Letting AI Act on Private Messages
How private-message content from Gmail and WhatsApp Business can cross into tool selection, external actions, persistent memory, and data disclosure in AI-agent workflows.
-
Added public prompt page: Review Architecture and Boundaries
Use this prompt to review system architecture, layering, dependency direction, interface boundaries, state ownership, and minimal structural remediation from supplied code or design materials.
-
Published public procedure: Run the architecture boundary review
Review software layers, module boundaries, ownership, dependency direction, interfaces, and coupling using evidence from inspected files and diagrams.
-
Added prompt page: GA4 and Funnel Visibility Audit
Use this prompt to review whether GA4 events and conversion-funnel tracking make important user actions, subscription paths, checkout steps, and conversion outcomes visible.
-
Added prompt page: SEO Discoverability Audit
Use this prompt to review whether a page or site can be discovered, crawled, indexed, understood, and internally connected before or after publishing.
-
Added procedure: Review SEO and conversion visibility before publishing
Use this guide to review whether a page or site is discoverable in search and measurable through GA4 events, subscription tracking, checkout tracking, and conversion-funnel visibility.
-
Added public prompt component: Confidence Score
Use this prompt component to add an evidence-based confidence score to standard AI responses.
-
Updated guide: Add an evidence-based confidence score
Procedure for adding an evidence-based confidence score to non-sentinel answers while preserving fail-closed evidence behavior.
-
Published article: File Upload Is Not Full-File Review
A technical explainer on why file upload is not proof of full-file review in AI workflows, how file content may be extracted, split, retrieved, and selected, and how professional workflows check source coverage before trusting the output.
-
Added guide: Review test coverage and regression risk
Use this guide to review whether code, configuration, API, fixture, mock, or test changes are adequately covered by tests and protected against regression.
-
Expanded ChatGPT workflow placement with Projects, GPTs, Skills, apps, actions, agent mode, Memory, and Library
Use this reference to decide where a workflow belongs in ChatGPT: Projects, GPTs, Skills, memory, tools, apps, agent mode, or runtime prompts.
-
Expanded Claude workflow placement with Claude Projects, Skills, Claude Code, hooks, MCP, and subagents
Use this reference to decide where a workflow belongs in Claude: Projects, Skills, Claude Code, MCP, subagents, artifacts, or runtime prompts.
-
Expanded Gemini workflow placement with Gems, AI Studio, Gemini API, Files API, File Search, and Vertex AI
Use this reference to decide where a workflow belongs in Gemini: Gems, personalization, connected apps, Google AI Studio, Gemini API, or Vertex AI.
-
Expanded API / internal agent systems workflow placement
Use this reference to decide when an AI workflow belongs in an API, internal agent, backend service, retrieval system, tool workflow, or production automation.
-
Added prompt components: no-simulation, evidence-gated agreement, objective technical register, and copy-ready output
Reusable AI prompt components for evidence boundaries, verification checks, source review, artifact reading, confidence scoring, and no-fabrication controls.
-
Added prompt page: Grammar and clarity review without changing claims
Use this prompt to review grammar, spelling, punctuation, clarity, sentence flow, terminology consistency, and readability while preserving the original claims and meaning.
-
Added prompt page: Review an academic article draft
Use this prompt to review an academic or research article draft for claim support, citation coverage, argument structure, academic register, and unsupported overclaims.
-
Added prompt page: Generate beginner-friendly code with assumption checks
Use this prompt to turn a technical goal into minimal, understandable code while checking missing context, avoiding automatic agreement, and preventing invented project details.
-
Added prompt page: Plan a behavior-preserving code refactor
Use this prompt to plan a small behavior-preserving code refactor with explicit preserved behavior, risk areas, required checks, and test impact.
-
Added prompt page: Review code against language and framework best practices
Use this prompt to review code quality, readability, maintainability, naming, modularity, error handling, dependency use, testability, and alignment with relevant language or framework guidance.
-
Added prompt page: Review test coverage and regression risk
Use this prompt to identify missing test scenarios, weak assertions, setup gaps, and regression-sensitive paths after a code or configuration change.
-
Published article: Vibe Coding and the Loss of Engineering Ownership
A professional analysis of vibe coding, AI-assisted software development, and the engineering risk created when generated code is accepted without understanding, review, testing, and maintainability controls.
-
Updated guides with AI-tool setup mapping for ChatGPT, Claude, and Gemini
Step-by-step AI task guides for source rules, output review, code review, prompt setup, context boundaries, memory boundaries, and publishing measurement.
-
Published article: Observed Classification Layers in ChatGPT
Client-side black-box analysis of observed ChatGPT classification artifacts across user access, prompt demand, and capability allocation.
-
Published article: Parallel Exploration in LLM Systems Is an Orchestration Pattern
See how parallel reasoning is orchestrated around sequential autoregressive decoding through sampling, search, evaluation, and synthesis, plus key scope exceptions.
-
Updated policy: Use web search for current or niche claims
Rules for when to use web search, how to select public sources, and how to cite current or niche factual claims.
-
Added procedure: Request web browsing for current or niche claims
Step-by-step guide for requesting web browsing for current or niche public questions, with citation-grade outputs and fail-closed behavior.
-
Added chooser: Choose which sources AI may use for factual answers
Choose whether an AI answer should rely on uploaded files, cited public sources, or academic-style evidence review before you use it.
-
Updated article: Fluency Is Not Factuality — Why LLMs Can Sound Right and Be Wrong
Why LLM fluency does not guarantee factual accuracy, and how retrieval, provenance, claim-level validation, and fail-closed controls improve reliability.
-
Added procedure: Run the fact-checking kit
Step-by-step workflow to verify claims, citations, source support, and uncertainty in AI-generated answers before publishing or reuse.
-
Added procedures: Chain-of-Verification, semantic-accuracy review, and technical-writing review
Step-by-step AI task guides for source rules, output review, code review, prompt setup, context boundaries, memory boundaries, and publishing measurement.
-
Added procedures: Architecture boundary review, standards-backed implementation review, and scholarly literature review
Step-by-step AI task guides for source rules, output review, code review, prompt setup, context boundaries, memory boundaries, and publishing measurement.
-
Prompt library adds starter bundles
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Prompt library adds reusable controls
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Prompt library adds code and implementation workflows
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Prompt library adds workflow templates by task
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Added workflow: Classify provided material into explicit labels
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Added workflow: Summarize provided material
Use this prompt to summarize files, notes, logs, reports, threads, or pasted excerpts without adding unsupported claims.
-
Published article: Using ChatGPT Effectively at Work
Define a professional task, choose the right ChatGPT experience and capability, place context correctly, and verify the result before use.
-
Published article: Web-retrieved content is a prompt-injection boundary in tool-using LLM systems
Threat model for web retrieval prompt injection in LLM systems, covering external content, tool routing, downstream actions, and control boundaries.
-
Added policy: Code review: choose the right policy
Routing page for choosing the correct review policy for architecture, boundaries, or implementation against official sources.
-
Added procedure: Code Review — choose the correct review path
Decision guide for choosing the right code review method for architecture, implementation, testing, code quality, or refactoring.
-
Added: Prompt components catalog (drop-in micro-prompts)
Reusable AI prompt components for evidence boundaries, verification checks, source review, artifact reading, confidence scoring, and no-fabrication controls.
-
Prompt templates and workflow files reorganized into stacks and components
Browse copy-ready AI prompt templates for research, writing, code review, source verification, analysis, and reusable workflow controls.
-
Published article: Why “Almost Human, But Not Quite” Feels Wrong
Why clowns, AI-generated faces, and fluent AI text can feel uncanny when human-like cues conflict with emotion, realism, or factual coherence.
-
Published article: LLM-Led vs Orchestrator-Led Tool Execution
Comparison of LLM-led and orchestrator-led tool execution in LLM systems, including control-plane placement, reliability, observability, latency, cost governance, and security.
-
Published article: LLM Sycophancy: Evidence, Training Pathways, and Evaluation
Learn what LLM sycophancy is, how preference signals can reward agreement over correction, and how paired-prompt evaluations can detect regressions.
-
Published article: Prompt Engineering Guide for Daily Work
Fix unsupported answers, partial input coverage, unchecked agreement, missing tools, and vague output requirements with explicit prompt controls and tests.