Vibe Coding Risks: Code Defects and Verification Gaps

By Published Updated

A working UI can hide defects in AI-generated code. Learn what to verify and how engineering ownership guides correction, acceptance, and maintenance.

Executive summary

This article examines vibe coding as a development workflow in which a user requests code, runs it, and asks for changes until the visible result seems satisfactory. The engineering risk is accepting that result without establishing whether the implementation satisfies the relevant requirements and system constraints.

The analysis covers both defects in generated code and gaps in the process used to evaluate it. It explains how functionality, authorization, integration, test quality, and maintainability need to be examined beyond an initial successful run. Engineering ownership connects these checks to a decision: who is responsible for assessing the change, resolving identified problems, and documenting limitations that remain?

1. Introduction: what remains unverified

A successful demonstration shows an observed behavior under particular conditions. It does not establish compliance with requirements that the demonstration never tested.

For example, a test showing that an authorized user can retrieve a record does not establish that other users are prevented from retrieving it. These are separate requirements. Authorization must be checked for the requested operation and resource, including requests that should be denied. OWASP’s authorization guidance explicitly calls for permission checks on every request and tests derived from the application’s access rules. OWASP Authorization Cheat Sheet.

This distinction provides the basis for examining the workflow. A missing access check is an implementation defect. An undefined access rule is a requirements gap. A test suite that exercises only permitted requests leaves a verification gap. Each requires a different correction.

The sections that follow apply this analysis to the technical weaknesses and missing engineering work that a working result can conceal. They examine what can fail, why the initial evaluation may miss it, and which checks are needed before acceptance. This extends the general requirement to review AI-generated code into specific engineering questions. GitHub’s own guidance calls for understanding suggestions, reviewing their functionality, security and maintainability, and using automated checks. GitHub Copilot best practices.

2. What vibe coding means

Vibe coding is an informal term, not an engineering standard. Collins defines it as the practice of writing computer programs by using natural-language prompts to make a generative AI system output the desired code. Merriam-Webster defines vibe coding as the practice of using an AI system to generate computer code in a given programming language.

This definition matters because it places vibe coding within a broader shift from direct programming to intent-based code generation. Instead of writing the logic directly, the user describes the intended outcome, and the model generates a possible implementation.

Code generation supplies a candidate implementation. Its suitability for the task must be established through review and testing against explicit requirements, with responsibility for resolving any discrepancies.

A distinction is therefore necessary between AI-assisted coding and vibe coding in the critical sense.

AI-assisted coding refers to the use of AI as a support tool inside an engineering workflow. The developer may use the model to generate, explain, or improve code, but still reads the code, understands the logic, validates the behavior, and owns the result.

Vibe coding in the critical sense describes a different pattern: the user describes intent, receives code, runs it, asks for iterative fixes, and approves the result mainly because it appears to work.

3. Why “the code works” is not enough

Visible UI behavior and completion of an application operation are separate properties. A button responding, a table updating, or a confirmation appearing establishes what the browser displayed. It does not by itself establish that the required service was called, the business rules were applied, or the intended change was stored.

For a feature that must save records to a server, updating an array in browser memory can make a new record appear without persisting it. A mocked API can also return predefined data without calling the service it replaces. Both can support interface development while leaving the required backend integration unimplemented or untested.

The gap can remain even when a network request is sent. JavaScript’s fetch() resolves with a response when the server returns an HTTP error status; it does not automatically reject that response. Code that displays success merely because await fetch() completed can therefore report success after a failed request. The implementation must inspect the response and interpret it according to the API contract.

Two client implementations handle HTTP 500: unchecked code displays Saved; a response.ok check displays the error and exits.
For the same HTTP 500 response, unchecked code displays Saved. Checking response.ok routes execution to the error handler and returns before success.

Verification must follow the operation beyond the visible interaction. For a server-backed save, check that the request reaches the intended service, carries the expected data, and produces the required stored state. Retrieve the record independently through the service and compare its values with the submitted data. A page refresh alone is insufficient if the display can be restored from browser storage or a mocked response. Also check that a rejected operation produces an error state rather than a success confirmation.

In the workflow examined here, accepting the result as soon as the interface changes leaves these completion conditions unverified. The following subsections examine additional failures that can remain even after the underlying operation is connected.

3.1. Failure handling can introduce incorrect behavior

Consider a request that creates a record in a remote service. The service may complete the operation even if its response never reaches the caller. Adding an automatic retry can appear to fix the timeout while creating a second record.

A demonstration in which every request receives a response will not expose this problem. The missing requirement concerns what should happen when the caller cannot determine whether an operation completed.

The implementation therefore needs an explicit retry policy and, where repeated requests must not repeat the effect, an idempotency mechanism. Idempotency means that repeating the same logical operation produces no additional side effects. Verification should simulate a lost response after completion, retry the operation, and inspect the resulting state for duplicates. Receiving a successful response on the retry is insufficient.

3.2. A local change can break an existing interface

An API contract defines the inputs, outputs, and error behavior that callers can rely on. A generated revision might change a response field from a number to a string while the page being demonstrated still displays it successfully. Another caller that performs calculations could then fail.

Review must examine the affected callers and the established contract. Integration tests should exercise those interactions and verify required types, fields, and error responses. A test of the revised function alone cannot establish compatibility with its consumers.

3.3. Ordinary inputs can conceal unsafe data handling

A search feature can return correct results while constructing SQL queries by concatenating user input. The defect becomes apparent when input changes the query’s meaning instead of remaining a value to search for. Testing ordinary search terms does not examine this boundary.

The required correction is to separate SQL instructions from input values through parameterized queries. Where identifiers such as column names cannot be bound as parameters, user choices should map to explicitly permitted identifiers.

Verification must inspect how input reaches the database and test that unusual input remains data. Adding validation to the visible form alone leaves requests made directly to the server outside that control.

3.4. Passing tests can leave the required behavior unchecked

A test may call a calculation function and assert only that its result is a number. The function could return an incorrect total and still pass. Executing the relevant code does not establish that the test checks its meaning.

Expected results must be justified by the requirements. For a calculation, that includes explicit examples with independently established totals and relevant boundary conditions. Assertions should check those values and any required changes to system state.

Code coverage helps identify unexecuted code, but it cannot establish that assertions detect incorrect behavior. Mutation testing can provide additional evidence by introducing small changes to the implementation and checking whether the tests detect them.

3.5. Successive fixes can leave behavior inconsistent

Suppose a revision adds a validation rule to one request handler while another handler retains a separate copy of the same rule. The demonstrated route may now behave correctly, but the two routes can disagree. A later change also requires finding and updating both implementations.

Review should examine the complete change and surrounding code, identify where the rule belongs, and check every affected caller. Where both routes must enforce the same rule, a shared implementation with explicit tests can reduce divergence. This makes maintainability concrete: required changes should have identifiable locations and verifiable effects.

4. How defects pass through the approval process

A defect can survive an approval process that includes tests when those tests reproduce the implementation’s mistaken assumptions. The resulting agreement establishes consistency between the code and its tests, while leaving compliance with the requirement unresolved.

When the test repeats the implementation’s error

Google’s testing guidance provides a concrete example. A navigation test constructs its expected URL by joining a base URL ending in / with a path beginning in /. The expected result therefore contains an unintended double slash in the path. If the implementation makes the same error, the equality assertion can pass.

The verification gap is in the expected result: it was computed using logic that can reproduce the defect being tested. Writing the intended destination explicitly makes the discrepancy visible. Once that expectation is corrected, an implementation producing the extra slash will fail the comparison. The implementation then requires correction and another test run.

How this applies to AI-generated code

This mechanism matters when an assistant supplies both an implementation and the material used to approve it. If its tests adopt an unchecked assumption from the implementation, the test suite can preserve that assumption. An accompanying explanation may describe the same behavior without establishing that the behavior is required.

Using the same assistant for code and tests does not itself make the tests invalid. What matters is how the expected behavior was established. A reviewer needs to check the expected results against the relevant requirements and confirm that the assertions would detect a violation.

The approval decision must therefore connect each material requirement to evidence that actually evaluates it. If a failing test exposes a disagreement between the implementation and its expected behavior, that disagreement requires investigation. Changing the expectation to match the generated output is justified only if the requirement supports the change. Otherwise, it removes the failure signal while leaving the implementation defect in place.

5. Research evidence: the gap between code that looks useful and code that has been validated

Research on LLM-based code assistants points to a recurring risk: models can produce code that appears useful while still containing security weaknesses or incorrect assumptions that users may not identify at approval time.

In their study of GitHub Copilot, Pearce et al. evaluated 89 scenarios related to Common Weakness Enumeration categories and generated 1,689 programs. Approximately 40% of the generated programs were found to be vulnerable. The central finding is not that Copilot always generates vulnerable code. The finding is that code generated by such tools can appear usable while still containing significant security weaknesses.

The user study by Perry et al. adds an additional layer to the argument. The issue is not only what the model generates, but how users accept and trust the output. Participants who had access to an AI code assistant wrote less secure code than those without access, and they were also more likely to believe that their code was secure. In other words, the risk is not only in model output. It is also in the approval behavior and confidence calibration of the user.

Together, these studies document security weaknesses in generated code and, in Perry et al.’s study, a gap between users’ confidence and the security of their solutions. They support examining both the implementation and the basis on which users judge it. The article’s practical response is to connect identified problems to specific corrective work and an accountable acceptance decision.

6. Engineering ownership: who is responsible for the change

Engineering ownership assigns responsibility for defining a change, evaluating its implementation, resolving findings, and maintaining the accepted code. For each responsibility, the team needs to identify who performs the work and who has authority to make the relevant decisions.

NIST SP 800-218 provides a foundation for this allocation. Its SSDF practices address roles throughout the software development life cycle, criteria for security checks, and the recording and assessment of findings. The following allocation applies those principles to the workflow examined here.

Responsibility Who takes responsibility What that responsibility requires
Define the required behavior The feature or product owner, with the relevant domain and technical specialists Resolve business rules, access rules, and interface constraints; establish the expected behavior against which the implementation will be assessed.
Implement and correct the change The developer responsible for submitting it Examine the generated code, identify the components it affects, implement the agreed behavior, and correct confirmed defects.
Verify the implementation Assigned reviewers with expertise relevant to the affected code and risks Assess the code and test expectations against the requirements; examine the recorded results and identify missing checks.
Authorize acceptance The person or group designated to approve the change for its intended use Decide whether the evidence meets the acceptance criteria and obtain a decision from the appropriate authority for any proposed exception.
Maintain the accepted code A named maintainer or owning team Take responsibility for subsequent defects and changes, with access to the code, tests, relevant design decisions, and outstanding issues.

These are responsibilities to allocate within the team’s existing workflow. One person may perform several of them where the review requirements permit; specialist assessment may be needed for changes involving unfamiliar security or architectural constraints.

The allocation should also make maintenance possible in practice. The maintainer needs to know where the affected behavior is implemented, which dependencies and callers it relies on, and which checks to run after changing it.. Relevant decisions should be retained with the project so that later work does not depend on reconstructing the original prompting conversation.

7. What must be established before acceptance

Acceptance applies to a specific code revision and intended use. The approver should be able to trace each material acceptance claim to the relevant review or test result, including the conditions and integrations actually checked. A test performed against a mocked service, for example, cannot establish that the real service integration was tested.

Resolve findings against the requirements

An open finding should identify the affected requirement, the observed discrepancy or missing evidence, and the person responsible for the next action.

If the expected behavior is unclear, the requirement owner must resolve it before the implementation can be judged against it. If the implementation violates an agreed requirement, the developer must correct it and the affected behavior must be checked again. If the evidence is insufficient, the reviewer must identify the additional assessment needed.

Closing a finding requires a recorded basis: evidence supporting the correction, or an explanation supported by evidence showing why the reported behavior does not violate the requirement. An unresolved disagreement should be referred to the person authorized to decide the requirement or acceptance condition.

Record the acceptance decision

The approval record should identify the assessed revision, intended use, supporting evidence, and remaining limitations. It should also identify the approver and the maintainer responsible for subsequent work.

Where the applicable acceptance process permits an exception, the authorized decision-maker should record the unmet condition, its consequences, the justification for proceeding, and any restrictions or follow-up work. That work needs an owner and, where applicable, a deadline or condition for reassessment.

An accepted exception remains a limitation; it should not be recorded as a verified correction. If a required condition remains unmet and no permitted exception has been authorized, acceptance should remain pending.

8. Conclusion

Vibe coding can produce a visible result before the underlying operation and its constraints have been adequately evaluated. Review must therefore establish both whether the implementation satisfies the requirements and whether the checks performed support that judgment.

Engineering ownership connects these assessments to a decision. It assigns responsibility for resolving implementation problems, settling requirements, and obtaining missing evidence, then determining whether the result is suitable for its intended use.

Completion is supported by the assessed behavior, the corrective work performed, and the limitations recorded in the acceptance decision.

References

  1. Collins Dictionary. “Vibe coding.”

  2. Merriam-Webster. “Vibe coding.”

  3. Google for Developers. “Gemini Code Assist and responsible AI.”

  4. NIST SP 800-218. “Secure Software Development Framework (SSDF) Version 1.1: Recommendations for Mitigating the Risk of Software Vulnerabilities.”

  5. OWASP. “Code Review Guide.”

  6. Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. “Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions.”

  7. Perry, N., Srivastava, M., Kumar, D., & Boneh, D. “Do Users Write More Insecure Code with AI Assistants?”