The Illusion of Green Tests: Why Happy-Path AI Code Breaks in Production
AI coding assistants are good at generating feature code and quick happy-path unit tests. Production software fails on edge cases, race conditions, null parameters, and third-party API timeouts, none of which a happy-path test ever exercises.
The more common failure, though, isn't missing tests. It's tests that pass and fail for reasons that have nothing to do with the code: a suite that stayed green for weeks and then broke five times in one afternoon because a CSS class got renamed for an unrelated reason, or because CI got a little slower and a fixed 3-second wait ran out before the page actually finished loading.
The principle of high-assurance QA: tests must be deterministic, isolated from external network dependencies, structured with Arrange-Act-Assert, and query the DOM through user-facing accessibility roles, not implementation details that are free to change.
The 4 Pillars of Modern Automated Testing
1. Playwright Accessibility-First Locators
Never target volatile CSS class names or XPath strings. Query in this priority order instead:
page.getByRole('button', { name: 'Submit' })page.getByLabel('Email address')page.getByPlaceholder('dev@company.com')page.getByText('Payment successful')page.getByTestId('checkout-card'), fallback for genuinely non-semantic widgets
2. Web-First Auto-Waiting Assertions
Eliminate hardcoded delays. Playwright's web-first assertions poll the DOM automatically until the condition is met or the timeout expires:
// Bad (flaky)
await page.waitForTimeout(3000);
expect(await page.isVisible('.success-badge')).toBe(true);
// Good (deterministic auto-waiting)
await expect(page.getByRole('alert')).toHaveText('Payment confirmed');
await expect(page.getByRole('button', { name: 'Download Invoice' })).toBeVisible();
3. Deterministic Third-Party API Mocking
Don't let unit or integration tests make live HTTP calls to payment gateways or cloud providers. Mock external services at the network layer:
vi.mock('stripe', () => ({
default: vi.fn().mockImplementation(() => ({
checkout: {
sessions: {
create: vi.fn().mockResolvedValue({ id: 'cs_test_123', url: 'https://pay.stripe.com' })
}
}
}))
}));
4. Polyglot Table-Driven Tests (Go / Python / Vitest)
Test boundary values, zero, negative, overflow, null, with parameterized, table-driven tests instead of one-off assertions:
func TestCalculateDiscount(t *testing.T) {
tests := []struct {
name string
amount float64
want float64
wantErr bool
}{
{"zero amount", 0, 0, false},
{"negative amount", -10, 0, true},
{"tier 1 discount", 100, 90, false},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, err := CalculateDiscount(tt.amount)
if (err != nil) != tt.wantErr || got != tt.want {
t.Errorf("got %v (err: %v), want %v", got, err, tt.want)
}
})
}
}
Automating Test Suite Generation with Agent Skills
The Agent Test Suite Architect skill audits your application logic and generates high-coverage, zero-flake test suites with a full Test Gap Analysis matrix:
"Using the agent-test-suite-architect skill, audit my billing service and write comprehensive Vitest tests with Stripe API mocks, targeting 85%+ branch coverage and reporting the actually measured result."
Frequently Asked Questions
Why are hardcoded sleeps (page.waitForTimeout) bad in Playwright?
Arbitrary sleeps slow down CI test runs and still fail when network latency exceeds the arbitrary threshold. Playwright web-first assertions automatically wait for elements to be attached, visible, and stable.
What is the recommended Playwright locator priority?
Query elements by accessibility hooks in strict order: getByRole() → getByLabel() → getByPlaceholder() → getByText() → getByTestId(). Avoid brittle CSS selectors.
How does Test Gap Analysis improve QA confidence?
It outputs a markdown table explicitly listing every untested conditional branch or line, explaining the risk level and recommending specific edge-case assertions.
Comments
Comments are reviewed before appearing publicly.
No comments yet — be the first.