Skip to content

Test Execution Handbook ​

Neemias ​

1. Purpose ​

This handbook converts test strategy into execution rules for contributors.

Use this with:

  • mtp.md
  • ../../RTM.csv
  • ../architecture/srs.md

2. Scope ​

Applies to:

  • unit, integration, and system test execution
  • accessibility and compliance verification steps
  • release-gate evidence requirements

3. Requirement Traceability Rule ​

Every behavior-changing PR should reference affected requirement IDs and test IDs.

Minimum PR traceability:

  • Requirement IDs from ../architecture/srs.md
  • Test case IDs in PR summary
  • ../../RTM.csv updates when mappings change

4. Test ID Naming ​

Use stable IDs:

  • TC-RBAC-###
  • TC-ATT-###
  • TC-STU-###
  • TC-SYNC-###
  • TC-CONFLICT-###
  • TC-TTL-###
  • TC-A11Y-###
  • TC-SEC-###
  • TC-COMP-###
  • TC-PERF-###

5. Planned Test Structure for Implementation ​

When code scaffold is introduced, follow this structure:

  • tests/unit/
  • tests/integration/
  • tests/system/
  • tests/accessibility/
  • tests/security/
  • tests/fixtures/

Note:

Tooling commands are intentionally not fixed in this document until the implementation stack is finalized.

6. Execution Order by Risk ​

Run tests in this order when validating release candidates:

  1. Authentication and role gating (SCR-001, SCR-002, FR-001 to FR-008)
  2. Offline queueing and sync conflict rules (FR-014, FR-015, SR-001 to SR-005)
  3. Data integrity and audit retention (DR-001 to DR-005)
  4. Deletion justification and admin-only constraints (FR-007, SCR-007)
  5. Accessibility and UX confirmation (NFR-001, FR-021, FR-022)
  6. Localization and defaults (FR-019, FR-020)

7. Evidence Requirements ​

Each executed test cycle should capture:

  • timestamp and environment profile
  • test ID and requirement mapping
  • pass or fail result
  • logs and screenshots for failures
  • defect ticket reference for failures

Mandatory artifact groups:

  • conflict resolution evidence
  • session expiry and re-auth evidence
  • deletion justification evidence
  • accessibility quick-audit evidence

8. Defect Triage Rules ​

Severity follows mtp.md:

  • Critical: security breach, data loss, conflict rule break, compliance breach
  • High: blocked business flow, incorrect sync outcome
  • Medium: recoverable behavior mismatch
  • Low: cosmetic and wording issues

Gate policy:

  • no open Critical defects for release
  • no open High defects in security, sync, or data integrity domains

9. Manual Accessibility Checklist (Minimum) ​

For each major flow (login, attendance, students, deletion, reports):

  • full keyboard traversal
  • visible focus indicator
  • semantic labels for controls
  • error association to fields
  • non-color-only status cues
  • screen-reader announcement for success and failure states

10. Session and Time-Travel Testing ​

Session expiry tests should use deterministic time controls where available.

Requirements:

  • do not wait real 24 hours in tests
  • advance clock abstraction in test harness
  • verify unsynced queue remains intact across expiry
  • verify protected submission blocked until re-authentication

11. Release Candidate Checklist ​

Before release candidate signoff:

  • requirement coverage report generated against ../../RTM.csv
  • priority suites in mtp.md executed
  • unresolved risks documented
  • accessibility checklist completed
  • decision-impacting findings reflected in ../architecture/adr/ when needed

12. Local E2E runbook ​

The Playwright suites run against a local stack and are deliberately not part of pnpm test:all (#750): the umbrella covers the unit tiers plus the Worker integration suite, while E2E needs a browser and (for the online suite) a Worker process.

Prerequisites ​

bash
pnpm install
pnpm --dir app exec playwright install chromium   # add --with-deps on Linux

No env exports are needed: each config sets its own VITE_* values, and Playwright boots the servers it needs as webServer entries (reuseExistingServer locally, so an already-running server is reused). Ports used: 5173 (Vite), 8788 (wrangler dev, online suite only).

Tiers and commands ​

TierCommandStackSpecs
Smoke (what CI runs)pnpm --dir app test:e2e:smokeVite dev on :5173, VITE_BACKEND_URL="" (offline island), demo seed, 1 worker, no retries, chromium9 specs / 18 tests (~20 s locally)
Safari (nightly CI, safari.yml)pnpm --dir app test:e2e:safariidentical to Smoke but on WebKit — needs playwright install --with-deps webkit (its Linux system libs differ from chromium's)11 specs / 22 tests: the 9 smoke specs plus 08-responsive.spec.ts and 11-accessibility.spec.ts (see below) (~29 s on an idle mac, ~3.5 min on the runner)
Full offlinepnpm --dir app test:e2esame stack, 3 workers locallyall app/e2e/specs/*.spec.ts
Online (what CI runs)pnpm --dir app test:e2e:onlinewrangler dev on :8788 with a fresh local D1 (app/e2e/scripts/e2e-online-worker.sh) + Vite with VITE_BACKEND_URL=http://localhost:8788, VITE_ONLINE_ONLY=true, 1 worker, 1 retry09-modo-facil-checkout.spec.ts, 12-modo-facil-notificacoes.spec.ts, 13-modo-facil-cadastro.spec.ts, 14-modo-facil-qr-checkin.spec.ts

Every test:e2e* script pins --project=chromium on purpose — playwright.config.ts also declares a webkit project (the PWA's primary browser: spec 14 calls iOS/Safari the fila principal), and both the smoke and online configs spread baseConfig. Without the pin, adding that project would silently run every suite on both engines and roughly double the e2e lane. To run Safari locally use test:e2e:safari; to run everything on both engines, pass --project explicitly. Note that desktop WebKit is not iOS Safari — it catches JS/CSS engine divergence, not the iOS install flow or storage eviction, which the manual device checklist owns (docs/quality/modo-facil-device-smoke.md; the canonical item list is the header of spec 14).

The Safari tier is the smoke subset PLUS two specs that no other tier runs (playwright.safari.config.ts imports SMOKE_SPECS from the smoke config and appends them, so the two lists cannot drift):

  • 08-responsive.spec.ts — layout at 375 px (phone) and 1024 px (tablet). This is a phone-first PWA, so a 375 px regression is user-facing; nothing else in CI covers it.
  • 11-accessibility.spec.ts — axe-core over /login and /dashboard, asserting zero WCAG 2.1 AA violations. Its own header records that Students/Attendance/Classes are excluded for color-contrast violations — those three pages are still ungated. Closing that gap is a separate, known task; do not read a green accessibility run as "the app is AA compliant".

Both are deliberately nightly-only: they are "deeper check" work, not the fast pre-push net, and they never ran in any CI before 2026-09-15 (both verified green on chromium and webkit first).

Notes:

  • 02-attendance.spec.ts targets a removed UI and is intentionally outside the smoke subset — rewrite or delete it before adding it back (playwright.smoke.config.ts header).
  • If Playwright reports Executable doesn't exist at .../chromium_headless_shell-<build>/..., the cached browser is older than the pinned Playwright version — rerun pnpm --dir app exec playwright install chromium.
  • The online suite needs workers/.dev.vars (dev admin token + SEED_ENABLED); its globalSetup seeds via POST /api/v1/_seed (a 409 ALREADY_SEEDED is fine). The file is gitignored, so CI synthesizes it in the e2e job.
  • The smoke subset plus the online suite are the CI safety net (both run in the e2e job); pnpm test:all never runs any Playwright spec. Both suites are skipped when the change is docs-only (if: steps.changed.outputs.code == 'true' in the e2e job of ci.yml).
  • The e2e job starts by freeing the E2E ports with app/e2e/scripts/free-e2e-ports.sh (:5173 Vite, :8788 wrangler dev). A job that is cancelled or hits timeout-minutes on the shared self-hosted runner leaves its servers alive, and the next run fails at webServer startup — http://localhost:5173 is already used (run 34787497864) — before a single test runs, and Playwright's retries never cover webServer boot. The script kills listener + supervisor in a loop (a plain kill only makes wrangler dev restart its workerd child) and exits 1 if a port still responds after 20 s, rather than let a stale server silently serve old code. Use it locally too before test:e2e:online: with reuseExistingServer: true a stale worker is otherwise silently reused and the fresh-D1 wipe is skipped.
  • The online specs need the real Worker (D1 + R2), so they cannot move into the offline smoke subset. Because they all log in from one IP, they exceed the login ceiling of 5/60 s; the local E2E Worker is booted with --var RATE_LIMIT_E2E_MAX:500 (ignored when ENVIRONMENT=production — see workers/src/middleware/rate-limit.ts). The per-route ceiling is not just about logins: the screens the specs drive poll GET /api/v1/students hard, and a fast run fired 54 requests in a single 60 s window (read the counts straight from the local rate_limits D1 table), which is why the earlier ceiling of 50 produced a deterministic 429 RATE_LIMIT_EXCEEDED for spec 14 while the same suite passed on a slower runner. If a spec gets a 429, raise the ceiling — never add a login-driven spec's traffic to production's real limit.
  • Diagnosing an online-suite failure: read the artifact, not the spec names. The Worker on :8788 is the suite's infrastructure dependency, and its death mid-run does not look like infrastructure to Playwright: on CI run 34853673151 wrangler dev crashed two minutes in and the three specs that followed failed as Nenhuma opção disponível para a turma (seed do D1 local vazio?) and connect ECONNREFUSED ::1:8788 — zero bugs, three "failures". So:
    • The e2e job uploads app/test-results/ as the e2e-online-diagnostics artifact on every run (if: always()), which includes /tmp/wrangler-e2e.log and /tmp/wrangler-e2e-migrations.log copied into test-results/wrangler/ — that is where the wrangler error message actually lives.
    • The online specs import test/expect from app/e2e/fixtures/online.ts instead of @playwright/test. Its onlineWorkerAlive fixture is auto, polls GET /api/v1/health before each test, and fails with Infra: … o wrangler dev caiu when the Worker is gone — so an infra death is labelled as one instead of being misattributed to the screen under test. New online specs must import from that fixture, not from @playwright/test.
    • retries: 1 covers runner contention (cold Vite, timeouts), not a dead Worker: every attempt of a spec whose Worker is gone still fails the health poll. The specs are retry-safe by construction (09 re-picks a free student, 13/14 use uniqueName and stop before the check-in is effective); check that property before adding a spec that mutates shared state.

Cost of these tiers, and why changing them is harder than it looks ​

The wall clock of the CI, where it actually goes (measured on the runner), the cost model of the unit tiers (≈1.3 s per jsdom file, of which only ~0.13 s is real test work), and every "obvious" optimization that was measured and rejected — workers: 2, higher maxWorkers, happy-dom, isolate: false, deps.optimizer, trimming specs — are documented in docs/research/ci-cost-model-contention-and-levers-2026.md. Read it before touching the tier configuration or adding a spec to the smoke subset: the short version is that the CI is contention-bound, not work-bound, so parallelism knobs do nothing, and the remaining levers are structural (a second runner, selective runs).

13. Ownership and Updates ​

Contributors updating requirements or architecture must update this handbook if execution expectations change.

Do not leave process-critical changes undocumented.

Distribuído sob licença MIT.