관리
← All articles

Asking Codex for Tests: Normal Cases, Boundaries, and Failure Conditions

This article was translated from its source language with AI assistance. Please check technical terms and equations against the original.

Asking Codex to add tests can produce code checking only one normal input. What matters is the behavior to guarantee, more than the number of test files. Defining inputs, expected results, and failure handling first also helps assess whether AI merely copied the implementation into its tests.

Asking Codex for Tests: Normal Cases, Boundaries, and Failure Conditions — Original concept illustration
Original concept illustration

This article uses a hypothetical reservation-quantity validator. Its numbers do not describe actual product limits or Codex usage. The request and case table are templates to adapt to your project, not claims that its tests were executed here.

Normal, boundary, and failure cases for quantity validation at a glance

Explanatory illustration — not an actual screen or test result.

1. Define the function's input/output contract

2. Write normal inputs and expected results

3. Distinguish immediately inside and outside boundaries

4. Specify wrong types, missing values, and failure results

5. Separate actual execution results from unexecuted items

1. Define the function's promise before testing

Let the example validate_quantity accept integers from 1 through 5 and return validation errors for other inputs. Success and error formats must follow the actual project contract. Define desired behavior before writing expected values, instead of treating outputs observed from the implementation as the expectations.

The example disallows 0, 6, negatives, fractions, strings, and missing inputs. Whether numeric-looking strings are automatically converted is another decision. “Success for an appropriate quantity” makes Codex guess what appropriate means, so fix the permitted range in words and a table.

OpenAI's Codex modernization example shows planning normal flows and edge cases with inputs and outputs for each scenario. Official prompting guidance also describes identifying the tested function and requesting normal and boundary cases. The template below applies those principles to a small original example.

2. Choose normal cases representing the intended use

Passing 3 and expecting success is a normal case. Rechecking the same middle value under several names does not add important conditions. Verify results users actually rely on, such as return type, normalized quantity, and required post-call state.

If the function only validates, there is no reason to check database changes. If a reservation API's contract stores successful requests once, check stored quantity and response together. Separating function-unit and API-through tests makes failure causes easier to find.

Type Illustrative input Contractual expectation
Normal middle value 3 Success; quantity remains 3
Minimum allowed value 1 Success
Maximum allowed value 5 Success
Below minimum 0 Validation error
Above maximum 6 Validation error
Different type String 3, decimal 3.5 Validation error without conversion

These expectations follow the example contract. A service allowing strings needs a different table. Leave unsettled contract details as questions or undecided items before testing; do not automatically adopt AI-chosen outputs as product requirements.

3. Check both immediately inside and outside boundaries

To test a maximum of 5, check success at 5 and failure at 6 together. Checking only 4 can miss an implementation wrongly rejecting 5. Likewise connect success at 1 with failure at 0 for the minimum. Preserve the reason for each boundary in test names or brief explanations.

Count and date boundaries have different units. Counts compare neighboring integers; reservation times require timezone, inclusivity, and input precision. Testing “just before the deadline” requires specifying the deadline first. Do not transfer this quantity table unchanged to every time condition.

A contract rejecting fractions also needs a separate failure case such as 1.5. Merely checking minimum and maximum can miss type violations. Check how the language handles booleans and integers, and independently decide whether the product accepts booleans as quantities.

4. Define error content and subsequent state for failures

Asking Codex for Tests: Normal Cases, Boundaries, and Failure Conditions — Original illustration of the key points
Original illustration of the key points

Specify the required error format rather than saying only “it should fail.” Tests differ for thrown exceptions, returned error objects, and API responses. Ask Codex to read the actual project's existing error-handling conventions.

If the contract prohibits saving invalid inputs, check that too. A hypothetical reservation API can require the reservation count to remain unchanged after a validation error. Checking only the message and ignoring storage tests only part of the expected result.

Decide whether the entire error sentence needs to be fixed. Test a guaranteed error code when required; frequently changing display text can make exact sentence comparisons burdensome. Choose assertion strength according to the contract.

5. Prepare a request for Codex

Include the function, related files, permitted inputs, expected results, existing test locations, and execution method. Use paths verified in the actual repository. Rather than making Codex infer whether new tools are needed, ask it to inspect existing tools and commands first.

Read validate_quantity and existing tests, then add tests using the same conventions. The contract permits only integers 1–5 and performs no automatic string conversion. Distinguish successes for 3, 1, and 5 from failures for 0, 6, negatives, fractions, strings, and missing input. Confirm boolean handling and error format from the existing contract; tell me first if unclear. Report contract differences before changing the implementation to fit tests. Run relevant tests and record commands, outcomes, and conditions that could not be executed.

This request does not fit every project. Nonexistent function names or commands can prevent verification from starting. Check arguments, calling paths, and preparation of existing dependencies, and do not include unnecessary deployment work.

6. Identify tests merely copying the implementation

When reading generated tests, check where expected values came from. Calculating expectations with the function being tested can calculate the same bug twice and pass. Obtain example limits from an independent contract table and directly verify important output fields.

Imagine an incorrect implementation rejecting 5. A test of the maximum allowed value should fail to enforce the range contract. A test checking only 3 may miss this mistake. This thought experiment checks what tests detect; it is not an actual code-change experiment.

If mocks are used, read what they replace. A repository mock does not prove actual storage on an external server. Separate checking a function's request to store correct inputs from checking whether the external system actually worked.

7. Distinguish execution results from file existence

Reported state Evidence to check Meaning
Tests written Added code and condition table Checking prepared
Tests discovered Runner-selected test list Whether targets are included
Execution passed Command, exit result, and pass count Result within the executed scope
Skipped Conditions and skipped items Unverified scope
Could not run Environment error and cause Verification remains

A successful command selecting zero tests does not verify the intended conditions. Read filters, working location, and actual discovery count. If dependencies block execution, report product-code failure separately from environment-preparation failure.

Official prompting guidance explains running small relevant tests and checks after changes, then reporting results. Rather than retaining only “success,” record which commands checked what, helping reviewers understand scope. Do not count unexecuted conditions as completed.

8. Compare contract, tests, and implementation after failure

Compare actual input, expectation, result, and the contract table. A test may have incorrect expectations, or the implementation may violate the contract. Changing only expectations to current outputs can erase intended behavior; explain the evidence for any change.

For a product decision changing the contract, record its reason and affected conditions before updating tests. For an implementation bug, fix the behavior and check that the same failure is resolved while nearby conditions remain correct. Investigate new failures in previously passing checks as well.

Final questions are “Does it represent normal use? Does it check both boundary sides? Does it check failure results and state? Is actual execution scope clear?” Expressing desired behavior as a contract and requesting independent verification connects the tests' purpose with Codex's result.

Official sources and writing standards

References checked: 2026-10-07. This explanation was written with AI assistance based on official sources actually opened. Separately marked calculations, code, and checking examples are illustrative, not results from directly testing or measuring the user environment. Recheck changes to features and references on the publication date.

Original illustrations created to help explain this article.

Original on Tistory ↗