Published Sep 29, 2026 ⦁ 16 min read
Master Advanced Pytest Techniques for Reliable Testing

Master Advanced Pytest Techniques for Reliable Testing

For many engineers, pytest starts as a cleaner alternative to unittest: fewer boilerplate lines, better assertions, and an easier path to writing tests. But once your codebase grows, the basics stop being enough. You need tests that are strict, expressive, debuggable, and fast enough to trust in CI.

That is the real value of advanced pytest usage.

In this session, longtime Python and pytest contributor Freya Bruhin walks through the techniques that separate "we have tests" from "our tests actively protect the codebase." The talk is not about introductory syntax. It is about using pytest deliberately: matching exceptions precisely, managing fixture scope without breaking isolation, making warnings fail builds, structuring parameterized tests for clarity, and debugging failures with less friction.

For mid-career developers moving toward data engineering or AI engineering, this matters more than it might seem. Production data systems are full of edge cases: schema drift, floating-point quirks, flaky integrations, slow pipelines, and stateful components. A sophisticated testing approach is often what turns a promising project into a dependable one.

Key Takeaways

  • Use pytest.raises(..., match=...) to test failure modes precisely, not just whether "some error" happened.
  • Prefer pytest.approx for floating-point comparisons instead of exact equality.
  • Treat warnings as failures by default so deprecations and risky behavior don’t silently accumulate.
  • Enable strict settings in pytest to catch unknown markers, invalid config, and unexpected xpass results early.
  • Use xfail for known bugs, not as a hiding place for flaky tests; flaky behavior should be handled explicitly.
  • Design fixture scope carefully: broader scope improves speed, but can quietly destroy test isolation.
  • Use layered fixtures when you want both caching and clean state between tests.
  • Make parameterized tests readable with custom IDs, per-case marks, and data classes for complex scenarios.
  • Learn pytest CLI debugging options like --lf, --ff, --stepwise, --tb=short, and --durations.
  • Treat conftest.py as part of your test architecture, not just a dumping ground for helpers.

Why Advanced Pytest Skills Matter More Than Ever

A lot of teams undervalue test design. They focus on code correctness, but not on the quality of the test suite itself. The result is familiar:

  • tests that pass for the wrong reason
  • fragile assertions that break on harmless changes
  • slow runs that developers stop trusting
  • skipped failures that hide real regressions
  • fixture state bleeding across test cases
  • CI noise that obscures the real signal

For engineers working in data-heavy systems, these problems are amplified. Data platforms and AI pipelines often combine numerical logic, external dependencies, stateful services, and asynchronous workflows. Poor testing practices in those environments don’t just slow development; they create real operational risk.

Bruhin’s talk is useful because it frames pytest not as a convenience library, but as an engineering tool for precision and feedback quality.

Start by Tightening Failure Assertions

One of the simplest advanced habits is also one of the most valuable: stop asserting only that an exception occurred. Assert why it occurred.

Use pytest.raises to validate intent

Most Python developers know this pattern:

with pytest.raises(ValueError):
    do_something()

That’s fine as a baseline. But it can also be dangerously broad. If your code can raise the same exception type for multiple reasons, the test may still pass even when you’re exercising the wrong path.

Bruhin’s point is practical: if two error conditions both raise ValueError, your test should distinguish between them.

Add match= to make tests stricter

pytest.raises supports a match argument that applies a regular expression to the exception message.

That matters because it makes your tests:

  • more self-documenting
  • less likely to pass accidentally
  • better at catching regressions in control flow

For example, a parser might raise the same error type for both invalid syntax and an unsupported value. Matching part of the message lets the test prove it hit the intended case.

This is especially useful when:

  • the exception class is intentionally broad
  • custom error hierarchies would be overkill
  • the message contains meaningful context

That said, Bruhin also hints at an important design principle: if you constantly need to inspect error text, your exception model may be too coarse. In mature systems, introducing custom exception types can reduce brittleness and improve clarity.

Inspect exception objects when needed

Using as excinfo gives you access to the actual exception object. That enables richer assertions on:

  • the message
  • custom attributes
  • structured error context

This is especially relevant in systems code, data validation, and parser-like logic, where the exception may include metadata such as expected tokens, invalid fields, or offending values.

A subtle but important implementation detail: assertions on excinfo must happen outside the with block, because execution stops when the exception is raised.

Don’t forget warnings

Bruhin also points out that pytest has an equivalent for warnings: pytest.warns. In mature codebases, warning behavior can be just as important as exceptions, especially around deprecations and migration work.

Floating-Point Equality Is a Trap

If your work touches analytics, pipelines, feature engineering, model metrics, or infrastructure that processes numeric values, exact float comparison is a common source of misleading failures.

Why 0.2 + 0.1 != 0.3

The issue is not Python-specific. Binary floating-point formats cannot exactly represent many decimal fractions. So values that look simple in source code become approximations in memory.

That means an assertion like this is risky:

assert result == 0.3

Even when the logic is correct, representation error may make the test fail.

Use pytest.approx

Bruhin recommends pytest.approx, which integrates cleanly with assertion output and expresses intent clearly:

assert result == pytest.approx(0.3)

This is preferable to hand-rolled tolerances because it is:

  • readable
  • idiomatic
  • built for test diagnostics

You can also specify custom tolerances when domain logic requires it. That’s especially useful for:

  • sensor data
  • scientific computing
  • hardware-adjacent systems
  • approximate metrics in ML workflows

For professionals moving into AI engineering, this habit is not optional. Many model, inference, and preprocessing tests should validate acceptable numerical closeness, not exact byte-level equality.

Make Warnings Loud Enough to Matter

One of the strongest opinions in the talk is also one of the most actionable: warnings should usually fail the test suite.

That may sound strict, but it solves a real problem. In many teams, deprecation warnings appear in output for weeks or months without anyone noticing. Then an upstream package releases a breaking version, and the team is surprised by a failure they were already warned about.

Turn warnings into errors

You can do this for a single run via the command line, but Bruhin recommends putting it into project configuration so it becomes a team default.

The bigger lesson here is architectural: tests should surface maintenance risk early, not merely validate current correctness.

For teams building long-lived internal platforms, this matters a lot. Data and AI stacks often depend on large dependency graphs. If you ignore warnings, you are effectively postponing dependency management until it becomes urgent.

Use targeted ignores, not blanket suppression

Bruhin describes a practical compromise:

  • treat all warnings as errors
  • add narrow ignore rules only for known third-party issues
  • scope those ignores tightly by message, category, and module

That’s a disciplined strategy because it preserves signal. Instead of muting the whole warning system, you document specific exceptions while keeping pressure on the rest of the codebase.

A smart extension for professional teams is to pair this with issue tracking. If you suppress a third-party warning, record why and revisit it periodically.

Turn on Strict Mode

pytest has accumulated enough flexibility over time that it can be permissive in ways that hurt maintainability. Bruhin recommends enabling strict behavior, especially in newer versions.

What strictness buys you

Strict settings can catch:

  • unknown config options
  • unregistered markers
  • expected-failure tests that suddenly pass

That last case deserves attention.

If a test is marked xfail because of a known upstream bug, and the underlying problem gets fixed, you want to know. Otherwise your suite keeps carrying dead annotations and your team misses a valuable signal: the ecosystem improved.

For teams working across fast-moving Python libraries, strictness is a force multiplier. It helps keep the test suite aligned with reality instead of historical assumptions.

Use Marks as Metadata, Not Just Labels

Many developers think of marks purely as tags like slow, integration, or webtest. Bruhin broadens that view. In pytest, marks also act as metadata carriers.

That mindset helps explain why features like parameterization, skipping, and expected failures are mark-based.

Practical mark patterns

You can apply marks at multiple levels:

  • individual test
  • class
  • module, via pytestmark

This is more than convenience. It lets you shape test behavior at the right level of abstraction. For example:

  • mark an entire module as requiring an optional dependency
  • skip a work-in-progress test file temporarily
  • apply environment-specific behavior to a whole test class

importorskip is underrated

If a dependency is optional, pytest.importorskip is often cleaner than manual conditionals. It expresses that the test is valid only when a module is available.

That pattern is especially relevant for data teams who support optional connectors, database backends, or acceleration libraries.

skip and xfail Are Not the Same Tool

This distinction is one of the most important conceptual points in the session.

Use skip when the test cannot run

A skipped test is appropriate when preconditions are missing:

  • required package not installed
  • external service unavailable
  • environment does not support the feature

In short: the test is not meaningful under current conditions.

Use xfail when the behavior is known-broken

An expected failure means the test should pass in principle, but currently doesn’t due to a known bug or limitation.

That difference matters because xfail still executes the test and tracks its status. It gives you visibility into whether the problem still exists.

Bruhin gives a compelling maintenance argument here: old bugs in dependencies or the standard library may be fixed years later. A strict xfail setup helps you discover that automatically.

Don’t use xfail to hide flaky tests

This is a subtle but important critique. Teams sometimes misuse xfail as a bandage for instability. That weakens trust in the suite and confuses bug tracking.

If a test is flaky, treat it as a flakiness problem:

  • fix the nondeterminism
  • isolate the environment issue
  • or, as a last resort, use a retry plugin designed for that purpose

That separation improves test hygiene and team communication.

Parameterization: Powerful, but Easy to Misuse

Parameterized tests are one of pytest’s best features, but once test cases grow beyond simple tuples, they can quickly become unreadable.

Bruhin’s advice is useful because it focuses on maintainability, not just syntax.

Stack parameterization carefully

Stacking multiple @pytest.mark.parametrize decorators creates combinations automatically. That’s elegant, but it can also create combinatorial explosion.

This is where experienced engineering judgment matters. More test cases are not always better if they:

  • slow collection dramatically
  • duplicate low-value coverage
  • make failures harder to reason about

A good heuristic: use Cartesian products when combinations reflect meaningful risk, not just because the syntax allows it.

Mark individual parameter cases

pytest.param lets you attach marks or IDs to specific test cases. This is a major upgrade over plain tuples because it enables:

  • per-case xfail
  • per-case skip
  • readable test names

That level of control becomes critical in mature systems where not every input case has the same status or expected behavior.

Give test cases custom IDs

Readable IDs improve failure triage more than most teams realize.

If a CI job says test_transform[case_17] failed, that may be meaningless. If it says test_transform[invalid_timestamp] failed, the debugging path is immediately clearer.

This is particularly valuable in data engineering, where a single test function may cover dozens of edge cases across formats, null handling, parsing, and constraints.

Use data classes for complex cases

One of the more advanced suggestions in the talk is to use Python data classes to structure parameter sets.

This can improve:

  • readability
  • type safety
  • editor autocomplete
  • default handling
  • self-documentation

Instead of passing opaque tuples full of booleans and values, you create named fields and a more descriptive representation.

That approach is especially effective when testing:

  • parser behavior
  • configuration-driven logic
  • validation rules
  • branching workflows

For engineers building portfolio-quality projects, this is a strong pattern. It shows thoughtfulness in test design, not just feature coverage.

Subtests: Useful, but Not Always Better

Pytest now supports subtests, and Bruhin presents them as a useful addition with tradeoffs.

When subtests help

Subtests are valuable when you want multiple checks inside one broader test flow and want execution to continue even if one subcase fails.

That can help in scenarios where:

  • setup is expensive
  • grouped checks belong together conceptually
  • you want a consolidated view of related assertions

Why parameterization is still often preferable

Bruhin generally prefers parameterization because parameterized cases are first-class tests:

  • collected independently
  • selectable from the CLI
  • better supported by tooling and plugins

That is a meaningful distinction. In professional environments, test discoverability and tool compatibility matter. Subtests provide flexibility, but parameterization usually provides cleaner integration.

Fixtures Are More Than Setup Helpers

Most Python developers use fixtures, but many stop at the basics. The deeper value of fixtures is that they provide dependency injection, lifecycle management, and reusable test architecture.

Bruhin emphasizes two fixture best practices that deserve wider adoption.

Add type hints and docstrings

Fixtures can become opaque quickly, especially in larger teams. Type annotations and docstrings reduce that opacity.

This improves:

  • editor support
  • onboarding
  • fixture discovery
  • long-term maintainability

For technical professionals trying to level up, this is an important mindset shift: test code deserves the same clarity standards as production code.

Understand scope as caching, not visibility

This is one of the most useful conceptual clarifications in the talk.

Fixture scope controls reuse duration, not who can access the fixture. A broader scope can significantly speed up tests by reusing expensive setup across cases.

But it introduces risk.

If the shared object is mutable, one test can influence another. The suite gets faster, but less trustworthy.

That tradeoff appears constantly in real systems:

  • shared Spark sessions
  • temporary databases
  • API clients
  • model artifacts
  • browser instances
  • cached datasets

A better pattern: layered fixtures

Bruhin recommends separating:

  1. the expensive setup fixture
  2. the per-test cleanup/reset fixture

This is a high-value pattern.

It lets you:

  • cache heavyweight initialization
  • preserve test isolation
  • avoid full re-creation costs between tests

In data and ML contexts, this is analogous to reusing infrastructure while resetting state:

  • reusing a database container but clearing tables
  • reusing a model object but resetting runtime config
  • reusing a filesystem sandbox but emptying working directories

It’s a practical compromise between speed and correctness.

Use Yield Fixtures for Cleanup

yield-based fixtures provide a clean setup/teardown pattern:

  • code before yield runs before the test
  • code after yield runs as teardown

This is one of the most elegant parts of pytest, especially when compared with manual cleanup logic.

It is a strong fit for:

  • temp resources
  • external connections
  • environment changes
  • monkeypatch restoration
  • service lifecycle management

The bigger advantage is not syntax. It is reliability of cleanup behavior, which becomes essential when tests interact with real systems.

Autouse Fixtures: Convenient, but Use Sparingly

Autouse fixtures run automatically wherever they are visible. This can be powerful for shared setup such as:

  • environment variable management
  • global patching
  • common reset behavior

But it also increases invisibility.

That’s the tradeoff: less repetition, more implicit behavior.

Bruhin mentions a less hidden alternative: using usefixtures markers. That can be a better choice when you want automatic fixture execution without making it entirely invisible.

For team environments, this is more than style. Excessive implicitness makes test debugging and onboarding harder. If an autouse fixture changes behavior globally, a new engineer may struggle to understand why a test works.

The request Fixture Unlocks Dynamic Behavior

The request fixture is where pytest starts to feel like a framework, not just a library.

It gives fixtures access to information about the test requesting them. That enables:

  • dynamic fixture selection
  • marker-driven configuration
  • introspection of the active test node
  • access to config state

Marker-driven fixtures are a powerful pattern

Bruhin demonstrates a pattern where a fixture adapts based on a test’s marker arguments. That’s a sophisticated way to pass metadata into setup without hardcoding behavior.

This can be especially useful in larger suites where the same fixture supports multiple modes:

  • different config profiles
  • alternate prompts or inputs
  • environment-specific settings
  • backend variants

For data platform testing, the same technique can support:

  • multiple storage engines
  • multiple schema modes
  • feature flags
  • strict vs permissive validation

The key architectural idea is elegant: tests declare intent through metadata, and fixtures interpret that intent.

Learn the Command Line if You Want to Debug Faster

One of the most pragmatic parts of the talk is Bruhin’s emphasis on the CLI. IDE integrations are convenient, but advanced debugging often requires direct command-line fluency.

That advice is especially relevant for engineers who work in CI pipelines, containerized environments, or remote development setups.

Output control options worth knowing

When many tests fail at once, default traceback output can be too verbose. Shorter traceback modes help you scan the failure landscape faster.

This is useful when:

  • upgrading Python
  • upgrading a dependency
  • refactoring a shared module
  • triaging a cascade of failures

Rerun smarter with failure-focused options

Bruhin highlights several powerful workflows:

  • rerun only last failures
  • run failures first
  • stop after the first failure
  • use stepwise progression to fix one failure at a time

These options reduce feedback time and mental overload. Instead of reprocessing the entire suite repeatedly, you work through failures in a controlled sequence.

That matters in real engineering teams, where test runs may take minutes rather than seconds.

Use --setup-show to understand fixture behavior

When fixture interactions become hard to reason about, --setup-show makes hidden lifecycle behavior visible.

This is one of the best tools for debugging:

  • fixture ordering
  • scope surprises
  • unexpected teardown timing
  • implicit dependencies

It is especially helpful when onboarding into an existing codebase with heavy fixture usage.

Measure slow tests with --durations

If a suite feels slow, don’t guess. Measure.

--durations helps identify expensive tests or phases, broken down by setup, call, and teardown. That often reveals whether the real problem is:

  • environment setup
  • the assertion path itself
  • cleanup logic

For performance-sensitive teams, this is a more disciplined way to prioritize optimization work.

Handle hanging tests proactively

A hanging test is one of the most frustrating failure modes in CI. Bruhin recommends timeout-related tooling so stalled runs produce meaningful tracebacks rather than just dead air.

That advice is deeply practical. In distributed systems, integration tests, or async workflows, hangs can be more common than outright exceptions. Better timeout diagnostics reduce wasted investigation time.

Plugins Extend Pytest Further Than Most Teams Realize

Bruhin briefly notes that the pytest ecosystem is enormous, with a large number of plugins available. The takeaway is not "install more plugins." It’s this:

If your team has a recurring testing pain point, there may already be a tested extension for it.

Examples mentioned include:

  • rerunning flaky failures
  • subtest support in older versions
  • richer reporting options

But there’s also a deeper point: conftest.py is effectively a local plugin. Many developers are already extending pytest without thinking of it that way.

That’s useful because it reframes test architecture. Instead of stuffing helpers into ad hoc files, teams can build shared testing behavior in structured, framework-native ways.

What This Means for Data and AI Engineers

Although the talk is about Python testing broadly, several lessons map directly to data engineering and AI engineering workflows.

1. Precision beats vague correctness

Data systems often fail at the edges:

  • invalid records
  • null handling
  • parser ambiguity
  • schema incompatibility
  • rounding and tolerance issues

Broad assertions miss those failures. Precise exception matching and deliberate parameterization catch them.

2. State isolation is critical

Shared state problems are common in:

  • notebook-driven experimentation
  • pipeline integration tests
  • feature stores
  • model-serving validation
  • local dev environments

Fixture scoping and cleanup strategy matter just as much in these contexts as in application development.

3. CI signal quality is a career skill

As you move into more senior technical roles, people will judge not just your code, but your systems thinking. A well-configured test suite that fails on warnings, surfaces flaky behavior honestly, and supports fast triage is evidence of engineering maturity.

4. Test readability affects team velocity

Clear parameter IDs, type-annotated fixtures, and explicit metadata may seem small, but they make a codebase easier to maintain across a team. That matters when projects move from solo builds to shared ownership.

A Practical Adoption Plan

If your current pytest usage is mostly basic, don’t try to adopt everything at once. A better sequence is:

Phase 1: Improve correctness

  • replace broad exception assertions with match=... where appropriate
  • switch float equality assertions to pytest.approx
  • start using explicit custom IDs in important parameterized tests

Phase 2: Improve safety

  • configure warnings as errors
  • enable strict pytest behavior
  • review current uses of skip and xfail

Phase 3: Improve maintainability

  • add type hints and docstrings to fixtures
  • refactor complex parameter sets into data classes
  • separate expensive setup from per-test reset logic

Phase 4: Improve debugging speed

  • learn --lf, --ff, --stepwise, --setup-show, and --durations
  • add timeout safeguards for hanging tests
  • identify where plugin support would reduce recurring pain

Final Thoughts

The best testing advice is rarely about writing more tests. It is about writing tests that produce better information.

That is the thread connecting the techniques in Bruhin’s session. Advanced pytest usage is not about clever tricks for their own sake. It is about making failures more meaningful, making state management more intentional, and making the test suite a more dependable engineering instrument.

For professionals aiming to grow into more advanced backend, platform, data, or AI roles, this is exactly the kind of skill that compounds. Libraries change. Frameworks evolve. But the ability to design trustworthy tests remains a durable advantage.

A team can tolerate imperfect code for a while. It struggles much more when it cannot trust the system meant to verify that code. pytest, used well, helps close that gap.

Source: "Building reliable data pipelines with polars and dataframely [PyCon DE & PyData 2026]" - PyData, YouTube, Aug 4, 2026 - https://www.youtube.com/watch?v=08tyYLgfaBg