Upgrading to Filters v4

Filters v4 makes every filter generic over its output type, so a type checker can follow a value all the way through a chain. That part is free — your own filters need no changes. But three changes can break code that worked in Filters v3, and one of them fails silently.

Note

This guide covers the v4 alpha. Runtime behaviour is settled; the typing surface may still shift before 4.0.0 final.

Installing the alpha

Pre-release versions aren’t selected by default, so ask for this one explicitly:

pip install --pre 'phx-filters==4.0.0a1'

Or, with uv:

uv add 'phx-filters==4.0.0a1'

If the alpha bites, going back is just as explicit:

pip install 'phx-filters<4'

Note

Filters v4 adds typing-extensions>=4.15.0 as a runtime dependency. If you vendor your dependencies or install from a private mirror, add it before you upgrade.

Note

The phx-filters[django] and phx-filters[iso] extras declare no upper bound on phx-filters, so both install alongside v4 without a resolver conflict, and neither looks likely to trip the breaking changes below. Neither has been exercised against v4 in anger yet, though, so test in a throwaway environment before you rely on it.

At a Glance

Ordered by how likely each is to affect your code, not by how loud the failure is — the last one gives no warning at all, but tripping it takes a subclass check on these specific filters, which most code never does:

Change

How it surfaces

How to find it

Chaining with None

TypeError when the chain is built. For chains defined at module scope, that’s on import — but only if your test suite imports that module.

A type checker flags every one immediately; if you don’t have one, consider adopting one. Otherwise, rg -U 'f\.\w+(?:\([^)]*\))?\s*\|\s*None' catches chains built off the conventional import filters as f alias, though not one assembled through an intermediate variable.

Split with an empty keys

ValueError when the filter is constructed. For one built at module scope, that’s on import.

rg -U 'Split\([^)]*,' — a candidate list to eyeball; check each hit for a keys argument

ByteString and Date are no longer subclasses

Nothing. No exception, no warning — a check silently returns False and your code takes the other branch.

rg -U '(?s)(isinstance|issubclass)\(.{0,120}?\b(ByteString|Date|Datetime|Unicode)\b'

Chaining with None

Important

In Filters v3, some_filter | None silently did nothing. In Filters v4 it raises TypeError:

TypeError: None is not compatible with Int in a filter chain; use NoOp
instead, or Optional[Int] in a type annotation.

The silent no-op hid a common typo, and it had no sensible type: a chain whose next link might be absent can’t be checked. Use filters.NoOp explicitly instead.

The TypeError fires when the chain is constructed, not when it runs, so a chain built at module scope raises on import.

Filters v3:

chain = f.Unicode | None | f.Strip

Filters v4:

chain = f.Unicode | f.NoOp | f.Strip

This most often shows up when a chain is assembled from parts that may legitimately be absent:

Filters v3:

def build_chain(extra=None):
    return f.Unicode | extra

Filters v4:

def build_chain(extra=None):
    return f.Unicode | (extra if extra is not None else f.NoOp)

Note

Only the | operator is affected. filters.BaseFilter.resolve_filter() and FilterCompatible still accept None, so a filter that accepts None from its own caller keeps working.

None on the left — None | f.Int() — raised TypeError in Filters v3 as well, and is unchanged.

Tip

The Optional named in the error message is typing.Optional, used in a type annotation — not the filters.Optional filter.

Split with an empty keys

Important

f.Split(pattern, keys=[]) — an empty keys, not None — now raises ValueError immediately:

ValueError: keys must not be empty; pass None for list output instead.

Filters v3 branched on whether keys was truthy, so an empty keys fell through to the list branch and returned a list. Filters v4 branches on keys is not None, which is what the documented behaviour always described: an empty keys caps the split at zero items, and nothing fits — so rather than let every apply() call fail on that, it’s rejected up front, where the mistake was actually made.

Note

Why not just fall back to a list, the way Filters v3 did? The two constructor overloads dispatch on the type of keys — None versus Sequence[str] — not its value, and there is no way to type “a non-empty sequence” separately from “a sequence.” For a computed keys, a type checker can’t know whether it will be empty at runtime, so it infers Split[dict[str, str]] for any non-None keys regardless. Falling back to a list on empty input would make that inference silently wrong: code trusting the dict type would only fail later, at whatever call site used the result — further from the mistake, and with no exception pointing back to it.

The ValueError fires when the filter is constructed, not when it runs, so a filter built at module scope raises on import.

If you were relying on the old behaviour, pass None — the default:

Filters v3:

>>> f.Split(":", keys=[]).apply("a:b")
['a', 'b']

Filters v4:

>>> f.Split(":").apply("a:b")
['a', 'b']

Tip

This is most likely to bite where keys is computed rather than written out — keys=[k for k in fields if ...] that happens to select nothing. In Filters v3 that quietly returned a list; it now raises as soon as that Split is built, which is the point.

ByteString and Date are no longer subclasses

Important

filters.ByteString no longer subclasses filters.Unicode, and filters.Date no longer subclasses filters.Datetime. Each pair is now two siblings sharing a private base class:

>>> issubclass(f.ByteString, f.Unicode)
False
>>> isinstance(f.ByteString(), f.Unicode)
False

This is the one change nothing will tell you about. There is no exception and no warning: a check that used to be True is now False, and whatever branch depended on it quietly stops running. Search for ByteString and Date wherever you use isinstance() or issubclass() — both are affected, so searching for only one of them will miss cases.

The subclass relationship was never meaningful — a filters.ByteString emits bytes where a filters.Unicode emits str, so it could not stand in for its parent. Once each filter declared an output type, a type checker could see the violation.

If you were testing for the concrete filter, name it directly:

# Unchanged, and now means what it says.
isinstance(some_filter, f.ByteString)

If you were testing for “any decoder” or “any date-like filter”, there is no public replacement — the shared base classes are private. Test against the pair:

isinstance(some_filter, (f.Unicode, f.ByteString))
isinstance(some_filter, (f.Datetime, f.Date))

Note

Only the class hierarchy changed. Both filters accept the same input and produce the same output as they did in Filters v3, so code that simply uses them needs no changes.

Type Parameters

Now for the good news 😺

Chains and filters.FilterRunner infer real types instead of Any:

import filters as f

# Inferred as ``int``, not ``Any``.
f.FilterRunner(f.Int).cleaned_data

# Inferred as ``str``.
f.FilterRunner(f.Unicode | f.Strip | f.NotEmpty).cleaned_data

The library ships a py.typed marker, so you get this without installing a separate stubs package.

Note

Not every filter can declare an output type. A validator that passes its input through — filters.NotEmpty, filters.Required, filters.Min — leaves the chain’s type alone, and filters.Optional widens it to include its default’s type. But a filter whose output shape is only known at runtime still yields Any, so a chain ending in one infers nothing: filters.JsonDecode, filters.FilterMapper, filters.FilterRepeater, filters.FilterSwitch, filters.Array, filters.Item, filters.Pick and filters.Omit.

filters.Type, filters.NamedTuple, filters.Choice, filters.Round and filters.Call take their output type from their constructor arguments, so they infer only when you instantiate them.

Your own filters need no changes. The type parameter defaults to Any, so a bare subclass keeps working exactly as before:

# Still valid; behaves as it did in Filters v3.
class Pkcs7Pad(f.BaseFilter):
    ...

To opt into inference, declare what your filter emits:

# A chain ending in ``Pkcs7Pad`` now infers ``bytes``.
class Pkcs7Pad(f.BaseFilter[bytes]):
    ...

Tip

Some filters shouldn’t declare a concrete output type — a validator that passes its input through unchanged, for example. See Writing Your Own Filters for the base class to pick in each case.

Macros work differently. A filters.filter_macro() decorates a function, so there is no base class to parameterise — annotate its return type instead:

# mypy infers ``Any`` from this macro.
@f.filter_macro
def String():
    return f.Unicode | f.Strip

# mypy infers ``str``.
@f.filter_macro
def String() -> f.BaseFilter[str]:
    return f.Unicode | f.Strip

Note

pyright reads the function body and infers str either way; mypy reads only the annotation. Adding it costs nothing and satisfies both.

Important

Assigning a chain of bare classes to a variable defeats inference under mypy, which reads Unicode | Strip as a PEP 604 type alias rather than a chain:

# mypy infers ``types.UnionType``, and the chain's output degrades to
# ``Any``. There is no error — you just lose the type.
chain = f.Unicode | f.Strip | f.NotEmpty

Passing a chain straight into a filter, as in the examples above, is unaffected. Where you do want to name it, either annotate the variable or instantiate any one of the filters:

chain: f.FilterChain[str] = f.Unicode | f.Strip | f.NotEmpty

# Or, equivalently:
chain = f.Unicode() | f.Strip | f.NotEmpty

pyright infers FilterChain[str] in every one of these forms, and runtime behaviour is identical throughout.

While You’re Here

filters.FilterMapper now accepts sequence input. Where every key in filter_map is a non-bool int, a list or tuple is filtered by position and returned as a list:

>>> runner = f.FilterRunner(
...     f.Split(":") | f.FilterMapper({0: f.Unicode | f.Strip, 1: f.Int})
... )
>>> runner.apply("  name  :42")
>>> runner.cleaned_data
['name', 42]

This lets filters.Split chain straight into per-position validation, with no list-to-dict filter in between. Such a mapper still accepts Mapping input as well, and one with any other key — a str, bool keys, or an empty filter_map — is unchanged.

Note

A filters.FilterMapper picks its output shape from the runtime value, so it can’t declare an output type: cleaned_data is Any here, whichever checker you run. The per-key chains inside it are still checked.

Known Limitations

  • Neither mypy nor pyright runs at full strictness yet, so some strict typing warnings or errors in your own filters go unreported. Full strictness is expected by the 4.0.0 final release; progress is tracked in issue #119.

  • There is no narrowing filter category yet, so filters.Required and NotEmpty(allow_none=False) do not narrow T | None to T. Tracked in issue #122.

  • The two checkers disagree in two places, both covered above: an unannotated macro, and a chain of bare classes assigned to a variable. pyright infers the real type in each; mypy needs the annotation.

If you hit something this guide doesn’t cover, post in the Filters issue tracker and I’ll have a look 🙂