Skip to content

pyreorder

pyreorder — AST-based structural sorter for Python source code.

Public API:

  • :func:sort_source — sort a source string in memory.
  • :func:would_change — check whether sorting would change a source string.
  • :class:Config / :func:load_config / :func:discover — configuration.

Modules:

  • cache –

    Content-hash cache for skipping already-sorted files.

  • classify –

    Classification of top-level libcst statements into section buckets.

  • cli –

    CLI module for pyreorder.

  • config –

    Configuration discovery and parsing for pyreorder.

  • migrate –

    Auto-migrate legacy pyreorder config and cache paths to the current scheme.

  • pipeline –

    Top-level orchestration: parse -> section reorder -> in-class sort -> render.

  • sorters –

    In-section sorting strategies.

  • transforms –

    Opt-in import transforms: inline-import hoisting and TYPE_CHECKING removal.

  • undersort –

    In-class method sorting (undersort semantics).

Classes:

  • Cache –

    On-disk content-hash cache keyed by config_signature:source_hash.

  • Config –

    Resolved pyreorder configuration.

  • SectionSorter –

    Reorder a module's top-level statements by configured section.

Functions:

  • discover –

    Walk up from start (cwd by default) to find a pyreorder config file.

  • hash_text –

    Stable hex digest of text.

  • load_config –

    Load a resolved :class:Config.

  • sort_source –

    Return source sorted according to cfg.

  • would_change –

    True if :func:sort_source would alter source.

Cache

Cache(path: Path | None, *, ttl_days: int = 30)

On-disk content-hash cache keyed by config_signature:source_hash.

The cache is loaded once at the start of a run and saved once at the end (if any new entries were recorded). Reads/writes are atomic on save (temp file + rename) to avoid corruption on crash.

Parameters:

  • path

    (Path | None) –

    Cache file location (None = in-memory/no-op).

  • ttl_days

    (int, default: 30 ) –

    Entries not seen within this many days are pruned on load. 0 disables TTL pruning.

Methods:

  • load –

    Read the cache file. Missing or corrupt files are treated as empty.

  • lookup –

    True when the file can be skipped (already in sorted state).

  • merge –

    Merge external cache entries (e.g. from parallel workers) into this cache.

  • record –

    Record (or refresh) the sorted-output hash for a source hash.

  • save –

    Atomically write the cache if it changed. No-op without a path.

Attributes:

  • path (Path | None) –

    Where this cache is stored on disk (None = in-memory/no-op).

Source code in src/pyreorder/cache.py
47
48
49
50
51
def __init__(self, path: Path | None, *, ttl_days: int = 30) -> None:
    self._path = path
    self._entries: dict[str, CacheEntry] = {}
    self._dirty = False
    self._ttl_days = ttl_days

path property

path: Path | None

Where this cache is stored on disk (None = in-memory/no-op).

load

load() -> None

Read the cache file. Missing or corrupt files are treated as empty.

Old-format caches (plain hash strings) are migrated to the new {hash, seen} format with seen set to now. Entries whose seen timestamp is older than ttl_days days are dropped.

Source code in src/pyreorder/cache.py
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
def load(self) -> None:
    """Read the cache file. Missing or corrupt files are treated as empty.

    Old-format caches (plain hash strings) are migrated to the new
    ``{hash, seen}`` format with ``seen`` set to *now*.  Entries whose
    ``seen`` timestamp is older than ``ttl_days`` days are dropped.
    """
    if self._path is None or not self._path.exists():
        self._entries = {}
        return
    try:
        data = json.loads(self._path.read_text(encoding="utf-8"))
        if not isinstance(data, dict):
            self._entries = {}
            return
        now = time.time()
        migrated: dict[str, CacheEntry] = {}
        for k, v in data.items():
            if not isinstance(k, str):
                continue
            if isinstance(v, str):
                # Old format: plain hash string -> migrate.
                migrated[k] = {"hash": v, "seen": now}
            elif isinstance(v, dict) and "hash" in v and "seen" in v:
                migrated[k] = v
        self._entries = migrated
    except (json.JSONDecodeError, OSError):
        # Corrupt cache: silently start fresh. Better a few redundant sorts
        # than a crashed run.
        self._entries = {}
        return
    # Prune expired entries.
    if self._ttl_days > 0:
        cutoff = time.time() - self._ttl_days * 86400
        before = len(self._entries)
        self._entries = {k: v for k, v in self._entries.items() if v.get("seen", 0) >= cutoff}
        if len(self._entries) != before:
            self._dirty = True

lookup

lookup(config_sig: str, source_hash: str, current_content_hash: str) -> bool

True when the file can be skipped (already in sorted state).

A hit requires the stored sorted-output hash to match the file's current content hash -- otherwise the file was edited after caching and must sort. On a hit the seen timestamp is refreshed.

Source code in src/pyreorder/cache.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
def lookup(self, config_sig: str, source_hash: str, current_content_hash: str) -> bool:
    """True when the file can be skipped (already in sorted state).

    A hit requires the stored sorted-output hash to match the file's current
    content hash -- otherwise the file was edited after caching and must sort.
    On a hit the ``seen`` timestamp is refreshed.
    """
    key = f"{config_sig}:{source_hash}"
    entry = self._entries.get(key)
    if entry is None or entry.get("hash") != current_content_hash:
        return False
    # Refresh last-seen timestamp.
    entry["seen"] = time.time()
    self._dirty = True
    return True

merge

merge(entries: dict[str, CacheEntry]) -> None

Merge external cache entries (e.g. from parallel workers) into this cache.

Source code in src/pyreorder/cache.py
132
133
134
135
136
137
def merge(self, entries: dict[str, CacheEntry]) -> None:
    """Merge external cache entries (e.g. from parallel workers) into this cache."""
    for key, value in entries.items():
        if self._entries.get(key) != value:
            self._entries[key] = value
            self._dirty = True

record

record(config_sig: str, source_hash: str, sorted_hash: str) -> None

Record (or refresh) the sorted-output hash for a source hash.

Source code in src/pyreorder/cache.py
113
114
115
116
117
118
119
def record(self, config_sig: str, source_hash: str, sorted_hash: str) -> None:
    """Record (or refresh) the sorted-output hash for a source hash."""
    key = f"{config_sig}:{source_hash}"
    existing = self._entries.get(key)
    if existing is None or existing.get("hash") != sorted_hash:
        self._entries[key] = {"hash": sorted_hash, "seen": time.time()}
        self._dirty = True

save

save() -> None

Atomically write the cache if it changed. No-op without a path.

Source code in src/pyreorder/cache.py
121
122
123
124
125
126
127
128
129
130
def save(self) -> None:
    """Atomically write the cache if it changed. No-op without a path."""
    if self._path is None or not self._dirty:
        return
    self._path.parent.mkdir(parents=True, exist_ok=True)
    tmp = self._path.with_suffix(self._path.suffix + ".tmp")
    payload: dict[str, CacheEntry] = self._entries
    tmp.write_text(json.dumps(payload), encoding="utf-8")
    tmp.replace(self._path)
    self._dirty = False

Config dataclass

Config(sections: list[str] = (lambda: list(DEFAULT_SECTIONS))(), strategies: dict[str, SectionStrategy] = dict(), class_methods_enabled: bool = True, class_methods_order: list[str] = (lambda: ['public', 'protected', 'private'])(), class_methods_type_order: list[str] = (lambda: ['instance', 'class', 'static'])(), constants_pattern: str = '^[A-Z_][A-Z0-9_]*$', dunder_exports_names: list[str] = (lambda: ['__all__'])(), unknown_section: str = 'other', hoist_inline_imports: bool = False, hoist_main_imports: bool = True, remove_type_checking: bool = False, fail_on_changed: bool = True, exclude: list[str] = list(), recursive: bool = True, cache_enabled: bool = True, cache_dir: Path | None = None, cache_ttl_days: int = 30, jobs: int = 0, parallel_backend: str = 'process', config_path: Path | None = None)

Resolved pyreorder configuration.

Attributes:

  • sections (list[str]) –

    Ordered list of section buckets; top-level statements are grouped into these in the order given.

  • strategies (dict[str, SectionStrategy]) –

    Per-section in-section strategy. Sections missing from this mapping default to "keep".

  • class_methods_enabled (bool) –

    Reorder methods within each class (undersort).

  • class_methods_order (list[str]) –

    Method visibility ordering.

  • class_methods_type_order (list[str]) –

    Method-type ordering within each visibility.

  • constants_pattern (str) –

    Regex for the module_constants classification.

  • dunder_exports_names (list[str]) –

    Dunder names (__all__, __version__, etc.) that classify into dunder_exports — a section placed after functions and before main_block — rather than module_constants.

  • unknown_section (str) –

    Bucket name for unrecognised top-level nodes. Nodes in a bucket that is absent from sections are appended at the end, preserving their original relative order.

  • hoist_inline_imports (bool) –

    Move imports nested inside function/class bodies to the top of the module (pre-pass before section sorting).

  • remove_type_checking (bool) –

    Delete if TYPE_CHECKING: guards and hoist the imports they contained to the top of the module (pre-pass).

  • fail_on_changed (bool) –

    When True (default), pyreorder run exits with code 1 when any file was modified. Pre-commit/CI friendly. Set to False via [cli] fail_on_changed = false or --no-fail.

  • exclude (list[str]) –

    Glob patterns to exclude during file discovery (merged with --exclude flags).

  • recursive (bool) –

    Whether file discovery descends into subdirectories (default True). --no-recursive overrides per-invocation.

  • cache_enabled (bool) –

    When True (default), skip parsing/sorting for files whose content hash matches a cached sorted-output hash. Disable via [cli] cache = false or --no-cache.

  • cache_dir (Path | None) –

    Directory storing the content-hash cache. Defaults to ~/.cache/pyreorder/<project-slug> (or ./.pyreorder-cache/ when configured). Overrides via [cli] cache_dir.

  • jobs (int) –

    Number of parallel workers (0 = serial, negative = auto-detect int(0.75 * cpu_count())). Configurable via [cli] jobs or --jobs/-j.

  • parallel_backend (str) –

    Parallel execution backend — "process" (multiprocessing, default) or "thread" (threading). Configurable via [cli] parallel_backend or --parallel-backend.

  • config_path (Path | None) –

    Where the config was loaded from (None = pure defaults).

Methods:

  • config_signature –

    Stable hash of the output-affecting fields + pyreorder version.

  • from_table –

    Build a :class:Config from a parsed pyreorder table.

  • strategy –

    In-section strategy for section ("keep" if unset).

config_signature

config_signature() -> str

Stable hash of the output-affecting fields + pyreorder version.

Used as part of the content-hash cache key. Changes to any field that affects sorted output, or to the pyreorder version, invalidate the cache. Non-output fields (fail_on_changed, exclude, recursive, unknown_section, cache_*, config_path) are excluded.

Source code in src/pyreorder/config.py
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
def config_signature(self) -> str:
    """Stable hash of the output-affecting fields + pyreorder version.

    Used as part of the content-hash cache key. Changes to any field that
    affects sorted output, or to the pyreorder version, invalidate the cache.
    Non-output fields (``fail_on_changed``, ``exclude``, ``recursive``,
    ``unknown_section``, ``cache_*``, ``config_path``) are excluded.
    """
    try:
        ver = _pkg_version("pyreorder")
    except PackageNotFoundError:
        ver = "0.0.0"
    relevant = {
        "sections": self.sections,
        "strategies": self.strategies,
        "class_methods_enabled": self.class_methods_enabled,
        "class_methods_order": self.class_methods_order,
        "class_methods_type_order": self.class_methods_type_order,
        "constants_pattern": self.constants_pattern,
        "dunder_exports_names": self.dunder_exports_names,
        "hoist_inline_imports": self.hoist_inline_imports,
        "hoist_main_imports": self.hoist_main_imports,
        "remove_type_checking": self.remove_type_checking,
        "version": ver,
    }
    payload = _json_dumps(relevant, sort_keys=True).encode()
    return sha256(payload).hexdigest()[:16]

from_table classmethod

from_table(data: dict[str, Any], *, path: Path | None = None, raw: dict[str, Any] | None = None, is_pyproject: bool = False) -> Config

Build a :class:Config from a parsed pyreorder table.

Source code in src/pyreorder/config.py
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
@classmethod
def from_table(  # noqa: PLR0915
    cls,
    data: dict[str, Any],
    *,
    path: Path | None = None,
    raw: dict[str, Any] | None = None,
    is_pyproject: bool = False,
) -> Config:
    """Build a :class:`Config` from a parsed ``pyreorder`` table."""
    cfg = cls(config_path=path)

    module = data.get("module", {}) or {}
    if isinstance(module, dict) and "sections" in module:
        sections = module["sections"]
        if isinstance(sections, list) and all(isinstance(s, str) for s in sections):
            cfg.sections = [s for s in sections if isinstance(s, str)]

    strategy = data.get("strategy", {}) or {}
    if isinstance(strategy, dict):
        for name, value in strategy.items():
            if not isinstance(name, str):
                continue
            if value in VALID_STRATEGIES:
                cfg.strategies[name] = value
            else:
                warnings.warn(
                    f"pyreorder: unknown strategy {value!r} for section {name!r}; ignoring",
                    stacklevel=2,
                )

    cm = data.get("class_methods", {}) or {}
    if isinstance(cm, dict) and cm:
        cfg._apply_class_methods(cm)
    elif raw is not None:
        cfg._apply_class_methods(_legacy_undersort_overrides(raw, is_pyproject=is_pyproject))

    classification = data.get("classification", {}) or {}
    if isinstance(classification, dict):
        if "constants_pattern" in classification:
            pattern = classification["constants_pattern"]
            if isinstance(pattern, str):
                cfg.constants_pattern = pattern
        if "dunder_exports_names" in classification:
            names = classification["dunder_exports_names"]
            if isinstance(names, list) and all(isinstance(n, str) for n in names):
                cfg.dunder_exports_names = list(names)
            else:
                warnings.warn(
                    "pyreorder: classification.dunder_exports_names must be a list of strings; ignoring",
                    stacklevel=2,
                )

    transforms = data.get("transforms", {}) or {}
    if isinstance(transforms, dict):
        if "hoist_inline_imports" in transforms:
            cfg.hoist_inline_imports = bool(transforms["hoist_inline_imports"])
        if "hoist_main_imports" in transforms:
            cfg.hoist_main_imports = bool(transforms["hoist_main_imports"])
        if "remove_type_checking" in transforms:
            cfg.remove_type_checking = bool(transforms["remove_type_checking"])

    cli = data.get("cli", {}) or {}
    if isinstance(cli, dict):
        if "fail_on_changed" in cli:
            cfg.fail_on_changed = bool(cli["fail_on_changed"])
        if "cache" in cli:
            cfg.cache_enabled = bool(cli["cache"])
        if "cache_dir" in cli:
            cache_dir = cli["cache_dir"]
            if isinstance(cache_dir, str):
                cfg.cache_dir = Path(cache_dir)
            else:
                warnings.warn(
                    "pyreorder: cli.cache_dir must be a string; ignoring",
                    stacklevel=2,
                )
        if "jobs" in cli:
            jobs = cli["jobs"]
            if isinstance(jobs, int):
                cfg.jobs = jobs
            else:
                warnings.warn(
                    "pyreorder: cli.jobs must be an integer; ignoring",
                    stacklevel=2,
                )
        if "parallel_backend" in cli:
            backend = cli["parallel_backend"]
            if isinstance(backend, str) and backend in ("process", "thread"):
                cfg.parallel_backend = backend
            else:
                warnings.warn(
                    "pyreorder: cli.parallel_backend must be 'process' or 'thread'; ignoring",
                    stacklevel=2,
                )

    if "cache_ttl_days" in cli:
        ttl = cli["cache_ttl_days"]
        if isinstance(ttl, int) and ttl >= 0:
            cfg.cache_ttl_days = ttl
        else:
            warnings.warn(
                "pyreorder: cli.cache_ttl_days must be a non-negative int; ignoring",
                stacklevel=2,
            )

    discovery = data.get("discovery", {}) or {}
    if isinstance(discovery, dict):
        if "exclude" in discovery:
            exclude = discovery["exclude"]
            if isinstance(exclude, list) and all(isinstance(p, str) for p in exclude):
                cfg.exclude = [str(p) for p in exclude]
            else:
                warnings.warn(
                    "pyreorder: discovery.exclude must be a list of strings; ignoring",
                    stacklevel=2,
                )
        if "recursive" in discovery:
            cfg.recursive = bool(discovery["recursive"])

    return cfg

strategy

strategy(section: str) -> SectionStrategy

In-section strategy for section ("keep" if unset).

Source code in src/pyreorder/config.py
129
130
131
def strategy(self, section: str) -> SectionStrategy:
    """In-section strategy for ``section`` (``"keep"`` if unset)."""
    return self.strategies.get(section, "keep")

SectionSorter

SectionSorter(cfg: Config)

Bases: CSTTransformer

Reorder a module's top-level statements by configured section.

Safety model: statements whose section is not listed in Config.sections (notably unrecognised runtime setup such as app = typer.Typer()) act as barriers and never move. Recognised statements only reorder within their contiguous barrier-free run, so pyreorder never moves code across a setup statement it might depend on. The module docstring and from __future__ imports are pinned at the top (Python requires __future__ first).

Source code in src/pyreorder/pipeline.py
30
31
32
def __init__(self, cfg: Config) -> None:
    self.cfg = cfg
    self._reorder_sections = set(cfg.sections)

discover

discover(start: Path | None = None) -> Path | None

Walk up from start (cwd by default) to find a pyreorder config file.

Source code in src/pyreorder/config.py
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
def discover(start: Path | None = None) -> Path | None:
    """Walk up from ``start`` (cwd by default) to find a pyreorder config file."""
    base = (start or Path.cwd()).resolve()
    for directory in [base, *base.parents]:
        for rel in PYREORDER_FILES:
            candidate = directory / rel
            if candidate.is_file():
                return candidate
        pyproject = directory / "pyproject.toml"
        if pyproject.is_file():
            try:
                data = _read_toml(pyproject)
            except (tomllib.TOMLDecodeError, OSError):
                continue
            tool = data.get("tool", {})
            if tool.get("pyreorder"):
                return pyproject
    return None

hash_text

hash_text(text: str) -> str

Stable hex digest of text.

Source code in src/pyreorder/cache.py
140
141
142
def hash_text(text: str) -> str:
    """Stable hex digest of *text*."""
    return sha256(text.encode("utf-8")).hexdigest()

load_config

load_config(*, explicit: Path | None = None, start: Path | None = None) -> Config

Load a resolved :class:Config.

explicit overrides discovery. start is the directory to search from (defaults to cwd) and is ignored when explicit is given.

Before discovery, :func:pyreorder.migrate.migrate_if_needed is called to rename any legacy config / cache paths (e.g. csort.toml -> pyreorder.toml). On a steady-state install with no legacy paths, the only cost is one stat() call on the migration sentinel.

Source code in src/pyreorder/config.py
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
def load(
    *,
    explicit: Path | None = None,
    start: Path | None = None,
) -> Config:
    """Load a resolved :class:`Config`.

    ``explicit`` overrides discovery. ``start`` is the directory to search from
    (defaults to cwd) and is ignored when ``explicit`` is given.

    Before discovery, :func:`pyreorder.migrate.migrate_if_needed` is called to
    rename any legacy config / cache paths (e.g. ``csort.toml`` ->
    ``pyreorder.toml``). On a steady-state install with no legacy paths, the
    only cost is one ``stat()`` call on the migration sentinel.
    """
    # Auto-migrate legacy config file / cache paths before discovery.
    migrate_if_needed()
    path = explicit or discover(start)
    if path is None:
        return Config()
    try:
        data = _read_toml(path)
    except tomllib.TOMLDecodeError as exc:
        warnings.warn(f"pyreorder: could not parse {path}: {exc}; using defaults", stacklevel=2)
        return Config(config_path=path)

    is_pyproject = path.name == "pyproject.toml"
    table = _pyreorder_table_from_data(data, is_pyproject=is_pyproject)
    cfg = Config.from_table(table, path=path, raw=data, is_pyproject=is_pyproject)

    return cfg

sort_source

sort_source(source: str, cfg: Config, *, filename: str = '<unknown>') -> str

Return source sorted according to cfg.

  1. parse with libcst;
  2. bail out untouched if the file is disabled (# pyreorder: off header);
  3. reorder top-level statements by section;
  4. (optional) reorder methods within each class.
Source code in src/pyreorder/pipeline.py
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
def sort_source(source: str, cfg: Config, *, filename: str = "<unknown>") -> str:
    """Return ``source`` sorted according to ``cfg``.

    1. parse with libcst;
    2. bail out untouched if the file is disabled (``# pyreorder: off`` header);
    3. reorder top-level statements by section;
    4. (optional) reorder methods within each class.
    """
    module = cst.parse_module(source)
    if undersort.file_disabled(module):
        return source
    module = transforms.apply_transforms(module, cfg)
    new_module = module.visit(SectionSorter(cfg))
    if cfg.class_methods_enabled:
        new_module = new_module.visit(undersort.MethodSorter(cfg.class_methods_order, cfg.class_methods_type_order))
    return new_module.code

would_change

would_change(source: str, cfg: Config, *, filename: str = '<unknown>') -> bool

True if :func:sort_source would alter source.

Source code in src/pyreorder/pipeline.py
165
166
167
def would_change(source: str, cfg: Config, *, filename: str = "<unknown>") -> bool:
    """True if :func:`sort_source` would alter ``source``."""
    return sort_source(source, cfg, filename=filename) != source