# Spellchecking The repository includes a spellchecker (`infra/pre_commit/spellcheck.py`) that scans the prose of the repository for misspelled words. It is backed by `pyspellchecker` and augmented with project-specific word lists. ## What is checked The checker covers `torx/`, `storx/`, `tests/`, `infra/`, `notebooks/`, `doc/`, plus `README.md` and `install.sh`. What counts as prose depends on the file type: | File type | Checked | |---|---| | `.py` | Comments (`# ...`), docstrings, and string constants | | `.md`, `.rst` | The prose, minus fenced code blocks and inline code spans | | `.sh`, `.yml`, `.toml`, `.in` | The `#` comments only | Identifiers are intentionally excluded: `ALL_CAPS`, `snake_case`, and `camelCase` tokens are discarded before checking, so variable and function names do not produce false positives. URLs, filenames, LaTeX math, and tokens containing digits are also filtered out. ## American English TorX uses American English. The dictionary is American, so most British spellings (`normalise`, `centre`, `behaviour`) are flagged as plain misspellings and suggest the American form. A handful of British spellings (such as `coloured` or `travelling`) are real words the checker recognizes on their own and will not catch, so review spelling by eye as well. ## Skipping a line Occasionally a line is genuinely not prose, such as a legacy identifier that has to keep its original spelling, or a deliberately lowercase example. Put the marker `spellcheck: ignore` in the line to exempt it from both checks: ```python # Legacy GRILLIX wrote 'params_penalisation'. spellcheck: ignore ``` Prefer this over adding a nonsense word to the dictionary, which would let the same mistake through everywhere else. ## Adding words to the dictionary The spellchecker merges two word lists at runtime: | File | Purpose | |---|---| | `infra/pre_commit/spelling_dict.txt` | Physics and domain-specific vocabulary - **edit this one** | | `infra/pre_commit/code_dict.txt` | Auto-generated from repository identifiers - do not edit manually | To permanently allow a word, add it (one per line) to `infra/pre_commit/spelling_dict.txt`. `infra/pre_commit/code_dict.txt` is regenerated from the codebase and any manual additions would be overwritten. Comments starting with `#` are supported and encouraged to keep the file organized. ### Capitalization requirements `infra/pre_commit/capitalize_dict.txt` lists words that must always appear capitalized in documentation (proper nouns and named quantities such as `Maxwellian`, `Boltzmann`, `Alfven`). Each entry should be the canonical capitalized form. Add entries here when a word is correct but must not appear in lowercase. ### Regenerating the identifier dictionary `infra/pre_commit/code_dict.txt` is not committed to the repository and must be generated before running the spellchecker for the first time. After large refactors, regenerate it to keep the word list current: ```bash python infra/pre_commit/extract_dict.py --root torx/ storx/ tests/ infra/ \ notebooks/ --out infra/pre_commit/code_dict.txt ``` This walks all `.py` files, splits `snake_case` and `CamelCase` names into components, and writes the result as a flat word list. ## Running manually From the repository root, inside the project environment: ```bash # Check only files changed relative to origin/main (fast, same as pre-commit) python infra/pre_commit/spellcheck.py torx storx tests infra notebooks doc \ README.md install.sh --extensions .py .md .sh .rst .yml .in --diff-only # Full scan of the whole repository (same as CI) python infra/pre_commit/spellcheck.py torx storx tests infra notebooks doc \ README.md install.sh --extensions .py .md .sh .rst .yml .in ``` `--extensions` defaults to `.py` if omitted, so pass it explicitly to also check markdown, rst, shell, yaml, and `.in` files. ## Running in your own environment If you are working in a personal environment where `pyspellchecker` is not already installed, install it first and then run the script directly: ```bash pip install pyspellchecker python infra/pre_commit/spellcheck.py torx storx tests infra notebooks doc \ README.md install.sh --extensions .py .md .sh .rst .yml .in --diff-only ``` No other dependencies are needed beyond the Python standard library. ## Pre-commit hook The `spellcheck` hook runs automatically on `git commit` for staged files matching `.py`, `.md`, `.sh`, `.rst`, `.yml`, or `.in`. It uses `--diff-only` so only changed files are checked. The hook is installed automatically when you create or reinstall the Python environment with `./install.sh`. If you are working in an existing environment where it has not been set up yet, install the dependencies and register the hook manually, after installing `pyspellchecker`: ```bash pip install pre-commit pre-commit install ``` ### Skipping the hook If you need to push to your own fork or a personal branch without the spellcheck passing (for example, while work is still in progress), you can bypass it with the `--no-verify` flag: ```bash git commit --no-verify -m "Your message" ``` To skip only the spellcheck hook but keep other hooks running, use the `SKIP` environment variable with the hook ID: ```bash SKIP=spellcheck git commit -m "Your message" ``` The hook ID (`spellcheck`) is defined in `.pre-commit-config.yaml`.