Switches the type system to the single-family Geist spec referenced
in ``Business/DataTools/geist_spec.md`` and the matching
``datatools_layout_redesign2.html`` mockup. Editorial-serif headings
are out; the product now reads as modern SaaS-tool typography per
the spec's positioning note (§10).
src/gui/theme.py (new)
Implements geist_spec.md §3 verbatim — preconnect + Google Fonts
link for Geist (400/500/600/700) and Geist Mono (400/500), the
canonical ``:root`` token table (§7) plus severity extensions,
and the type scale (§4): h1 32/600/-0.035em, h2 22/600/-0.025em,
h3 18/500/-0.018em, h4 15/500/-0.012em, body 14/400, caption
12.5/400, mono 0.92× ss02. ``apply_theme()`` is the single entry
point.
Two deviations from the spec, both anticipated by spec §6.1:
- ``font-family: var(--font-sans) !important`` on the base rule.
Streamlit applies ``font-family: "Source Sans"`` directly to
``[data-testid="stMarkdownContainer"]`` and a few widget
wrappers at equal-or-higher specificity than the spec's
selector list, so plain inheritance loses the cascade.
- The base selector list explicitly enumerates
``stSidebarNav``, ``stMarkdownContainer``, ``stVerticalBlock``
and a few siblings so Streamlit's per-widget font reset
doesn't reach descendant text.
src/gui/components/_legacy.py
- ``_DESIGN_TOKENS_CSS`` no longer redeclares fonts or the
heading rules — those are theme.py's job (spec §9 says the
spec is type-only; everything below is component chrome).
- Token references switched from ``--dt-*`` to the spec names
(``--ink``, ``--bg``, ``--surface``, ``--border``, ``--accent``,
``--font-sans``, ``--font-mono``, …).
- Sidebar section-label rule tightened to 11.5px / 500 to match
the "Eyebrow" row in spec §4.
- Primary-button text color now also targets every descendant
(``button[kind="primary"] *``) so the inner
``stMarkdownContainer > p`` doesn't pick up
``color: var(--ink)`` from the base rule and render
near-invisible ink-on-ink.
- ``hide_streamlit_chrome`` now calls ``apply_theme`` before
injecting component CSS so the base tokens are defined first.
Acceptance criteria from spec §8 verified at 1920×1050:
- h1 computes ``font-family: Geist``, ``font-weight: 600``,
``letter-spacing: -1.12px`` (= 32px × -0.035em), size ``32px``.
- Body ``<p>`` inside ``stMarkdownContainer``: Geist 400 / 14px.
- Caption: Geist 400 / 12.5px.
- Inline mono filenames: Geist Mono in accent-fill chip.
- No Source Sans Pro leaks into any text the user reads.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
🌐 Language: English · Español
DataTools
Local CSV / Excel cleaning. CLI + browser GUI, no cloud, no install ceremony. GUI ships with English and Spanish language packs.
Tools
| # | Tool | Status |
|---|---|---|
| 01 | Find Duplicates — exact + fuzzy match, 5 normalizers, survivor rules, audit | Ready |
| 02 | Clean Text — whitespace, smart chars, BOM, line endings, case ops | Ready |
| 03 | Standardize Formats — dates, phones, emails, addresses, names, currencies, booleans | Ready |
| 04 | Fix Missing Values — disguised-null detection, profile, mean/median/mode/ffill/bfill/interpolate, drop strategies | Ready |
| 05 | Map Columns — fuzzy auto-rename, target schema with type coercion, required fields with defaults, drop/reorder | Ready |
| 06 | Find Unusual Values | Coming Soon |
| 07 | Combine Files | Coming Soon |
| 08 | Quality Check | Coming Soon |
| 09 | Automated Workflows — chain tools with recommended (not forced) order, save/load JSON, automate weekly cleanups | Ready |
Download (non-technical users)
Pre-built installers — no Python required:
| Platform | Download | First-launch note |
|---|---|---|
| macOS | DataTools-X.Y.Z-mac.dmg |
Drag DataTools.app into /Applications, then double-click. |
| Windows | DataTools-X.Y.Z-win-setup.exe |
Run the installer; launches from Start Menu. |
| Linux | DataTools-X.Y.Z-linux-x86_64.AppImage |
chmod +x the file, then double-click. |
Latest release: see GitHub Releases (or the Gumroad listing). The installers are ~150–200 MB; the launcher boots a local server at http://127.0.0.1:8501 and opens your browser. Nothing is sent to the cloud.
Install from source (developers)
pip install -r requirements.txt
Python 3.10+ required.
Run
GUI (recommended):
streamlit run src/gui/app.py
CLI — seven entry points:
python -m src.cli customers.csv [--apply] # dedup
python -m src.cli_text_clean messy.csv [--apply] # text clean
python -m src.cli_format intl.csv [--apply] # format standardize (auto-streams >100 MB)
python -m src.cli_missing holes.csv [--apply] # missing values
python -m src.cli_column_map vendor.csv [--apply] # column mapper
python -m src.cli_pipeline any_file.csv [--apply] # chain tools end-to-end
python -m src.cli_analyze any_file.csv [--json] # scan only
Every CLI runs preview-only by default; add --apply to write output.
Language
The GUI sidebar has a language picker. Packs ship for English and Español (src/i18n/packs/); the choice persists for the session. Adding a language: drop a <code>.json next to en.json mirroring its key tree, then list it in LANGUAGES. See Developer Guide §i18n.
Review & Normalize gate
Every uploaded file passes through a CSV-normalization gate before any tool sees it. The analyzer flags ~15 issue types (whitespace, NBSP / zero-width chars, BOM, encoding, smart punct, dirty headers, null sentinels, mojibake, …) tagged by confidence (high / medium / low) and fix action. The GUI shows each finding with Auto-fix / Skip / Customize, a live before/after preview, and an encoding-override picker. Tool pages refuse to load until the gate passes.
Output
Every run writes:
{input}_<tool>.csv— the cleaned data{input}_changes.csv(text cleaner) or{input}_match_groups.csv(dedup) — audit traillogs/<tool>_YYYYMMDD_HHMMSS.log— debug-level run log
Original input file is never modified.
Docs
- User Guide — install, GUI workflow, gate
- CLI Reference — every flag with recipes
- Requirements — file sizes, encodings, detectors, perf targets
- Technical — architecture, gate internals, fix registry
- Developer Guide — adding fixes / detectors / standardizers
Dependencies
pandas, openpyxl, rapidfuzz, phonenumbers, typer, loguru, charset-normalizer, streamlit. Optional: ftfy for mojibake repair.
License
Proprietary.