Go to file

Michael 00d3f28865 feat(pipeline): plain-English per-step result summaries

Replaces the raw-JSON summary column in the Results table with the mockup's
plain-English phrasing: "312 duplicates removed across 147 groups
(18,442 → 18,130 rows)", "1,204 cells cleaned in name & city", etc.
(correct singular/plural via a small _n helper).

Adds step_phrase() and step_status() to pipeline_modules.py. step_status
derives the status pill (✓ ok / ⚠ ok · N skipped / ✗ error / ⏭ skipped) and,
for warn/error steps (e.g. format_standardize unparseable cells, column_map
coercion failures / missing required targets), an inline detail callout
rendered directly below the results table — surfacing non-fatal issues in
context without a dedicated always-empty column.

Extends tests/gui/test_pipeline_builder.py with phrasing + status assertions.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

2026-06-22 18:21:17 +00:00

.github/workflows

build: drop the local Python release method, return to CI-only installer builds

2026-06-22 17:47:36 +00:00

.streamlit

feat(brand): rebrand to UNALOGIX DataTools + Clean. Normalize. Transform.

2026-05-19 01:45:38 +00:00

build

build: drop the local Python release method, return to CI-only installer builds

2026-06-22 17:47:36 +00:00

docs

build: drop the local Python release method, return to CI-only installer builds

2026-06-22 17:47:36 +00:00

landing

docs+code: rename tool labels everywhere

2026-05-16 19:50:09 +00:00

layout-review

refactor(layout-review): consolidate tool-header actions + align reconcile downloads

2026-06-08 16:50:25 +00:00

marketing

docs(i18n): document language packs across user, dev, and marketing docs

2026-05-13 15:16:24 +00:00

samples

feat: 3 new tools, format streaming, distribution-ready demo + landing pages

2026-05-01 22:31:26 +00:00

scripts

feat(license): datatools-admin CLI for the mint API

2026-05-14 00:47:01 +00:00

server

feat(server): Gumroad webhook receiver + Postmark email (PR 2)

2026-05-14 01:33:43 +00:00

src

feat(pipeline): plain-English per-step result summaries

2026-06-22 18:21:17 +00:00

test-cases

test(junk-corpus): pathological-input stress suite for the analyzer

2026-05-16 21:35:22 +00:00

tests

feat(pipeline): plain-English per-step result summaries

2026-06-22 18:21:17 +00:00

.gitignore

build: bundle Tesseract 5.5.0 + tessdata into every release artifact

2026-06-02 18:20:33 +00:00

DECISIONS.md

feat(gui): port journey-level nav + local-first pill to the live app

2026-06-08 17:01:57 +00:00

LICENSE_TESSERACT.txt

docs: reflect bundled Tesseract on every install surface

2026-06-02 18:20:50 +00:00

pytest.ini

fix: clear all latent deprecation + resource warnings

2026-05-13 16:28:48 +00:00

README.es.md

build: drop the local Python release method, return to CI-only installer builds

2026-06-22 17:47:36 +00:00

README.md

build: drop the local Python release method, return to CI-only installer builds

2026-06-22 17:47:36 +00:00

requirements-dev.txt

build(pdf): bundle PDF deps in installers + pin versions + smoke tests

2026-05-19 23:10:43 +00:00

requirements.txt

refactor(pdf): rip out templates; heuristic scan + selectable table

2026-05-19 23:57:30 +00:00

run_tests.py

feat(gate): CSV-normalization gate with confidence-tiered findings

2026-04-29 20:35:27 +00:00

streamlit_app.py

feat: 3 new tools, format streaming, distribution-ready demo + landing pages

2026-05-01 22:31:26 +00:00

tox.ini

test: single-command runner, cross-platform automation, fixture auto-discovery

2026-04-29 16:01:06 +00:00

README.md

🌐 Language: English · Español

DataTools

Local CSV / Excel cleaning. CLI + browser GUI, no cloud, no install ceremony. GUI ships with English and Spanish language packs.

Tools

#	Tool	Status
01	Find Duplicates — exact + fuzzy match, 5 normalizers, survivor rules, audit	Ready
02	Clean Text — whitespace, smart chars, BOM, line endings, case ops	Ready
03	Standardize Formats — dates, phones, emails, addresses, names, currencies, booleans	Ready
04	Fix Missing Values — disguised-null detection, profile, mean/median/mode/ffill/bfill/interpolate, drop strategies	Ready
05	Map Columns — fuzzy auto-rename, target schema with type coercion, required fields with defaults, drop/reorder	Ready
06	Find Unusual Values	Coming Soon
07	Combine Files	Coming Soon
08	Quality Check	Coming Soon
09	Automated Workflows — chain tools with recommended (not forced) order, save/load JSON, automate weekly cleanups	Ready

Every tool page has an in-tool Help popover (right of the title) with a compact When-to-use / Steps / Examples / Tip card. Copy lives in the language packs (tools.<id>.help_md).

Download (non-technical users)

Pre-built bundles — no Python install, no admin rights, no internet at runtime. Each release ships an installer per OS that wires up Desktop + Start Menu / Launchpad shortcuts.

Platform	Installer
macOS	`DataTools-X.Y.Z-mac.dmg` — open, drag DataTools.app into /Applications, launch from Launchpad.
Windows	`DataTools-X.Y.Z-win-setup.exe` — run installer (per-user, no admin). Desktop shortcut + Start Menu entry created.
Linux	`DataTools-X.Y.Z-linux-x86_64.AppImage` — `chmod +x`, double-click. The AppImage is already portable.

Latest release: see GitHub Releases (or the Gumroad listing). Each bundle is ~300 MB unpacked; on first launch the app starts a local server at http://127.0.0.1:8501 and opens your default browser. Nothing leaves your machine.

Tesseract OCR is bundled. Scanned-PDF support in the PDF Extractor works out of the box on all three platforms — no separate Tesseract install required. License attribution: see LICENSE_TESSERACT.txt.

First-launch warnings (one-time):

macOS unsigned builds: right-click → Open → confirm. (Signed builds skip this.)
Windows SmartScreen: click More info → Run anyway.

Detailed install + troubleshooting walkthrough: User Guide §1.

Install from source (developers)

pip install -r requirements.txt

Python 3.10+ required.

Run

GUI (recommended):

streamlit run src/gui/app.py

CLI — seven entry points:

python -m src.cli            customers.csv [--apply]   # dedup
python -m src.cli_text_clean messy.csv     [--apply]   # text clean
python -m src.cli_format     intl.csv      [--apply]   # format standardize (auto-streams >100 MB)
python -m src.cli_missing    holes.csv     [--apply]   # missing values
python -m src.cli_column_map vendor.csv    [--apply]   # column mapper
python -m src.cli_pipeline   any_file.csv  [--apply]   # chain tools end-to-end
python -m src.cli_analyze    any_file.csv  [--json]    # scan only

Every CLI runs preview-only by default; add --apply to write output.

Language

The GUI sidebar has a language picker. Packs ship for English and Español (src/i18n/packs/); the choice persists for the session. Adding a language: drop a <code>.json next to en.json mirroring its key tree, then list it in LANGUAGES. See Developer Guide §i18n.

Review & Normalize gate

Every uploaded file passes through a CSV-normalization gate before any tool sees it. The analyzer flags ~15 issue types (whitespace, NBSP / zero-width chars, BOM, encoding, smart punct, dirty headers, null sentinels, mojibake, …) tagged by confidence (high / medium / low) and fix action. The GUI shows each finding with Auto-fix / Skip / Customize, a live before/after preview, and an encoding-override picker. Tool pages refuse to load until the gate passes.

Output

Every run writes:

{input}_<tool>.csv — the cleaned data
{input}_changes.csv (text cleaner) or {input}_match_groups.csv (dedup) — audit trail
logs/<tool>_YYYYMMDD_HHMMSS.log — debug-level run log

Original input file is never modified.

Docs

User Guide — install, GUI workflow, gate
CLI Reference — every flag with recipes
Requirements — file sizes, encodings, detectors, perf targets
Technical — architecture, gate internals, fix registry
Developer Guide — adding fixes / detectors / standardizers

Dependencies

pandas, openpyxl, rapidfuzz, phonenumbers, typer, loguru, charset-normalizer, streamlit. Optional: ftfy for mojibake repair.

License

Proprietary.

Languages

Python 87.3%

HTML 10%

CSS 1.8%

Shell 0.4%

JavaScript 0.2%

Other 0.2%