A linter for computer-vision datasets.
Point it at a folder of images and labels. It tells you what's broken,
duplicated, mislabeled, or suspicious, before it wastes a training run.
cvflow-demo-github.mp4
Getting started · Quickstart · Installation · Your first run · Dataset layout
Using it · The terminal report · The dashboard · Fixing boxes in the browser · Command reference · Common recipes
Reference · What it checks · Every finding, by code · In CI · From Python · Troubleshooting
Contributing · Testing & development · How it fits together · Design philosophy
Three commands. No config, no account, no setup.
pip install cvflow # 1. install
cvflow inspect # 2. no dataset? it fetches coco128 and checks that
cvflow inspect ./dataset --serve # 3. or point it at yours and open the dashboardThat's the whole tool. ./dataset is any folder holding a YOLO or COCO dataset;
CVFlow works out which one it is, and whether it holds boxes, polygons or
oriented boxes.
Step-by-step, from nothing
-
Check you have Python 3.9 or newer. In a terminal (Command Prompt on Windows, Terminal on macOS/Linux):
python --version
No Python? Install it from python.org/downloads.
-
Install CVFlow:
pip install cvflow
-
Try it. With no dataset of your own, just run it: CVFlow downloads and unzips the
coco128sample for you.cvflow inspect --serve
Already have a dataset? Point at its folder instead:
cvflow inspect ./my-dataset --serve -
Your browser opens
http://localhost:8000with the dashboard. PressCtrl+Cin the terminal when you're finished.
cvflow: command not found? Use python -m cvflow instead of cvflow.
Same tool, works even when your PATH isn't set up:
python -m cvflow inspect ./dataset --serveCVFlow needs Python 3.9 or newer and pulls in two runtime dependencies
(pyyaml and pillow).
pip install cvflow
cvflow --versionIf your shell reports cvflow: command not found, the console script isn't on
your PATH. Use the module form instead; it's the same program:
python -m cvflow inspect ./datasetTo work on CVFlow itself, install it from a checkout with the dev extras. See Testing & development.
git clone https://github.com/RizwanMunawar/cvflow
cd cvflow
pip install -e ".[dev]"Run inspect with no path and CVFlow downloads Ultralytics' coco128 sample
(about 7 MB), unpacks it into a cache directory, and inspects that, so you can
see exactly what the tool does before arranging any data of your own.
cvflow inspectThe sample is fetched once and reused on every later run. Set CVFLOW_CACHE to
choose where it lands; otherwise CVFlow uses the platform cache home:
%LOCALAPPDATA%\cvflow\cache on Windows, and $XDG_CACHE_HOME/cvflow (falling
back to ~/.cache/cvflow) everywhere else.
CVFLOW_CACHE=/data/cvflow-cache cvflow inspectNothing is downloaded when you pass a path of your own:
cvflow inspect ./datasetTwo rules of thumb are worth internalizing before you point it at real data:
- Structure and annotation checks run from the labels alone. They work even if the image files live somewhere else.
- Corrupt-image, duplicate, and leakage checks need the pixels. Point CVFlow at the folder that contains both the labels and the images. When the images can't be found, CVFlow says the image checks were skipped rather than inventing findings.
Reading and hashing every image takes minutes on a real dataset, so CVFlow
prints a note to stderr before it starts. Pass --no-images to skip the pixel
work entirely when you only want a fast structural pass.
CVFlow reads the two most common detection formats. The closer your folder is to one of the layouts below, the more it can audit. In particular it needs to find the actual image files on disk to check for corrupt images, duplicates, and split leakage.
An Ultralytics-style data.yaml next to mirrored images/ and labels/ folders:
dataset/
├── data.yaml
├── images/
│ ├── train/
│ │ ├── frame_0001.jpg
│ │ └── frame_0002.jpg
│ └── val/
│ └── frame_9001.jpg
└── labels/
├── train/
│ ├── frame_0001.txt # matches images/train/frame_0001.jpg
│ └── frame_0002.txt
└── val/
└── frame_9001.txt
data.yaml names the classes and points at each split:
path: . # optional; base for the paths below (relative to this file)
train: images/train
val: images/val
# test: images/test # optional
names: # a list works too: [person, helmet]
0: person
1: helmetA label file holds one box per line in normalized YOLO format: class id, then box center and size, each as a fraction of the image (0–1):
# class_id cx cy w h
0 0.512 0.437 0.104 0.216
1 0.300 0.300 0.050 0.080
A few things worth knowing:
- CVFlow finds a label by taking the image path and swapping
images/→labels/and the extension →.txt. Keep that mirroring intact. - An image with no label file, or an empty one, is treated as a background image (no objects); that's a WARNING to sanity-check, not an error.
- Supported image extensions:
.jpg .jpeg .png .bmp .webp .tif .tiff.
No data.yaml? CVFlow falls back to a plain images/ + labels/ pair. It
picks up train/, val/, test/ subfolders if present, otherwise treats
everything as one split. Class names come from classes.txt (one name per line)
if it exists, and are inferred from the label files otherwise.
Annotation JSON(s) under annotations/, images under images/:
dataset/
├── annotations/
│ ├── instances_train.json
│ └── instances_val.json
└── images/
├── train/
│ └── 000000000001.jpg
└── val/
└── 000000009001.jpg
Each JSON is standard COCO: images, annotations, and categories:
Notes:
- COCO
bboxis absolute pixels[x, y, width, height]with(x, y)at the top-left. CVFlow normalizes it using each image'swidth/height, so a YOLO box and a COCO box end up meaning the same thing. - The split is inferred from the JSON filename: anything containing
train,val, ortest. A single JSON with no such hint loads as one unnamed split. - To run the image-level checks (corrupt / duplicate / leakage), CVFlow needs the
pixels. It looks for each
file_nameunder the dataset root, thenimages/, thenimages/<split>/and<split>/. If it can't find them it still audits structure, annotations, and statistics, and tells you image checks were skipped rather than inventing false positives.
Rule of thumb: structure and annotation checks run from the labels alone; corrupt-image, duplicate, and leakage checks need the image files reachable from the dataset root. Point CVFlow at the folder that contains both.
The default output is a prioritized text report.
cvflow inspect ./datasetCVFlow Dataset Health
────────────────────────────────────────────────────
Format YOLO
Images 12,482
Annotations 12,103
Classes 8
Splits train, val
Health Summary
────────────────────────────────────────────────────
ERROR 17
WARNING 184
INFO 6
Most Important Problems
────────────────────────────────────────────────────
1. [ERROR] 17 images are unreadable or corrupt.
2. [ERROR] 21 bounding boxes extend outside the image boundaries.
3. [WARNING] 143 unusually small bounding boxes detected.
4. [WARNING] 127 highly similar image pairs between 'train' and 'val' (possible leakage).
5. [WARNING] Class 'helmet' represents only 0.4% of annotations.
It has six sections: the overview, dataset statistics (suppressed with
--no-stats), the health summary, the most important problems, a
findings-by-type breakdown, and per-issue detail for the top findings.
Severity is used honestly, and it's worth trusting:
| Severity | Means | Example |
|---|---|---|
ERROR |
Something is objectively broken. | A corrupt image, a box with negative area. |
WARNING |
Worth a look; CVFlow can't decide for you. | A tiny box, a near-duplicate pair, split leakage. |
INFO |
An observation, not a problem. | A rare class, a heavily imbalanced distribution. |
Statistical oddities and duplicates are never hard errors. They're candidates for review, and they're phrased that way on purpose.
The progress note goes to stderr and the report to stdout, so
cvflow inspect ./dataset > report.txt captures just the report. It's written
as UTF-8 even when redirected, so that's safe on Windows consoles too.
The exit code is 0 when nothing's wrong and non-zero when there are errors, so
cvflow inspect drops straight into CI. Add --strict to fail on warnings too.
See Exit codes.
The terminal report tells you what's wrong. The dashboard shows you, on one page, with the images attached, and lets you fix boxes without leaving it.
cvflow inspect ./dataset --servecoco128: 128 images, 929 annotations, 71 classes (YOLO object detection, no splits).
98 findings: 0 errors, 49 warnings, 49 info.
Open http://localhost:8000 for the detail: every finding, its image, and the fix.
Your browser opens on the page; Ctrl+C in the terminal stops the server.
No extra install. The charts are Chart.js and the type is Archivo, the Ultralytics brand face, both vendored inside the Python package and inlined into the page. No Node, no npm, no CDN:
pip install cvflowis the whole setup and the page works offline.
One page, four tabs, everything hover- and keyboard-readable, light and dark:
| Tab | What it answers |
|---|---|
| Overview | How healthy is this dataset? What should I fix first, and how much accuracy is it worth? |
| Classes | Which classes dominate? How long is the tail? |
| Geometry | How big are the boxes, what shape, how many per image, and where in the frame do they sit? |
| Findings | Every finding, filterable by severity, check, split or free text. |
Every tab is rendered once at load and then shown or hidden, so switching costs nothing.
| Option | Effect |
|---|---|
--serve |
Start the dashboard instead of printing the text report. |
--port N |
Bind this exact port. Omitted, CVFlow scans upward from 8000 for the first free one. |
--host HOST |
Interface to bind. Defaults to 127.0.0.1 (loopback only). |
--no-browser |
Print the URL instead of opening a browser. Use it over SSH, in a container, or in a tmux pane. |
--html FILE |
Write the same page to a self-contained file. Can be combined with --serve. |
# Loopback, opens your browser
cvflow inspect ./dataset --serve
# On a remote box: print the URL, bind a fixed port
cvflow inspect ./dataset --serve --no-browser --port 9000
# Reachable from your network (see the note below)
cvflow inspect ./dataset --serve --host 0.0.0.0 --port 9000The server holds one page in memory. There is no static root and no directory
listing: with the editor attached it answers the page plus four /api/
endpoints; without one it serves the page and 404s everything else. Binding
0.0.0.0 exposes both the page and, with it, read and write access to the labels
in your dataset folder to anyone who can reach that port. Do that on a trusted
network only.
- Hover any card for a plain-English explanation of what it shows.
- ⤢ blows a chart up full screen; ↓ saves it as PNG or as the JSON behind it.
- The rail button collapses the sidebar.
- Click any check or any image (a bar in Findings by type, or one in Images with the most findings) to filter the findings view to it.
- Findings render as cards: severity, the headline, why it was flagged, a thumbnail of the image, and the suggested next step. Filter by severity, check, split or free text, and sort by any of them.
- Arrow keys move between tabs when a tab has focus; Escape closes the open overlay, tooltip, or editor.
Accuracy headroom (on the Overview tab) is a rough estimate of the accuracy you could recover by fixing each class of problem, with its formula shown on screen. Tick items off as you fix them and the remaining headroom updates. It's labeled an estimate because it is one. Treat it as a way to rank what to fix first, not as a number to report.
Click any finding that points at an image (or any bar in Images with the most findings) and the photo opens with its boxes drawn on top. The box the finding is about is drawn in red; every other class keeps its own colour. Down the right side: what was flagged, why, the suggested next step, and one-click fixes (Clamp into frame, Delete this box, Reassign class, Add a box). You can also:
- drag a box to move it, drag its corner to resize,
- drag on empty canvas to draw a new one,
- change the class of the selected box, or delete it (
Delete/Backspace), - press
Escapeto close the editor, - Save labels to write the corrections straight back to the label file.
Coordinates are clamped to the frame and reordered on the way in, so a box dragged past an edge or drawn backwards still lands as valid data rather than as a new finding. Saving also updates the in-memory dataset, so the dashboard and any later save agree with what's now on disk.
Writing is allowed only for YOLO detection datasets served with --serve.
Everything else opens read-only, and the panel says which case you're in:
| Dataset | Editable? | Why |
|---|---|---|
YOLO, detect |
✅ Yes | One small text file per image, so a change is contained and easy to review in git. |
YOLO, segment or obb |
❌ Read-only | Those labels hold more than a box; rewriting a polygon as a rectangle would silently throw the mask away. |
| COCO, any task | ❌ Read-only | The annotation JSON is shared across the whole split and is other tooling's to own. |
Any dataset via --html |
❌ Read-only | The static file has no backend to write with. |
A save writes one label file: one line per box, class_id cx cy w h, normalized
and mirrored from the image path. Images that had no label file get one created
at that mirrored path.
Every path the editor touches (reading an image, resolving a label file, writing
one) is resolved and confined to the dataset root. The editor writes in place and
keeps no backup, which is the point: your labels live in git, one file per image,
so a bad edit is a git diff away from being seen and a git checkout away from
being undone. If your dataset isn't under version control, make a copy before an
editing session.
Prefer a file you can archive or attach to a PR? --html writes the same
self-contained page to disk, no server involved:
cvflow inspect ./dataset --html report.htmlNo server, no external requests, nothing to install to read it. It's the right
output for attaching to a PR, archiving with a dataset version, or uploading as
a CI artifact. The one thing the file can't do is edit: that needs the backend
--serve provides, so the annotation editor opens read-only.
cvflow inspect [path] [options]
path Dataset root. Omitted, CVFlow downloads and
inspects the coco128 sample.
-f, --format {yolo,coco} Force a format instead of auto-detecting.
--no-images Skip checks that read image bytes (much faster;
also skips corrupt/duplicate/leakage detection).
--no-stats Hide the dataset-statistics section.
--strict Exit non-zero on warnings, not just errors.
dashboard:
--serve Open the findings in a browser dashboard instead
of printing a text report.
--port N Port for --serve (default: first free from 8000).
--host HOST Interface for --serve (default: 127.0.0.1).
--no-browser With --serve, print the URL instead of opening it.
--html FILE Write the dashboard to a self-contained HTML file.
cvflow --version Print the version.
cvflow inspect --help Show this list in your terminal.
Notes on individual options:
--formatonly overrides detection. Reach for it when a folder looks like both formats, or when detection guesses wrong.--no-imagesturns off every check that opens an image file: corrupt images, exact and near duplicates, and split leakage. Structure, annotation geometry, and statistics still run.--strictchanges the exit code only. It doesn't add checks or change the report.--serveand--htmlcan be combined. The file is written first, then the server starts.--portis only a starting point when omitted: CVFlow scans upward from 8000 for the first free port. Given explicitly, that exact port is bound.
| Code | Meaning |
|---|---|
0 |
Clean: no errors (and no warnings under --strict). |
1 |
Problems found: any ERROR, or any WARNING under --strict. |
2 |
Bad usage: no subcommand given. |
3 |
The path does not exist. |
4 |
The dataset couldn't be loaded (unknown format, malformed, or the sample download failed). |
INFO findings never affect the exit code.
# Kick the tyres with no dataset of your own
cvflow inspect --serve
# Fast structural check on a huge dataset (skips reading image bytes)
cvflow inspect ./dataset --no-images
# Gate a CI job on data quality
cvflow inspect ./dataset --strict
# Force the format when auto-detection can't tell
cvflow inspect ./dataset --format coco
# Just the findings, no statistics section, saved to a file
cvflow inspect ./dataset --no-stats > report.txt
# A self-contained page to attach to a PR
cvflow inspect ./dataset --html report.html
# Share the dashboard with a teammate on your network
cvflow inspect ./dataset --serve --host 0.0.0.0 --port 9000Point CVFlow at a dataset and it answers the questions you'd otherwise check by hand, one script at a time:
CVFlow reads object detection, instance segmentation and oriented box (OBB) datasets. Polygons and oriented boxes are audited through their axis-aligned extent, and the task is named in the sidebar and the terminal, so you always know how your labels were read.
- Is anything broken? Corrupt/unreadable images, missing or invalid annotation files, broken paths, bad image dimensions, duplicate filenames.
- Are the annotations sane? Boxes outside the image, negative or zero-area boxes, absurdly tiny or full-frame boxes, duplicate overlapping boxes, class IDs that don't exist.
- Is the distribution weird? Class balance, objects per image, box sizes, aspect ratios, and images that are statistical outliers.
- Do I have duplicates? Exact copies (by hash) and near-duplicates (by perceptual hash, with a similarity score).
- Are my splits leaking? The same (or nearly the same) image in more than one split. This one bites hardest on datasets cut from video.
Every finding carries a severity, a plain-English reason, where it is, the evidence behind it, and a suggested next step. CVFlow won't tell you your dataset is wrong; it shows you what looks off and lets you make the call.
Every finding carries a stable code. Checks run in six families and are
reported most-severe-first. Checks marked needs pixels open the image files,
so they're skipped by --no-images and when the images can't be found under the
dataset root.
Structural problems: files that are missing, unreadable, or internally inconsistent.
| Code | Severity | Triggered when |
|---|---|---|
corrupt-image |
ERROR | The image file exists but can't be decoded. Needs pixels. |
broken-image-path |
ERROR | The annotations reference an image that isn't on disk. |
invalid-image-dimension |
ERROR | The recorded width or height is zero or negative. |
invalid-annotation-file |
ERROR | A YOLO label file has malformed lines (not class_id cx cy w h with numeric values). |
invalid-class-id |
ERROR / WARNING | A box uses a class id that isn't in the class map. ERROR when the id is negative, WARNING otherwise. Skipped entirely when the dataset defines no class names. |
duplicate-filename |
WARNING | The same image filename appears more than once, a common sign of a merge that overwrote samples. |
empty-image |
WARNING | An image has no annotations. Legitimate for background samples, so it's a prompt to confirm, not a failure. One finding per image up to a limit, then the tail is summarized in one further finding. |
images-not-found |
INFO | No image file could be located under the dataset root, so every pixel-reading check was skipped. Reported once, instead of flooding the report with false broken-image-path findings. |
Box geometry, on every task. Polygons and oriented boxes are audited through their axis-aligned extent.
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
box-out-of-bounds |
ERROR | A normalized coordinate falls outside 0–1. |
out_of_bounds_eps (1e-3) tolerance before it counts |
degenerate-box |
ERROR | Width or height is zero or negative. | n/a |
tiny-box |
WARNING | A side is below 1% of the image. Degenerate boxes are excluded; they're reported by their own check. | tiny_box_side (0.01) |
huge-box |
WARNING | The box covers more than 90% of the frame. | huge_box_area (0.9) |
duplicate-annotation |
WARNING | Two same-class boxes in one image overlap with IoU at or above the threshold. | duplicate_iou (0.95) |
Task-specific geometry. Each of these is a no-op unless the dataset's task matches, so a detection dataset never sees them.
Segmentation only:
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
sparse-polygon |
ERROR | A polygon has fewer than three points, so it can't enclose an area. | n/a |
empty-mask |
ERROR | The polygon encloses zero area. | n/a |
rectangular-mask |
WARNING | The mask fills essentially all of its own bounding box: it's a rectangle wearing a polygon's clothes. | rectangular_mask_fill (0.98) |
sliver-mask |
WARNING | The mask fills very little of its own extent, which usually means a broken or collapsed contour. | sliver_mask_fill (0.05) |
Oriented boxes (OBB) only:
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
non-rectangular-obb |
WARNING | A corner deviates from square by more than the tolerance, so the four points aren't a rectangle. | obb_corner_tolerance (5.0 degrees) |
unrotated-obb |
INFO | Every oriented box in the dataset is axis-aligned. The dataset is labeled as OBB but carries no rotation, so it may have been exported as plain boxes. Reported once for the dataset. | obb_flat_tolerance (1.0 degrees) |
Distribution anomalies. All of these are candidates for review, never verdicts,
which is why none of them is an ERROR.
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
objects-per-image-outlier |
WARNING | An image holds far more objects than the dataset average: above mean + sigma × std, and above an absolute floor so small datasets don't produce noise. |
objects_outlier_sigma (3.0), objects_outlier_floor (10) |
rare-class |
INFO | A class accounts for less than 1% of all annotations. | rare_class_fraction (0.01) |
class-imbalance |
INFO | The most-common class outnumbers the least-common by more than 100×. Reported once for the dataset. | class_imbalance_ratio (100.0) |
The same module also computes the descriptive statistics printed in the report
and charted in the dashboard: class counts, objects per image, box areas, and
aspect ratios. Those are suppressed with --no-stats, but the checks above still
run.
Redundant samples. Both checks hash the image files, so both need pixels.
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
exact-duplicate |
WARNING | Two or more images share an identical SHA-256, so they are byte-for-byte copies. One finding per group, listing up to ten paths. | n/a |
near-duplicate |
WARNING | Two images' perceptual hashes (64-bit dHash) differ by at most 5 bits. A distance of 0 is skipped, since exact-duplicate already covers it. Reporting stops after 100 pairs so a heavily duplicated dataset can't flood the report. |
near_duplicate_max_hamming (5), max_reported_duplicate_pairs (100) |
Near-duplicate detection is a pairwise comparison, which is the slow part of a
run on a large dataset. --no-images skips it.
| Code | Severity | Triggered when | Threshold |
|---|---|---|---|
split-leakage |
WARNING | Visually near-identical images appear in two different splits: their perceptual hashes differ by at most 5 bits. One finding per pair of splits, counting every match and naming the closest one. Needs pixels. | leakage_max_hamming (5) |
This check only runs when the dataset has at least two splits. It's the one that bites hardest on datasets cut from video, where consecutive frames end up on both sides of a train/val boundary and inflate your metrics.
The CLI deliberately exposes only --no-images; every other threshold lives on
CheckConfig and is reachable from Python:
from cvflow.analysis import AnalysisEngine, CheckConfig, default_checks
from cvflow.loaders import load_dataset
config = CheckConfig(
tiny_box_side=0.005, # your objects really are that small
near_duplicate_max_hamming=2, # only flag very close pairs
rare_class_fraction=0.02,
)
issues = AnalysisEngine(default_checks(config)).run(load_dataset("./dataset"))Every field, with its default, is documented inline on CheckConfig in
src/cvflow/analysis/engine.py. Checks read only
the fields they care about, so adding one doesn't ripple outward.
Because the exit code is meaningful, cvflow inspect drops straight into a
pipeline as a data-quality gate.
Show the workflow steps
- name: Check dataset quality
run: |
pip install cvflow
cvflow inspect ./dataset --strictA few things make CI runs pleasanter:
- Use
--no-imageswhen the images aren't checked out, or when the run has to stay under a couple of minutes. You lose corrupt/duplicate/leakage detection. - Use
--html report.htmland upload the file as a build artifact. It's self-contained (one HTML file, no server, no external requests), so it attaches to a PR or an artifact store as-is.
- name: Dataset report
run: cvflow inspect ./dataset --strict --html report.html
- uses: actions/upload-artifact@v4
if: always()
with:
name: dataset-report
path: report.htmlThe CLI is thin wiring over a small public API, so you can run the same pipeline yourself: to filter findings, feed them into another tool, or tune a threshold the CLI doesn't expose.
Show the Python API
from cvflow.analysis import AnalysisEngine, CheckConfig, compute_statistics, default_checks
from cvflow.loaders import load_dataset
from cvflow.model import Severity
from cvflow.report import render_report
dataset = load_dataset("./dataset") # fmt="coco" to force a format
config = CheckConfig(check_images=True, tiny_box_side=0.005)
issues = AnalysisEngine(default_checks(config)).run(dataset)
errors = [issue for issue in issues if issue.severity is Severity.ERROR]
print(f"{len(errors)} errors out of {len(issues)} findings")
print(render_report(dataset, issues, stats=compute_statistics(dataset)))load_dataset raises cvflow.exceptions.DatasetError (or its
UnsupportedFormatError subclass) when the path can't be read as a dataset.
Findings come back sorted most-severe-first, and every Issue carries a code,
severity, message, why, location, evidence, and suggestion. Nothing
in this API prints or exits; that's the CLI's job alone.
To run a subset of the check families rather than all of them, assemble the list
yourself. integrity_checks, annotation_checks, shape_checks,
statistics_checks, duplicate_checks and leakage_checks each return the
checks for one family:
from cvflow.analysis import AnalysisEngine, CheckConfig, annotation_checks, integrity_checks
config = CheckConfig()
engine = AnalysisEngine([*integrity_checks(config), *annotation_checks(config)])cvflow: command not found: use python -m cvflow instead.
error: path does not exist (exit 3): check the path; CVFlow doesn't
search for it.
The format couldn't be detected (exit 4): the folder doesn't look like
either layout. Pass -f yolo or -f coco to force one, or move the labels and
images into the expected structure.
Every image reports broken-image-path: the labels reference images CVFlow
can't find from the dataset root. Point it at the parent folder that holds both,
or fix the paths in data.yaml.
"No image files were found next to the annotations": an INFO finding, not
a failure. The structural and annotation checks still ran; only the pixel checks
were skipped.
The run seems to hang: it's hashing images. The stderr note says so before
it starts; --no-images skips that phase.
The browser didn't open: the URL is printed either way; open it by hand.
--no-browser makes that the intended behavior.
"Address already in use": you passed --port for a port something else
holds. Drop --port to let CVFlow find a free one.
A teammate can't reach the dashboard: the default host is loopback. Re-run
with --host 0.0.0.0, and read the warning about what that
exposes.
Save is greyed out in the editor: the dataset isn't YOLO detection, or
you're looking at an --html file. See What can be edited.
git clone https://github.com/RizwanMunawar/cvflow
cd cvflow
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"All four must pass; CI runs exactly these:
ruff check . # lint
ruff format --check . # formatting (ruff format . auto-fixes)
mypy # strict type-checking
pytest # testsHandy pytest variations while you work:
pytest -q # quiet
pytest tests/test_annotations.py # one file
pytest -k "leakage or duplicate" # by name
pytest --cov=cvflow --cov-report=term-missing # with coverage, as CI runs it
pytest -x -vv # stop at the first failure, verboseThe suite is pure-Python and fast, with no fixtures to download and no GPU. Tests
mirror the package structure one file per module, and
tests/conftest.py builds minimal-but-valid YOLO and COCO
datasets in a temp directory, so loader and check tests exercise real file
parsing without shipping binary fixtures.
CI runs lint and type-checking on Python 3.11, and the test suite on 3.9, 3.10,
3.11 and 3.12. See .github/workflows/ci.yml.
Build a tiny broken dataset by hand and watch CVFlow catch it. This is the fastest way to see the whole pipeline work after a change:
mkdir -p demo/images/train demo/labels/train
python -c "
from PIL import Image
Image.new('RGB', (640, 480), (90, 110, 140)).save('demo/images/train/frame_0001.jpg')
Image.new('RGB', (640, 480), (40, 60, 80)).save('demo/images/train/frame_0002.jpg')
"
cat > demo/data.yaml <<'YAML'
path: .
train: images/train
names:
0: person
1: helmet
YAML
# one good box, one that runs off the right edge, one that's absurdly tiny
printf '0 0.512 0.437 0.104 0.216\n1 0.980 0.500 0.200 0.100\n' > demo/labels/train/frame_0001.txt
printf '0 0.500 0.500 0.002 0.003\n' > demo/labels/train/frame_0002.txt
cvflow inspect ./demoYou should see one ERROR (box-out-of-bounds) and one WARNING
(tiny-box), and the process should exit 1:
Health Summary
────────────────────────────────────────────────────
ERROR 1
WARNING 1
INFO 0
echo $? # 1 (errors were found)Then check the other two outputs against the same folder:
cvflow inspect ./demo --serve # the dashboard, with the editor live
cvflow inspect ./demo --html report.html && open report.html # the static pageOpen a box in the dashboard, drag it back inside the frame, hit Save labels,
and re-run cvflow inspect ./demo. The box-out-of-bounds error should be gone
and the exit code back to 0.
- Keep PRs small and focused: one logical change per PR.
- Add or update tests for behavior changes.
- Prefer language like "potential issue" / "worth reviewing" over declaring something definitively wrong. See Design philosophy.
- Add a dependency only when it clearly earns its place.
More detail in CONTRIBUTING.md. For dataset-related bug
reports, a minimal reproducible dataset layout (a few files) is enormously
helpful.
Dataset ─▶ Loaders ─▶ Normalized model ─▶ Analysis engine ─▶ Issues ─┬─▶ Report (text)
(cvflow.model) ├─ integrity │
├─ annotations └─▶ Design (dashboard)
├─ shapes
├─ statistics
├─ duplicates
└─ leakage
Data flows one direction. Each stage depends only on the stage before it through a stable interface, never on another stage's internals, so a new format, rule, or output slots in without touching the rest.
src/cvflow/
model/ normalized, format-agnostic domain model (Issue, Severity, …)
loaders/ dataset format loaders (YOLO, COCO)
analysis/ the analysis engine and every check
imaging.py the only place Pillow is used: readability, hashing, dHash
report/ findings → prioritized text report
design/ findings → the browser dashboard and its editor
sample.py downloads coco128 when inspect is run with no path
cli/ the command-line entry point
tests/ pytest suite mirroring the package structure
Module-by-module detail
The format-agnostic vocabulary everything else speaks in:
Severity:ERROR/WARNING/INFO, ordered by seriousness.Issue: the unit of feedback, carryingcode,severity,message,why,location,evidenceandsuggestion.Location: where a finding was detected (path, split, annotation index).BoundingBox: an annotation in canonical normalizedxyxycoordinates, withfrom_yolo/from_cococonstructors and geometry helpers.ImageItem/Dataset: an image (path, split, dims, boxes) and the collection of them (format, root, class-name map, split/count helpers).DatasetStatistics/Summary: computed descriptive statistics, kept in the model so the reporter never has to importanalysis.
Because the analysis engine only ever sees this model, loaders and checks evolve independently of one another.
Translate an on-disk dataset into the normalized model. DatasetLoader is the
base interface (detect() / load()); a small registry exposes
load_dataset(path, fmt=None) with auto-detection and a clear
UnsupportedFormatError. YoloLoader and CocoLoader ship today.
Loaders are deliberately lenient: they represent whatever is on disk (including out-of-range or malformed values) and leave judgment to the analysis engine.
They also record the task they read: detect, segment or obb. One YOLO
text format serves all three, told apart by how many numbers follow the class id;
polygons and oriented boxes are reduced to their axis-aligned extent so every
check works on one geometry, while Dataset.task remembers what they came from.
That matters on the way back out: the editor refuses to rewrite anything but
plain detection labels.
Check: the base class;run(dataset) -> Iterable[Issue].AnalysisEngine: runs a list of checks and returns findings sorted most-severe-first.CheckConfig: one small options object (image-byte toggle plus the thresholds listed under Every finding, by code). Checks read only the fields they care about.default_checks(config): assembles the full check set forcvflow inspect.
Each family (integrity, annotations, shapes, statistics, duplicates, leakage)
is self-contained and emits Issues, never printing or setting global policy.
imaging is a thin, import-guarded wrapper over Pillow: readability/verify,
size, streamed SHA-256 (file_hash), perceptual dHash (perceptual_hash), and
hamming_distance. Isolating Pillow here keeps the rest of the codebase free of
a hard image dependency and gives duplicate/leakage detection one source of truth
for hashing.
analysis.paths.resolve_image_path() locates the actual image file for an
ImageItem across the layouts different formats use (absolute YOLO paths, bare
COCO file names). Shared by every check that reads pixels, so path handling lives
in one place.
render_report() turns a list of Issues (plus optional DatasetStatistics)
into a prioritized text report: overview, statistics, health summary,
most-important problems, findings-by-type, and per-issue detail. Reporting is
separate from analysis so new output formats can be added without touching the
checks.
Everything the user sees in a browser. It's a sibling of cvflow.report, not
a layer above it: both consume Issues and neither knows about the other.
payload.build_payload(): the data contract. Turns aDataset, its findings, and its statistics into one plain JSON-serializable dict, with every aggregate precomputed in Python: class ranking and cumulative coverage, objects/size/shape histograms, the box-center heatmap, per-code counts, and the images carrying the most findings. The page never reasons about the dataset itself; adding a chart usually means adding one function here.assets/:dashboard.html,dashboard.css,dashboard.js, the UI as editable design artifacts. No framework, no build step, no external requests.assets/vendor/: Chart.js and the Geist fonts, inlined into the page at render time. Versions and licenses inassets/vendor/README.md.editor.Editor: the write path. Resolves an image the browser asks for, returns its bytes and boxes, and writes edited boxes back to the YOLO label file. Every path is confined to the dataset root, and only YOLO detection is written.dashboard.render_dashboard(): inlines the assets and the payload into a single self-contained HTML file (--html, or served in memory).server.serve_dashboard(): a loopbackhttp.serverholding one page in memory. With anEditorattached it also answers four endpoints:GET /api/editor(is this dataset writable, and which classes?),GET /api/image,GET /api/annotationsandPOST /api/annotations. Without one it serves the page and 404s everything else.
sample downloads and unpacks Ultralytics' coco128 into a cache directory when
cvflow inspect is run with no path. Standard library only, cached after the
first fetch, and every archive member is resolved and checked before extraction
so a crafted zip can't write outside the cache. Nothing here runs unless the path
is omitted.
cli is the user-facing entry point. It parses arguments, wires
load → analyze → report, and chooses the process exit code from the findings. It
owns no analysis logic of its own.
- New format → add a
DatasetLoaderand register it. - New check → add a
Checkand include it indefault_checks(); the engine, CLI, and reporter pick it up automatically. - New threshold → add a field to
CheckConfig. - New output → add a renderer in
cvflow.report. - New UI → edit the assets in
cvflow.design; add a field to the payload only if the page needs data it doesn't already have.
None of these require changes outside their own module. Add tests alongside, and run the four checks above before you push.
Don't tell developers their dataset is wrong. Show them what looks suspicious, explain why, and let them decide.
That principle is baked into the tool. Severity is used honestly: ERROR means
something is objectively broken, WARNING means "worth a look", and INFO is an
observation. Statistical oddities and duplicates are never hard errors; they're
candidates for review, phrased that way on purpose. The dashboard's accuracy
estimate is labeled an estimate, with its formula on screen.
Four rules hold the codebase together:
- Deterministic first. Validation, statistics, hashing, and lightweight CV techniques form the core. AI/model-assisted checks would be additive and optional, never required.
- One-directional data flow. Loaders → model → analysis → issues → report.
- Findings, not verdicts. Every issue carries severity, reason, evidence, and a suggestion; cautious language by default.
- Lean dependencies. Two runtime dependencies (
pyyaml,pillow), each added when a feature genuinely needed it. The dashboard adds none: it's standard-library rendering plus static assets.
- ✅ Project foundation: CLI, packaging, model, tests, CI
- ✅ Dataset loaders: YOLO & COCO → one normalized model
- ✅ Integrity analysis: corrupt images, missing/invalid annotations
- ✅ Annotation analysis: bounding-box validation & anomalies
- ✅ Dataset statistics: distributions & outlier detection
- ✅ Duplicate detection: exact + perceptual hashing
- ✅ Split-leakage detection: cross-split similarity
- ✅ Dashboard: one browser page for the whole report
- ✅ Visualization: eyeball the flagged samples
CVFlow is released under the MIT License.
{ "images": [{ "id": 1, "file_name": "000000000001.jpg", "width": 640, "height": 480 }], "annotations":[{ "id": 1, "image_id": 1, "category_id": 1, "bbox": [100, 120, 80, 160] }], "categories": [{ "id": 1, "name": "person" }] }