CSV Diff
CSV Diff Toolkit
Compare two versions of a CSV at the row and cell level. Drop in the old and new files and
it picks the key column itself — the columns both files share appear as chips under the bar,
with the one it chose marked, so changing it is a click rather than a guess at a name. The
result is an added / removed / changed breakdown with cell-level highlighting. More useful
than plain diff for tabular data, which flags reordered rows as different and
cannot point at the cell that changed — the typical use is sanity-checking a data migration
before cutover.
Before you start
You need:
- Two CSVs — the "old" (left) and "new" (right) version of the same dataset. They should share a header row and, ideally, share most columns.
- A key column (or composite) that identifies a record in both files — usually an
id,skuoremail. You do not have to name it: the tool reads the headers of both files and picks one, and the chips under the bar let you change it. Only if nothing usable is found does it fall back to comparing whole rows.
The two files don't have to have the same columns or the same row order. The diff matches rows by key, then compares cells of the matched rows.
How to use it
- Paste or drop your old CSV into the left pane.
- Paste or drop your new CSV into the right pane.
- Check the key. The chips under the bar are the columns both files have; the dashed one is what the tool picked. Click another to switch, click a second to make it a composite key (
country+sku), or click whole row to compare everything. - The result updates as you paste. Compare re-runs it if you want.
- Review the result below the panes. Colour legend:
- green — row exists only in the new file (added).
- red — row exists only in the old file (removed).
- yellow — row exists in both, but at least one cell differs. The specific differing cells are highlighted.
- Click Download diff for a CSV with a
change_typecolumn.
How the key is chosen
A column can be a key if it appears in both files, is never blank, and never repeats — that is exactly the property that lets a row in one file be matched with a row in the other. Chips for columns that fail the test are dimmed, and hovering one says why.
One exception, deliberately: a column named like an identifier (id,
customer_id, sku, email…) is chosen even when it repeats,
as long as it is never blank. A repeating id is a problem in the data, and matching
on some other column that happens to be unique would hide it behind a diff that looks fine. The
status line says how many duplicate keys there were and which side they were on.
Click whole row to compare entire rows as a set instead. That is the right choice when a file has no stable identifier, but it is weaker: a single changed cell reads as a delete plus an add, and there is no cell-level highlight.
What "changed" really means
- Comparison is string-exact.
"1.0"and"1"are different. Normalise types in the source if that's noise. - Only columns that exist in both files are compared for changes. A column added in the new file is reported as part of the "added" row context, not as a cell change for every existing row.
- If the same key appears more than once in one file, only the last row with that key takes part in the comparison — a set of rows keyed by the same value has no single answer to "did it change". The status line counts them and says which file they were in, so you can fix the source before trusting the diff.
Example
Old CSV:
id,name,city
1,Alice,Berlin
2,Bob,Paris
3,Carol,Rome
New CSV:
id,name,city
1,Alice,Berlin
2,Bob,Lyon
4,Dana,Madrid
Diff on id:
- id=1 — unchanged.
- id=2 — changed:
cityParis → Lyon. - id=3 — removed.
- id=4 — added.
Tips & common pitfalls
- Trim noisy columns first. Timestamps and "last updated" columns change on every row and drown out real diffs. Strip them or add them to a separate compare.
- Sort both files by key before download if you plan to eyeball the diff output — the tool itself doesn't need sorted input, but a sorted CSV is much easier to review.
- Big + wide files are slow because every matched row builds a string key and a per-column comparison. If you hit a wall, cut the file by column or by row with Split CSV first.
- Reordered columns are invisible to the diff — the tool lines up columns by name, not by position. If a column was renamed, it looks like "old column removed, new column added".
- Watch out for whitespace: trailing spaces in one file make the cell look "changed" even when it isn't. If the source has that problem, run the file through Dedupe CSV with trim first, or preprocess with a script.
Troubleshooting
Every row shows up as changed.
Most likely the files have different line endings or an encoding mismatch, or the key column has a subtle difference (BOM, trailing space). Open both files in a plain editor and compare the first row byte-for-byte.
Rows I expect to match are flagged as "added + removed".
Your key isn't actually unique, or its values don't match exactly across files. Try a composite key or normalise the key column first (trim, lowercase).
The tool says "duplicate key" but I only have one row with that id.
Check for whitespace or invisible characters in the key cell. Two rows with keys "42" and "42 " collide after trimming.
Frequently asked questions
Why not just use diff?
Command-line diff does byte-level line comparison. It gets confused by reordered rows or different line endings, and it can't tell you which cell changed in a row. This tool understands keys and cells.
Can I diff on a composite key?
Yes. Enter comma-separated column names, e.g. country, sku. Rows match when the full tuple matches.
Is my data uploaded?
No. Both files stay in your browser. See the privacy policy.
Can I export only the changed rows?
Yes — Download diff exports all rows with a change_type column (added / removed / changed / unchanged). Filter that column in any spreadsheet or with awk.