Guide
One sentence carries this page: software is not a weaker human reader, it is a different reader — and the ways it differs are not the ways people assume.
The assumption worth dropping first is that automated reading is a harder version of the same thing, so a cover that would defeat a determined person will comfortably defeat a script. Sometimes true, often irrelevant, because the risks are in different places. A person who cannot read your blurred account number gives up and tells you the image is unclear. A parser does neither of those things. It writes down a number.
Note what this page is not about. It is not about whether an effect can be reversed, and it is not about image matching, which is a separate machine-reading problem with its own answer — see does blurring a face protect against reverse image search. This is about text and fields: what software extracts from a frame you have already redacted, what it does with an extraction it is not sure about, and what its requirements do to your choices before you draw a single box.
Four ways a software reader differs from a person
1. It reads the entire frame at one price. A person looks where the interesting thing is. Attention is scarce, so we skim, and every judgement of the form nobody would bother with that is really a judgement about someone else’s effort. Automated reading has no notion of importance and no fatigue. Nine-pixel footer text, a sidebar, a watermark, the tail of a URL in a visible address bar, a notification that slid in at the top of the capture, the half-row clipped at the bottom edge: all read with exactly the attention given to the headline. This inverts a habit. The margins of the frame do not matter less to a machine, they matter the same, which in practice means more, because they are the parts you did not examine.
2. It does not report that something is illegible. The least intuitive difference and the most consequential. A recogniser is built to produce output; returning nothing is a failure state it avoids. Presented with a smeared or blocky glyph it will usually emit a character, perhaps with an internal confidence score that no one downstream displays, and from that point on the character is indistinguishable from one lifted off clean print. Both outcomes are bad and neither is visible to you. A wrong guess writes a false value into a record you cannot inspect. A right guess writes the value you were trying to hide into the same record. Constraint sharpens the guess considerably: a field with a known shape — a date, a postcode, an account format, a number carrying its own check digit — has a small set of legal values, so format can finish a character that the surviving pixels only began. Softened text in a structured field is a much weaker cover than softened text in free prose, and it looks identical.
3. It creates a copy that is not an image. One upload, two artefacts: the file, and whatever was read out of the file. The second lives in a database column, a search index, a log line or a cache, and it is searchable in ways a picture never was. This matters because every intuition about deleting things attaches to files. Withdrawing an attachment, an email recall, a retention window, a colleague saying they had deleted it — all of these describe the picture. None of them describe a text field that was created the moment the picture arrived.
4. It has requirements, and can refuse. A human recipient with a document missing the field they wanted will ask you about it, or work around it, or notice that the rest is enough. Automated intake applies rules. A required field that is not present, or a document whose structure it cannot recognise, produces an error rather than a conversation, and that error is the beginning of the sequence in the next section.
The rejection loop, where the actual disclosure happens
Follow it through once, because almost nobody plans for it and it is extremely common.
You do the careful thing. You cover the account number, the address, the other line items, the details that have nothing to do with the claim. You upload. Some minutes or days later something comes back: document unreadable, or required information missing, or please resubmit a clear copy. You are now in a position where every easy option is worse than what you just did. Resubmit with less covered, guessing at what offended the parser. Or send the original by email to whoever is now handling it, which is the option most people take, because it ends the problem immediately and feels cooperative. The original then sits in a mailbox, gets forwarded internally, and is attached to a ticket, in full, with none of the judgement you applied the first time.
The redaction was not defeated. It was made expensive, and the cost was paid by abandoning it. That is worth stating plainly: on an automated intake path, a cover that breaks the parse is more dangerous than no cover at all, because it manufactures a request for the unredacted file.
Which means the work starts before the image is open. Find out what the system actually needs. The requirements text next to the upload button usually says more than people read. The form itself is a strong signal: if there is a field asking you to type the reference number, the system wants the number and the image is corroboration, so covering it in the picture will fail a match rather than protect anything. If the instructions are vague, ask, and wait for the answer rather than covering optimistically and hoping. And if the field the system genuinely requires is the sensitive one, accept that the image is the wrong artefact for this: use the typed field where the value at least lands in a place designed for it, or move to a channel that has an agreement behind it, rather than uploading a picture and hoping the parser is lenient.
What to cover, and what to leave alone so the document still parses
Use the flat fill, not the soft ones. With a human reader there is a real argument for blur or pixelation: it keeps the page legible as a page and signals what kind of thing was there. With a software reader that argument collapses, because the residual structure the effect preserves is precisely what a recogniser scores its candidates against. Black Box replaces the region with a flat near-black and there is nothing left to score. The empirical version of this question, for the most machine-oriented target there is, is worked through in can a blurred or pixelated QR code still be scanned, and the answer there generalises: partially destroyed structure is structure.
Do not cover the scaffolding. Field detection frequently works from context rather than from the value: the label to the left, the box that encloses it, the rules of a table, a header band, a logo, the edges of the page. A generous rectangle that takes out the value along with its label and the table rule around it can turn a document the system knows how to read into one it does not. Cover values snugly and leave structure visible. This is the opposite of the advice that suits a human reader, where covering an entire block is usually the safer instinct.
Cover both copies of anything written twice. A machine-readable block and the characters printed beside it are two encodings of one fact. Reference numbers reappear in headers, footers, window titles and file names. A summary line restates a detail line. Cover one and leave the other and you have covered nothing, from the point of view of whichever reader you forgot. Then invert the check: is there a machine-readable region the receiving system needs in order to process the document at all? Covering that is a rejection waiting to happen.
Leave nothing half covered. For a person, half a digit is ambiguous and mildly annoying. For a recogniser working in a constrained field, half a digit plus the format is frequently the whole digit. Covers need to overlap the glyphs, including descenders and the thin gap where the next character begins, not merely reach them.
Check that small covers landed. The editor above silently discards a rectangle drag under about four image pixels and an oval under six. Machine-read documents are exactly where this bites, because the fields worth hiding on a dense form are small, and a cover that did nothing looks the same as a cover you never drew.
Prefer a narrower capture over a heavily covered one. If the system needs one field from a document full of others, a capture that contains only what is needed parses more reliably than a full page under a grid of boxes, and it removes the scaffolding problem entirely. Fewer covers, fewer landmarks destroyed, fewer things to get wrong.
Where the extracted text goes
You cannot audit this from outside, which is the reason to reason about it at all rather than hoping.
Assume extraction happened. If a service tells you what your document said, or fills fields for you, or lets you search your own attachments by their contents, it read the file. If it does none of those things visibly, it may still have read it, because extraction is cheap and is often applied by default for indexing and for accessibility. The safe default is that anything legible in what you sent now exists as text.
Assume the text outlives the file. Deletion that you can observe is deletion of the file. Whether the derived text goes with it depends on systems nobody exposes to you. If deletion genuinely matters, say both things explicitly when you ask — the attachment and anything extracted from it — and understand you may not be able to verify either.
Assume copies fan out. Previews, thumbnails, a rendered version in an automatically created ticket, an email notification that helpfully attached the file, a mirror in a second system that subscribes to the first. Each is generated from the version you sent, which is the one point in the chain you control.
Read the terms in front of you rather than assuming. Policies on retention and on whether uploads may be used for other purposes vary between services and change over time, so a claim about any particular one does not belong on a page like this. What does belong: the asymmetry. You get one chance to decide what is legible in the file you hand over, and everything after that is out of your hands.
A running order when software is the first reader
1. Establish the requirement before opening the image. Which fields does this system need, and which of them are already typed into the form? Ask if it is not stated. This step prevents the rejection loop and nothing else does.
2. Choose or capture the narrowest frame that satisfies it. A view that never contained the sensitive field beats a view where you covered it.
3. Cover with Black Box, value by value, snugly. Leave labels, table rules, borders and page structure intact so the parse still finds its landmarks.
4. Sweep for second copies. Machine-readable blocks, headers, footers, window titles, summary lines, anything restated.
5. Inspect at full size. Confirm every cover landed and overlaps the glyph edges, and scroll the whole image — the preview panel is capped in height and scrolls, so a tall document has regions you have not looked at, and a machine will read those regions whether or not you did.
6. Read your export back the way the machine will. Exhaustively, edge to edge, at full size, ignoring what you know is underneath. If your device offers text selection or lookup on an image, or you have a document scanner app to hand, run it over the export and see what comes back. That is the nearest you can get to the reader’s own view, and it is the only check on this page that produces evidence instead of an opinion.
7. Rename the file before uploading, and note which version you sent. The file name travels with the upload and is stored as text like everything else. Keeping a note of what you sent is what lets you answer a later rejection without reaching for the original.
What the editor above does and does not do here
This section is read from this page’s own code, because the boundary matters when the reader is software.
Three modes, and only one of them suits a machine reader. Black Box fills the selected region with a flat near-black. Blur rebuilds the region from a copy reduced by a factor of ten and scaled back up. Pixelate averages it into blocks of at least six pixels, each taking the colour of its centre pixel. All three genuinely replace the pixels inside the selection, so none of them can be peeled back — but two of them leave structure, and structure is what a recogniser needs. If a parser is in the path, use the first one.
It has no text recognition of any kind. There is no OCR and no content analysis anywhere in the code. It cannot tell you what a machine would read off your image, cannot locate fields, and cannot warn you that a cover is too small or that a value appears twice. Every check described above is yours to run; the tool executes covers and nothing else.
It is entirely region-based. Three selection shapes — rectangle, oval and freehand lasso — and every operation replaces pixels inside the selection you drew. There is no crop, no resize, no rotate, no text tool and no whole-image operation of any kind, so narrowing the frame has to happen at capture time, and substituting a value with a placeholder has to happen upstream in whatever produced the document.
Very small drags are silently ignored. Under about four image pixels for a rectangle and six for an oval, nothing is applied and nothing is said. Worth re-reading when the targets are form fields.
The preview is scaled and scrolls; the export is not. Edits run at true pixel size while the canvas is displayed scaled to fit a panel capped at seventy per cent of the window height, with its own scrollbar. The export contains the whole image at the source’s exact pixel dimensions, including the part that was below the fold while you worked.
Download writes a fresh flat PNG. The file is exported from the canvas as an image named hideshot- followed by a timestamp, re-encoded, with no layers, no objects, no text content and no tags carried over from the source. Rename it before you upload it anywhere.
Nothing you open here leaves your browser. The image is read locally, the editing happens on a canvas in the open tab, and the page stores nothing between sessions — no localStorage, no sessionStorage, no IndexedDB, no account. The only upload in this entire story is the one you make to the other service, and the point of the page is that you get to choose exactly what goes into it.
Common mistakes and misconceptions
Assuming a machine reader is just a persistent person. It is worse in some ways and better in others, and the differences are not on the axis of effort.
Expecting an illegible field to come back as illegible. It comes back as a value. That value is then treated as data.
Blurring a structured field. A date, a postcode or a checksummed number has few legal values, so the format can complete what the pixels left.
Treating a margin as unimportant. Cheap for you to ignore, free for software to read.
Covering generously over labels, rules and borders. Good instinct for a human reader, and the fastest route to a document the parser cannot recognise.
Not finding out which fields are required first. The single step that prevents the rejection loop, and the one almost everyone skips.
Answering a rejection with the original file. The most common way careful redaction ends in a full disclosure. Decide in advance that this is not the response.
Covering the printed reference but not the machine-readable block, or the reverse. Two encodings of one fact; covering either alone covers nothing.
Covering a machine-readable block the system needs. Another rejection, from the opposite direction.
Believing a deleted attachment means deleted data. Extraction happens on arrival and the text lives somewhere else.
Forgetting the file name. It uploads with the file and is stored as text, so it is read as reliably as anything in the frame.
Only checking the visible part of a tall document. The preview scrolls; the export does not omit what you did not scroll to.
Judging the export while remembering the original. Your eye fills in the covered field. A recogniser has only what is there, which is the view you need to take.