Example
Before
REPORT TITLEThis para- graph was copied from a PDF. 1REPORT TITLE This stays separate. 2REPORT TITLE Final page. 3
After
This paragraph was copied from a PDF. This stays separate. Final page.
What to check
- Tables, columns, poetry, and code still need review.
- Repeated headers and footers are detected only when page boundaries are present.
- A hyphen before an uppercase letter is kept.
How this cleaner works
PDF pages often store text as positioned lines rather than flowing paragraphs. When that text is copied, every visual line can become a hard newline. TextClean joins prose while preserving likely headings, lists, table-like rows, and indented code. It can also remove page-number-only lines.
When copied pages include page separators, TextClean compares the first and last lines across three or more pages. Repeated boundary text can be removed as a likely header or footer. This is a conservative pattern check, so the original input remains visible for review.
Dehyphenation is deliberately conservative. A letter followed by a hyphen and newline is rejoined only when the next line begins with a lowercase letter. That fixes common splits such as “para- graph” while retaining many genuine compounds and heading boundaries. Always review technical terms, tables, citations, poetry, and multi-column source material.
Useful for
- Journal articles copied as text
- Reports with hard-wrapped paragraphs
- Ebooks or manuals after text extraction
Common questions
Can I upload a PDF?
No. This tool cleans text you have already copied or extracted. It does not read files or perform OCR.
Will it preserve tables?
Not reliably. Tables and multi-column layouts depend on line structure, so use the input and output panes to review them manually.
What happens to real paragraphs?
Blank-line-separated blocks remain separate. Lines inside each block are joined with spaces.