unstructured

mirror of https://github.com/Unstructured-IO/unstructured.git synced 2025-11-06 21:29:42 +00:00

Author	SHA1	Message	Date
Steve Canny	31bef433ad	rfctr: prepare to add orig_elements serde (#2668 ) Summary The serialization and deserialization (serde) of `metadata.orig_elements` will be located in `unstructured.staging.base` alongside `elements_to_json()` and other existing serde functions. Improve the typing, readability, and structure of that module before adding the new serde functions for `metadata.orig_elements`. Reviewers: The commits are well-groomed and are probably quicker to review commit-by-commit than as all files-changed at once.	2024-03-20 21:27:59 +00:00
Wahab Alshahin	e4e25c9feb	Add clean_ligatures to core cleaners (#1326 ) # Background [Ligatures](https://en.wikipedia.org/wiki/Ligature_(writing)#Ligatures_in_Unicode_(Latin_alphabets)) can sometimes show up during the text extraction process when they should not. Very common examples of this are with the Latin `f` related ligatures which can be very subtle to spot by eye (see example below), but can wreak havoc later. ```python "ﬀ": "ff", "ﬁ": "fi", "ﬂ": "fl", "ﬃ": "ffi", "ﬄ": "ffl", ``` Several libraries already do something like this. Most recently, `pdfplumber` added this sort of capability as part of the text extraction process, see https://github.com/jsvine/pdfplumber/issues/598 Instead of incorporating any sort of breaking change to the PDF text processing in `unstructured`, it is best to add this as another cleaner and allow users to opt in. In turn, the `clean_ligatures` method has been added in this PR - with accompanying tests. # Example Here is an example PDF that causes the issue. For example: `Beneﬁts`, which should be `Benefits`. [example.pdf](https://github.com/Unstructured-IO/unstructured/files/12544344/example.pdf) ```bash curl -X 'POST' \ 'https://api.unstructured.io/general/v0/general' \ -H 'accept: application/json' \ -H 'Content-Type: multipart/form-data' \ -H 'unstructured-api-key: ${UNSTRUCTURED_API_KEY}' \ -F 'files=@example.pdf' \ -s \| jq -C . ``` # Notes An initial list of mappings was added with the most common ligatures. There is some subjectivity to this, but this should be a relatively safe starting set. Can always be expanded as needed.	2023-09-07 21:30:18 +00:00
Charles	de855bb4ed	enhancement: new extract function for detecting image URLs (#1212 ) - Adds new feature discussed in GitHub Issue #1117 and in slack	2023-08-30 11:29:15 -07:00
John	6e5d27c6c3	fix pdf partition of list items being detected as titles in OCR only mode (#1119 ) Closes Github issue #1010 adds group_bullet_paragraph func to handle grouping of bullet items that are split across multiple lines	2023-08-15 09:35:54 -07:00
shreyanid	433d6af1bc	fix: format Arabic and Hebrew annotated encodings (#823 ) * add modified arabic and hebrew encodings * added calls to format_encoding_str so encoding is checked before use * added formatting to detect_filetype() * explicitly provided default value for null encoding parameter * fixed format of annotated encodings list * adding hebrew base64 test file * small lint fixes * update changelog * bump version to -dev2	2023-06-27 18:15:02 -07:00
Matt Robinson	3f80301964	fix: handling for emails without datetimes (#724 ) * add empty filetype * add empty handling to partition * changelog and version * handling for when there is no datetime * changelog and version	2023-06-12 17:11:04 +00:00
Matt Robinson	137b4b9a2e	feat: cleaning brick for normalizing bytes string output (#481 ) * add cleaning brick for emojis * changelog and versoin * docs for bytes_string_to_string * different test for bytes_string_to_string	2023-04-13 19:39:08 +00:00
Matt Robinson	c99c099158	feat: enable grouping broken paragraphs in `partition_text` (#456 ) * cleaning brick to group broken paragraphs * docs for group_broken_paragraphs * add docs for partition_text with grouper * partition_text and auto with paragraph_grouper * version and changelog * typo in the docs * linting, linting, linting * switch to using regular expressions	2023-04-06 18:35:22 +00:00
Matt Robinson	9b5cae49e1	fix: allow `replace_mime_encodings` to accept and `encoding` kwarg (#453 ) * changelog and version * added test	2023-04-05 22:53:38 +00:00
natygyoon	e0eb66de52	feat: add staging brick to clean non-ascii characters from unicode (#366 )	2023-03-14 21:31:51 -07:00
Tom Aarsen	5eb1466acc	Resolve various style issues to improve overall code quality (#282 ) * Apply import sorting ruff . --select I --fix * Remove unnecessary open mode parameter ruff . --select UP015 --fix * Use f-string formatting rather than .format * Remove extraneous parentheses Also use "" instead of str() * Resolve missing trailing commas ruff . --select COM --fix * Rewrite list() and dict() calls using literals ruff . --select C4 --fix * Add () to pytest.fixture, use tuples for parametrize, etc. ruff . --select PT --fix * Simplify code: merge conditionals, context managers ruff . --select SIM --fix * Import without unnecessary alias ruff . --select PLR0402 --fix * Apply formatting via black * Rewrite ValueError somewhat Slightly unrelated to the rest of the PR * Apply formatting to tests via black * Update expected exception message to match 0d81564 * Satisfy E501 line too long in test * Update changelog & version * Add ruff to make tidy and test deps * Run 'make tidy' * Update changelog & version * Update changelog & version * Add ruff to 'check' target Doing so required me to also fix some non-auto-fixable issues. Two of them I fixed with a noqa: SIM115, but especially the one in __init__ may need some attention. That said, that refactor is out of scope of this PR.	2023-02-27 11:30:54 -05:00
Mallori Harrell	d7a00046a9	feat: Add new functionality to parse text and header of emails (#111 ) * partition_text function	2023-01-09 17:08:08 +00:00
Sebastian Laverde Alfonso	5a47eb06e9	feat: new bricks for removing and extracting ordered bullets (#128 ) * feat: new cleaning brick for ordered bullets * test: add test for cleaning ordered bullets * feat: new brick for extracting ordered bullets * test: add test for extracting ordered bullets * docs: update CHANGELOG and bump new dev version * chore: change extract ordered bullets return type to tuple * chore: made tidy * chore: regex to split on pattern instead of built-in * chore: catch ValueError, made tidy and fix incompatible type * chore: assertion statements in one line of code * docs: add documentation for new clean and extract bricks to bricks.rst * docs: refactor CHANGELOG 0.3.5.dev5 to dev6 with new bullets * docs: update CHANGELOG 0.3.6-dev0 changes and bump version Co-authored-by: Sebastian Laverde <sebastian@unstructured.io>	2023-01-05 17:06:26 +01:00
Matt Robinson	445533745c	feat: helper functions to identify and extract phone numbers (#124 ) * added pattern for finding phone numbers * added cleaning brick for extracting phone numbers * add docs * changelog and bump version * switch to us phone numbers * bump dev version	2023-01-03 13:31:05 -05:00
Matt Robinson	7a74cdda86	feat: add `partition_email` cleaning brick (#104 ) * fix for processing deeply embedded list elements * fix types in mime encodings cleaner * first pass on partition_email * tests for email * test for mime encodings * changelog bump * added note about \n= * linting, linting, linting * added email docs * add partition_email to the readme * add one more test	2022-12-19 18:02:44 +00:00
Matt Robinson	b1cce16c16	feat: `translate_text` cleaning brick (#101 ) * initial implementation for translate brick * more input validation * tests for translate brick * added docs * bumped version * chinese and arabic tests * re-run pip-compile * add torch to dependencies * cleanup doc string * fix long string * fix typo in docs * take out empty string check * return string if string is empty * added huggingface into make install	2022-12-15 15:35:15 -05:00
Matt Robinson	c62f18c0d0	feat: Add html escape quotes to cleaning brick (#84 ) * feat: Add html escape quotes to cleaning brick * bump changelog	2022-11-29 10:58:31 -05:00
Matt Robinson	300c564c62	feat: Cleaning bricks to extract text before/after a pattern (#63 ) * brick to extract text before * brick for extract text after * tests for extract before and after * updated docs * changelog and bump version * fix typo * fix another typo * positive -> non-negative	2022-11-10 21:35:37 +00:00
Matt Robinson	f3756abc90	feat: Cleaning bricks for removing prefixes and postfixes (#62 ) * added prefix and postfix cleaners * added test for pre and postfix cleaners * added docs for prefix and postfix bricks * changelog and bump version * add dev to version	2022-11-10 12:24:58 -05:00
Matt Robinson	5f40c78f25	Initial Release	2022-09-26 14:55:20 -07:00

20 Commits