{"id":12524004,"name":"easyocrsharp","ecosystem":"nuget","description":"High-accuracy native .NET OCR powered by EasyOCR's neural models running on ONNX Runtime. No Python required.","homepage":"https://github.com/FarhanLodi/EasyOcrSharp","licenses":"MIT","normalized_licenses":["MIT"],"repository_url":"https://github.com/FarhanLodi/EasyOcrSharp","keywords_array":["OCR","EasyOCR","ONNX","TextRecognition","CRAFT","CRNN","Handwriting","TrOCR","Barcode","QR","Redaction","FieldExtraction","Tables","Layout","DocumentAnalysis","PP-Structure","SearchablePDF","Unicode","CLI",".NET10","AOT"],"namespace":null,"versions_count":18,"first_release_published_at":"1900-01-01T00:00:00.000Z","latest_release_published_at":"2026-09-14T14:25:48.947Z","latest_release_number":"3.1.1","last_synced_at":"2026-09-14T14:33:56.485Z","created_at":"2025-11-26T17:00:24.390Z","updated_at":"2026-09-15T20:50:42.186Z","registry_url":"https://www.nuget.org/packages/easyocrsharp/","install_command":"Install-Package easyocrsharp","documentation_url":null,"metadata":{"license_info":{"type":"expression","text":"MIT","version":null},"license_url":"https://licenses.nuget.org/MIT","require_license_acceptance":false,"icon":"icon.png","readme":"README.md","repository":{"type":"git","url":"https://github.com/FarhanLodi/EasyOcrSharp","branch":"refs/heads/main","commit":"c27670c581293167b07a4fea4c256296ed3655f7"},"development_dependency":false,"serviceable":false,"framework_assemblies":[],"package_types":[],"release_notes":"3.1.1: Imaging moves to EasyImageSharp 1.1.0 and ONNX Runtime to 1.30.0. The PP-StructureV3 engine behind AnalyzeDocumentAsync now mirrors PaddleOcrNet 2.0.3-2.1.0: whole-page OCR matched to layout blocks, XY-Cut++ reading order (Auto; right-to-left pages keep direction-aware XY-cut), per-class layout thresholds, LayoutNms on by default, seal text recovery with polygon rectification, multi-line table cells kept in DOCX/XLSX, table-orientation correction, MarkdownRenderOptions, and the Python PaddleOCR 3.x accuracy fixes. FIXED: EasyOcrSharp.Gpu deployed the CPU onnxruntime native for portable builds and publishes, so CUDA was never available - a buildTransitive target now swaps in the GPU runtime (opt out with EasyOcrSharpGpuPreferGpuRuntime=false); OcrResult.UsedGpu reported the requested accelerator instead of the one used; line grouping could merge whole paragraphs into one unreadable box on wide pages; UVDoc unwarp fed RGB instead of BGR and blurred the page through a fixed 488x712 round trip; Arabic-script readings are reordered for display as EasyOCR does; multi-page images are no longer capped by EasyImageSharp's new total-pixel default; the model download ceiling is raised to 2 GB so the full-precision TrOCR decoder can download. ADDED: EasyOcrService.ActiveExecutionProvider, OcrResult.ExecutionProvider, GetRuntimeInfo()/OcrRuntimeInfo, EasyOcrServiceOptions.DetectorModelPath, and actionable GPU fallback diagnostics (CUDA 13 + cuDNN 9). CHANGED DEFAULTS (AnalyzeDocumentAsync only): LayoutNms true, ReadingOrder Auto = XY-Cut++, Markdown omits page furniture. See CHANGELOG.md. 3.1.0: production-operations pass - all additive, every new guard off by default, nothing renamed or removed. OBSERVABILITY: easyocr.operations, .duration, .lines and .pages are now tagged with easyocr.operation, easyocr.outcome, easyocr.provider, easyocr.languages and error.type, so latency and error rate can be sliced by entry point, language and CPU-vs-GPU. Crucially they are recorded on FAILURE too: metrics were previously emitted only on the success path, so a deployment failing every request reported ZERO operations rather than a 100 percent error rate - the one shape of failure a dashboard cannot see. Timeouts and shed load get their own outcomes rather than folding into error, so a burst the concurrency limit handled as designed does not fire the error-budget alert. New saturation instruments easyocr.queue.wait, easyocr.operations.active and easyocr.operations.queued, plus metrics and spans on the previously uninstrumented PDF, searchable-PDF, multi-frame and batch paths. Tag, outcome and operation constants are public as EasyOcrDiagnostics.TagNames, .Outcomes and .OperationNames so alert rules need no magic strings. BACK-PRESSURE: new EasyOcrServiceOptions.MaxConcurrentOperations, QueueTimeout and OperationTimeout, with typed OcrBusyException (503-shaped: at the limit and the queue wait elapsed) and OcrTimeoutException (504-shaped: one operation blew its budget), both deriving EasyOcrSharpException so existing catch-all handlers keep working. Concurrency, not session count, is what sets peak memory, because every concurrent run allocates its own tensors. The gate is taken once per outermost operation; PDF, multi-frame and batch runs gate per page instead, so a long document cannot hold a slot for its whole duration. STARTUP AND READINESS: AddEasyOcrWarmUp(languages) registers an IHostedService that loads models off the startup path (a host that cannot reach the model mirror still starts and reports state instead of crash-looping) and publishes progress through EasyOcrWarmUpState; EasyOcrHealthCheckOptions with DeepProbe, ProbeInterval and ProbeTimeout adds an opt-in health check that actually runs a tiny synthetic page through the real pipeline, caches the verdict, and reports the execution provider that resolved - a truncated model file, a half-copied cache or a GPU whose provider fails at session init previously reported Healthy and then failed every request. The probe never triggers a download, so offline deployments are unchanged, and readiness stays not-ready while warm-up is still running. FIXED: EasyOcrDiagnostics reported version 2.2.1 while the package was 3.x, mislabelling every metric and span it emitted. See CHANGELOG.md. 3.0.1: three bugs reported against PaddleOcrNet, whose PP-StructureV3 engine 3.0.0 brought in-tree, apply to the copy of that engine here and are fixed. (1) Every document-analysis language pack decoded one character class off: the per-script PP-OCRv5 dictionaries (cyrillic, latin, arabic, devanagari, korean, japan, th, el, te, ta, eslav) open with an empty line that IS the CTC blank, so the class they omit is the trailing space; CharacterDictionary prepended a second blank and shifted every class by one, turning Russian into mojibake. BuildVocab now detects the empty first line. The dictionary files were always correct, so nothing re-downloads; the default ch/en/ja recognizers were never affected. (2) OcrResult.ToJson, StructureResult.ToJson and the CLI's --format json escaped everything outside Basic Latin into \\uXXXX; they now serialize through EasyOcrJson.Encoder and write Cyrillic, Greek, Arabic, CJK and the rest verbatim, while still escaping HTML-sensitive characters. New ToJson(JsonSerializerOptions) overloads on both types let callers pick their own encoder, indentation or naming policy without giving up trim/AOT safety. (3) The layout confidence floor was a private const; it is now DocumentAnalysisOptions.LayoutScoreThreshold, default unchanged at 0.5 and still exclusive. That report also turned up two gaps: the layout detectors emit a fixed top-k with no NMS, so duplicate regions reached the caller - a post-processing chain now collapses overlaps, drops sub-6px slivers, reference markers and whole-page image false positives (FilterOverlappingRegions, on by default), with optional LayoutNms, LayoutUnclipRatio and LayoutMergeMode; and PP-DocLayoutV3's predicted reading-order index, previously discarded, now orders the blocks (DocumentAnalysisOptions.ReadingOrder, XY-cut still the fallback). Plain-text OCR is unaffected; every change is additive. 3.0.0: BREAKING - two dependencies leave. (1) Imaging moves to EasyImageSharp (MIT, same author, no build-time licence key and no commercial tier for consumers to inherit); Image\u003cRgb24\u003e and friends appear on the public API (IEasyOcrService, RedactionResult.Image, RedactionOptions.FillColor, DrawAnnotations, BarcodeScanner, the PDF page handlers), so those types now come from a different assembly. (2) The PP-StructureV3 document-structure engine behind AnalyzeDocumentAsync is no longer the third-party PaddleOcrNet package - it is built into this one, under EasyOcrSharp.Structure. Between them, no split-licensed imaging library remains anywhere in the dependency graph. Structure migration is one using directive: PaddleOcrNet.Structure becomes EasyOcrSharp.Structure; StructureResult, StructureBlock and StructureBlockType keep their names, members and behaviour, and everything else that package exposed was engine internals and is now internal. StructureBlock.Lines is now EasyOcrSharp.Models.OcrLine - the same line type the rest of the API returns. Structure models and their cache directory are unchanged, so nothing re-downloads; EASYOCRSHARP_STRUCTURE_MODEL_BASE_URL / _CACHE override them, with the old PADDLEOCRNET_* names still honoured. New dependency: Clipper2 (previously transitive). No OCR behaviour, method name, parameter, default or result shape changed. Migration is one find-and-replace in your using directives: the old imaging namespaces, root plus .PixelFormats / .Processing / .Formats.*, become the matching EasyImageSharp ones. Input format coverage is unchanged or wider (PNG, JPEG incl. progressive and CMYK, WebP, GIF, BMP, TIFF incl. CCITT G3/G4 and JPEG-in-TIFF, TGA, Netpbm, QOI, ICO); WebP is decode-only, so there is no WebP encoder. ImageInfo exposes FrameCount rather than a frame-metadata collection. Deskew preprocessing now uses EasyImageSharp's projection-profile deskew - same estimator, scored on ink coordinates instead of rotating the page once per candidate angle, so it is markedly faster and leaves an already-straight page untouched. See CHANGELOG.md. 2.3.0: thirteen additive capabilities. Recognition detail: RecognitionOptions.WordLevelDetail (default None) populates OcrLine.Words/Characters with true per-word and per-character geometry from the recognizer's CTC alignment, and hOCR/ALTO/TSV emit real word boxes when present. Documents: Unicode searchable PDF via an embedded Type0/CIDFontType2 Identity-H subset font with ToUnicode CMap (PdfOcrOptions.TextLayerFont/TextLayerFontPath, PdfOcrResult.TextLayerFontStatus; no font is bundled — supply one or rely on the system font probe, with graceful Latin-1 fallback); multi-frame TIFF input (ExtractTextFromFramesAsync/StreamTextFromFramesAsync); recovered tables as data (TableHtmlParser, ToRows/ToDataTable/ToCsv). New modes: handwriting recognition via TrOCR (EasyOcrServiceOptions.Handwriting + RecognizeHandwritingAsync; hosted models download on first use, int8 by default, or point at your own export) and barcode/QR reading via ZXing.Net (BarcodeScanner.ReadBarcodesAsync, combined text+barcode pass). Post-processing: redaction that permanently paints over matched text with Luhn-checked card and mod-97 IBAN presets; SymSpell-style post-OCR correction gated on confidence with IBAN/MRZ checksum repair; anchor-based field extraction with invoice presets; CER/WER accuracy metrics. Integration: streaming ExtractTextStreamAsync (IAsyncEnumerable), an easyocrsharp dotnet-tool CLI, and a Dockerized ASP.NET Core sample. All additive and opt-in; no public method renamed and no existing default changed. New dependency: ZXing.Net. 2.2.4: document-structure analysis and document preprocessing. New AnalyzeDocumentAsync (all input overloads) recovers layout regions, tables as HTML, formulas as LaTeX, seals and reading order via PP-StructureV3 (PaddleOcrNet engine), with ToMarkdown()/ToJson() export and DocumentAnalysisOptions (table model choice, per-feature toggles, languages, page orientation/unwarp). New PreprocessingOptions.Sharpen/SharpenAmount (unsharp mask), DocumentOrientation (PP-LCNet single-pass 90/180/270° page fix) and DocumentUnwarp (UVDoc page dewarp) — all default-off. All additive and opt-in; no public method or default changed. 2.2.3: dependency updates (ONNX Runtime 1.27.0, Microsoft.Extensions 10.0.9) plus the hardening + performance + accuracy pass. Security: image decompression-bomb guard (MaxImagePixels), PDF page/size guards (MaxPages, MaxPageMegapixels), HTTPS-only model source + fail-closed checksum verification, model file-name traversal check. Thread-safety: recognizer cache no longer poisoned by a cancelled/failed load; DisposeAsync drains in-flight OCR before releasing sessions. Performance: contiguous Buffer.Span tensor reads, PerspectiveWarp sub-rect copy, CPU intra-op=1 under box-level parallelism, pooled scratch buffers, new WarmUp() to remove cold-start latency. Accuracy: column/font-aware reading order, IoU NMS box de-dup, dominant-script bias on multi-language requests. Added typed exceptions (ModelDownload/Checksum/OfflineModelMissing/PdfProcessing/ImageTooLarge) and OcrResult.SourceWidth/Height. No public method renamed; all additive or safer defaults. See CHANGELOG.md. 2.2.1: clearer typed errors for bad PDFs. 2.2.0: PDF I/O, exporters, telemetry, resilient downloads, auto-GPU, EasyOCR parity.","dependency_summary":{"total_dependency_groups":1,"target_frameworks":["net10.0"],"total_dependencies":9},"verified":false},"repo_metadata":{"id":325918003,"uuid":"1103225672","full_name":"FarhanLodi/EasyOcrSharp","owner":"FarhanLodi","description":"High-accuracy, fully offline OCR for .NET powered by EasyOCR's neural models running natively on ONNX Runtime. EasyOcrSharp delivers EasyOCR-grade accuracy without Python, PyTorch, native OCR binaries, or external services, with automatic model downloads, GPU acceleration, and support for 86 languages.","archived":false,"fork":false,"pushed_at":"2026-09-14T16:11:50.000Z","size":33145,"stargazers_count":11,"open_issues_count":0,"forks_count":2,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-09-14T19:37:27.505Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"C#","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/FarhanLodi.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null},"funding":{"github":[],"patreon":null,"open_collective":null,"ko_fi":null,"tidelift":null,"community_bridge":null,"liberapay":null,"issuehunt":null,"lfx_crowdfunding":null,"polar":null,"buy_me_a_coffee":null,"thanks_dev":null,"custom":["https://www.paypal.com/paypalme/FarhanLodi"]}},"created_at":"2025-11-24T15:38:18.000Z","updated_at":"2026-09-14T16:09:14.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/FarhanLodi/EasyOcrSharp","commit_stats":null,"previous_names":["farhanlodi/easyocrsharp"],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/FarhanLodi/EasyOcrSharp","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FarhanLodi%2FEasyOcrSharp","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FarhanLodi%2FEasyOcrSharp/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FarhanLodi%2FEasyOcrSharp/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FarhanLodi%2FEasyOcrSharp/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/FarhanLodi","download_url":"https://codeload.github.com/FarhanLodi/EasyOcrSharp/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FarhanLodi%2FEasyOcrSharp/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37331643,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-15T02:00:05.939Z","response_time":113,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"repo_metadata_updated_at":"2026-09-15T20:50:42.185Z","dependent_packages_count":0,"downloads":3186,"downloads_period":"total","dependent_repos_count":0,"rankings":{"downloads":57.0496005920092,"dependent_repos_count":6.493362334851244,"dependent_packages_count":17.417129718336845,"stargazers_count":null,"forks_count":null,"docker_downloads_count":null,"average":26.986697548399096},"purl":"pkg:nuget/easyocrsharp","advisories":[],"docker_usage_url":"https://docker.ecosyste.ms/usage/nuget/easyocrsharp","docker_dependents_count":null,"docker_downloads_count":null,"usage_url":"https://repos.ecosyste.ms/usage/nuget/easyocrsharp","dependent_repositories_url":"https://repos.ecosyste.ms/api/v1/usage/nuget/easyocrsharp/dependencies","status":null,"funding_links":["https://www.paypal.com/paypalme/FarhanLodi"],"critical":null,"issue_metadata":null,"versions_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/versions","version_numbers_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/version_numbers","latest_version_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/latest_version","dependent_packages_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/dependent_packages","related_packages_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/related_packages","codemeta_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/packages/easyocrsharp/codemeta","maintainers":[{"uuid":"farhanlodi31","login":"farhanlodi31","name":null,"email":null,"url":null,"packages_count":24,"html_url":"https://www.nuget.org/profiles/farhanlodi31","role":null,"created_at":"2025-11-26T17:02:37.595Z","updated_at":"2025-11-26T17:02:37.595Z","packages_url":"https://packages.ecosyste.ms/api/v1/registries/nuget.org/maintainers/farhanlodi31/packages"}]}