Conversation
…r#4448) When tesseract is built with --disable-legacy and the requested traineddata contains no LSTM component, init_tesseract_lang_data previously printed a warning and downgraded the engine mode to OEM_TESSERACT_ONLY before returning true. The legacy engine code is compiled out in this build, so no recognizer is actually loaded. Recognition then runs with no engine, leaving best_choice NULL on every word, and pass-1 post-processing in recog_all_words segfaults when it dereferences best_choice (control.cpp:356). Under DISABLED_LEGACY_ENGINE there is no legacy fallback to take, so return false from init_tesseract_lang_data with a clearer error message that names the offending tessdata path. The caller in init_tesseract already handles a per-language load failure by logging "Failed loading language '%s'" and, if no language loads at all, "Tesseract couldn't load any languages!", so multi-language inputs (-l eng+gsta) degrade gracefully. The default build (legacy enabled) is unchanged: the existing fallback to OEM_TESSERACT_ONLY remains in place.
Up to standards ✅🟢 Issues
|
| Metric | Results |
|---|---|
| Complexity | 0 |
| Duplication | 0 |
NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.
Brief comment explaining the contract with init_tesseract: a per-language load failure is logged by the caller via "Failed loading language '%s'", and if no language loads at all the existing "Tesseract couldn't load any languages!" path triggers. No behavior change.
|
@gaurav0107 thanks for fixing this! I'm not so familiar with Tesseract code so I didn't understand how to take amido's recommendation info account. :) |
|
Hi, Did you check it works as intended? You also should check the multi tessdata scenario, where one of them has lstm and the other one does not. |
|
@amitdo Thanks for the review. Yes, I tested the single-language case in a You are right that I have not yet exercised the multi-tessdata scenario (one language with LSTM, one without) in the same build. I will run that case and report back before this is merged; the expectation is that the LSTM-capable language continues to load and only the no-LSTM one is skipped. |
|
Gentle ping — I believe this is ready: the review comments are addressed (the single-language |
@gaurav0107, did you run that case? What was the result? |
|
Sorry, I missed your comment. |
|
I don't think he tested this case. I didn't test it either, but still decided to merge it. |
|
That's okay. I think that we are getting an increasing number of pull requests which were made by or with the help of AI bots, this one probably, too. Ideally this help should be indicated in the commit message. |
tesseract 5.5.3 Created-by: HarmonybrewBot Commit-by: HarmonybrewBot Merged-by: HarmonybrewBot Description: Created by `brew bump` --- Created with `brew bump-formula-pr`. - [ ] `resource` blocks have been checked for updates. <details> <summary>release notes</summary> <pre>## What's Changed * Fix typos by @szepeviktor in tesseract-ocr/tesseract#4497 * Fix missing closing tags in multi-page PAGE XML output by @stweil with @Copilot in tesseract-ocr/tesseract#4506 * Fix Apple CMAKE_SYSTEM_PROCESSOR not set when crosscompiling by @ppavacic in tesseract-ocr/tesseract#4503 * Remove outdated docker-compose.yml (Fixes #4492) by @Mohataseem89 in tesseract-ocr/tesseract#4509 * docs: document memory ownership and lifecycle in C-API by @markbus-ai in tesseract-ocr/tesseract#4511 * Add Portuguese language support and section descriptions (Windows installer) by @eduardomozart in tesseract-ocr/tesseract#4516 * Update Slovak language string in tesseract.nsi by @eduardomozart in tesseract-ocr/tesseract#4518 * fix(nsi): mask LANGID to 16-bit for reliable auto-detection by @eduardomozart in tesseract-ocr/tesseract#4513 * Add GitHub Copilot instructions by @stweil with @Copilot in tesseract-ocr/tesseract#4508 * Correct mutex call to prevent multiple instances by @eduardomozart in tesseract-ocr/tesseract#4517 * Update versions of GitHub actions by @stweil in tesseract-ocr/tesseract#4522 * Bump microsoft/setup-msbuild from 2 to 3 by @dependabot[bot] in tesseract-ocr/tesseract#4533 * Use asciidoctor instead of asciidoc-py for manpage generation by @amitdo with @Copilot in tesseract-ocr/tesseract#4534 * Bump actions/upload-artifact from 4 to 7 by @dependabot[bot] in tesseract-ocr/tesseract#4545 * autotools: Fix linker warning on macOS by @stweil in tesseract-ocr/tesseract#4559 * Fix compiler warning on macOS by @stweil in tesseract-ocr/tesseract#4558 * Cmake cleanup by @zdenop in tesseract-ocr/tesseract#4562 * Remove SW_BUILD option from CMake and CI by @amitdo with @Copilot in tesseract-ocr/tesseract#4566 * Modernize code by @stweil in tesseract-ocr/tesseract#4561 * Remove some files by @stweil in tesseract-ocr/tesseract#4569 * Modernize more code by @stweil in tesseract-ocr/tesseract#4568 * Fix compiler warnings (-Wold-style-cast) by @stweil in tesseract-ocr/tesseract#4571 * Fix some compiler warnings (-Wunused-parameter) by @stweil in tesseract-ocr/tesseract#4570 * ci: Improve installer for windows by @stweil in tesseract-ocr/tesseract#4572 * Fix several issues reported by Codacy (including real bugs) by @stweil in tesseract-ocr/tesseract#4573 * Bump actions/checkout from 6 to 7 by @dependabot[bot] in tesseract-ocr/tesseract#4574 * Fix crash when LSTM is missing in disabled-legacy build (#4448) by @gaurav0107 in tesseract-ocr/tesseract#4563 * ci: Remove unused ilammy/setup-nasm by @stweil in tesseract-ocr/tesseract#4575 * autotools: Simplify Makefile rules by @stweil in tesseract-ocr/tesseract#4579 * Fix memory-safety issues in .traineddata deserialization by @stweil in tesseract-ocr/tesseract#4581 * Handle send() failures in SVNetwork::Flush by @Ramya-9353 in tesseract-ocr/tesseract#4576 * Small code improvements by @stweil in tesseract-ocr/tesseract#4585 * doc: Add comprehensive PARAMETERS section to the tesseract man page by @stweil with @Copilot in tesseract-ocr/tesseract#4526 * Fix integer overflow in LSTM Convolve and Reconfig deserialization by @stweil in tesseract-ocr/tesseract#4588 ## New Contributors * @szepeviktor made their first contribution in tesseract-ocr/tesseract#4497 * @ppavacic made their first contribution in tesseract-ocr/tesseract#4503 * @Mohataseem89 made their first contribution in tesseract-ocr/tesseract#4509 * @markbus-ai made their first contribution in tesseract-ocr/tesseract#4511 * @eduardomozart made their first contribution in tesseract-ocr/tesseract#4516 * @gaurav0107 made their first contribution in tesseract-ocr/tesseract#4563 * @Ramya-9353 made their first contribution in tesseract-ocr/tesseract#4576 **Full Changelog**: https://github.com/tesseract-ocr/tesseract/compare/5.5.2...5.5.3</pre> <p>View the full release notes at <a href="https://github.com/tesseract-ocr/tesseract/releases/tag/5.5.3">https://github.com/tesseract-ocr/tesseract/releases/tag/5.5.3</a>.</p> </details> <hr> See merge request: Harmonybrew/homebrew-core!15091
Summary
Fixes #4448. When tesseract is compiled with
--disable-legacyand theuser supplies a
.traineddatafile that does not contain an LSTMcomponent (e.g. the
gsta.traineddataattached to the issue), tesseractcrashes with a segmentation fault during pass-1 post-processing.
Root cause
Tesseract::init_tesseract_lang_data(src/ccmain/tessedit.cpp:171–178)checks whether an LSTM component is available. If not, it prints
"Error: LSTM requested, but not present!! Loading tesseract."anddowngrades
tessedit_ocr_engine_modetoOEM_TESSERACT_ONLY. UnderDISABLED_LEGACY_ENGINEthe legacy engine code is compiled out, so this"fallback" never actually loads a recognizer — but the function still
returns
true. Recognition runs with no engine, every word ends up withbest_choice == nullptr, andrecog_all_wordssegfaults atcontrol.cpp:356when it dereferencesbest_choice->permuter().Reproduction (from issue #4448)
The valgrind trace pinpoints the NULL deref at
tesseract::WERD_CHOICE::permuter() const (ratngs.h:332)reached fromrecog_all_words (control.cpp:356).Fix
Inside the existing
elsebranch (LSTM unavailable), split behavior bybuild mode:
DISABLED_LEGACY_ENGINE: emit a clearer error message thatnames the tessdata path, and
return false. The caller(
init_tesseract) already handles a per-language load failure with"Failed loading language '%s'", and if no language loads at all theexisting
"Tesseract couldn't load any languages!"path triggers, somulti-language inputs (
-l eng+gsta) degrade gracefully.OEM_TESSERACT_ONLYis unchanged.The diff is 8 added lines, 0 removed, confined to one TU.
Relationship to PR #4449
@sebras opened #4449 with a one-line NULL guard at the crash site in
control.cpp. This PR addresses the deeper fix that @amitdo describedin the #4448 thread: refuse to claim a successful init when no engine
can run. The two changes are orthogonal — #4449 is defense in depth at
the consumer side, this PR fixes the producer side. Either can land
first; landing both is fine.
Test plan
git diffreview — change is minimal and confined to one#ifdef-guarded branch.unittest(default build) — behavior unchanged in the#ifndef DISABLED_LEGACY_ENGINEbranch; existing tests pass.unittest-disablelegacy(autotools,--disable-legacy) —exercises the new early-return path; existing tests use tessdata
with LSTM, so they do not trip it.
cmake,cifuzz,CodeQL,Codacy— single conditionalreturn; no new flagged patterns expected.
gsta.traineddata, the segfault isreplaced by a clean error message and a non-zero exit code.