Files
prop-ai-hr/docs/superpowers/plans/2026-07-12-personal-scanned-pdf-ocr.md

17 KiB

Personal Scanned PDF OCR Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Make image-only PDFs in the personal assistant asynchronously OCR every page in bounded batches and become searchable without truncating pages.

Architecture: Keep Tika as the fast path. When a PDF has no text layer, create owner-scoped OCR job/page rows and let the existing scheduled ingestion worker process at most 20 pages per claim. A focused renderer converts PDF pages to bounded JPEG images; a focused vision gateway reuses the enabled OpenAI-compatible vision/chat model. Final fragments are published only after the job reaches a terminal result.

Tech Stack: Java 17, Spring Boot 3.5, JdbcTemplate, PDFBox (already transitively available through Tika; declare explicitly), JUnit 5/Mockito, MySQL 8, uni-app Vue 3/TypeScript.


File map

  • Create backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfPageRenderer.java: PDF page counting and bounded JPEG rendering only.
  • Create backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalVisionOcrService.java: resolve enabled vision/chat runtime and call OpenAI-compatible image OCR.
  • Create backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfOcrService.java: owner-scoped job/page lifecycle, 20-page claims, retry, aggregation.
  • Modify PersonalIngestionWorker.java: retain Tika fast path; hand empty PDFs to PersonalPdfOcrService.
  • Modify PersonalAssistantDto.java, PersonalSpaceService.java, PersonalAssistantController.java: expose progress and retry-failed-pages contract.
  • Modify PersonalCleanupService.java: remove OCR page/job rows when deleting a personal item.
  • Modify backend/script/sql/aihr_personal_knowledge_mysql8.sql: add OCR job/page tables.
  • Modify mobile-uni/src/services/personal-assistant.ts and mobile-uni/src/pages/user/assistant/item.vue: show progress and retry failed pages.
  • Add focused tests beside existing personal assistant tests.

Task 1: Schema and API contract

Files:

  • Modify: backend/script/sql/aihr_personal_knowledge_mysql8.sql

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/domain/PersonalAssistantDto.java

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalSchemaContractTest.java

  • Step 1: Write the failing schema test

Add assertions that the SQL contains both OCR tables, the owner isolation keys, page uniqueness, and the 200-page progress columns:

assertTrue(sql.contains("CREATE TABLE IF NOT EXISTS `aihr_personal_ocr_job`"));
assertTrue(sql.contains("CREATE TABLE IF NOT EXISTS `aihr_personal_ocr_page`"));
assertTrue(sql.contains("UNIQUE KEY `uk_personal_ocr_job_item` (`tenant_id`, `owner_user_id`, `item_id`)"));
assertTrue(sql.contains("UNIQUE KEY `uk_personal_ocr_page_number` (`tenant_id`, `owner_user_id`, `item_id`, `page_number`)"));
assertTrue(sql.contains("`processed_pages` int NOT NULL DEFAULT 0"));
  • Step 2: Run the schema test and verify RED

Run:

mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalSchemaContractTest test

Expected: FAIL because the OCR table strings do not exist.

  • Step 3: Add the two tables and progress DTO

Add aihr_personal_ocr_job with job status/counters/lease fields and aihr_personal_ocr_page with page status/text/attempt fields. Extend ItemResponse with an optional nested record:

public record OcrProgressResponse(boolean required, String status, int totalPages,
                                  int processedPages, int successPages, int failedPages,
                                  List<Integer> failedPageNumbers) {}

Append OcrProgressResponse ocr to ItemResponse so absence remains null for non-OCR items.

  • Step 4: Run the schema test and verify GREEN

Run the same Maven command. Expected: PASS.

  • Step 5: Commit
git add backend/script/sql/aihr_personal_knowledge_mysql8.sql \
  backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/domain/PersonalAssistantDto.java \
  backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalSchemaContractTest.java
git commit -m "feat(personal): add scanned PDF OCR schema"

Task 2: Bounded PDF page renderer

Files:

  • Modify: backend/ruoyi-modules/ruoyi-aihr/pom.xml

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfPageRenderer.java

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalPdfPageRendererTest.java

  • Step 1: Write failing renderer tests

Create in-memory PDFs with PDFBox and assert:

assertEquals(8, renderer.pageCount(eightPagePdf));
assertEquals(List.of(0, 1), renderer.render(eightPagePdf, 0, 2).stream().map(RenderedPage::pageIndex).toList());
assertThrows(PdfPageLimitException.class, () -> renderer.requireSupportedPageCount(201));

Also assert every rendered image is image/jpeg, non-empty, and below the renderer byte limit.

  • Step 2: Run renderer tests and verify RED
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalPdfPageRendererTest test

Expected: test compilation fails because PersonalPdfPageRenderer is absent.

  • Step 3: Implement the renderer

Declare org.apache.pdfbox:pdfbox explicitly at the version resolved by Tika. Implement constants MAX_PAGES=200, BATCH_SIZE=20, render at bounded DPI, scale oversized pages down, JPEG encode with a fixed quality, and return:

public record RenderedPage(int pageIndex, byte[] bytes, String mimeType) {}

Reject malformed/encrypted PDFs using controlled PdfRenderException codes; never log PDF bytes or extracted content.

  • Step 4: Run renderer tests and module tests

Expected: renderer tests PASS; existing parser tests remain PASS.

  • Step 5: Commit
git add backend/ruoyi-modules/ruoyi-aihr/pom.xml \
  backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfPageRenderer.java \
  backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalPdfPageRendererTest.java
git commit -m "feat(personal): render bounded PDF OCR pages"

Task 3: Shared vision OCR boundary

Files:

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalVisionOcrService.java

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalVisionOcrServiceTest.java

  • Step 1: Write failing gateway tests

Test runtime resolution order (vision before chat), disabled cost guard, missing runtime, HTTP failure, and normalized OCR text. Use an injected HTTP caller instead of a real provider.

assertEquals("第一条\n第二条", service.recognize(jpeg, "image/jpeg", 3));
assertThrows(OcrUnavailableException.class, () -> disabledService.recognize(jpeg, "image/jpeg", 3));
  • Step 2: Run and verify RED
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalVisionOcrServiceTest test

Expected: compilation failure because the service is absent.

  • Step 3: Implement the gateway

Query enabled model configuration with category order vision, then chat. Build the same OpenAI-compatible multimodal request used by the existing knowledge OCR, with temperature 0 and the exact extraction prompt:

忠实提取本页全部可见文字,保留标题、段落和表格行顺序;不要总结、解释或补写。无可识别文字时返回空字符串。

Honor AIHR_AI_RUNTIME_ENABLED and AIHR_AI_CHAT_ENABLED; bound connect/request timeouts and response bytes.

  • Step 4: Run and verify GREEN

Run the focused test. Expected: PASS.

  • Step 5: Commit
git add backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalVisionOcrService.java \
  backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalVisionOcrServiceTest.java
git commit -m "feat(personal): add vision OCR gateway"

Task 4: OCR job orchestration and publication

Files:

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfOcrService.java

  • Create: backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalPdfOcrServiceTest.java

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalIngestionWorker.java

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalIngestionWorkerTest.java

  • Step 1: Write failing orchestration tests

Cover these exact behaviors with owner-scoped SQL verification:

// Empty PDF text creates a job instead of PERSONAL_PARSE_EMPTY.
assertTrue(worker.processNext());
verify(ocr).enqueue(eq(item), any(byte[].class));

// One claim never exceeds 20 pages.
assertEquals(20, service.claimBatch(jobId).pageNumbers().size());

// Partial success publishes only after terminal aggregation.
assertEquals("READY", terminalItemStatus);
assertEquals(List.of(4, 7), failedPageNumbers);

Add tests for 201 pages, all pages failing, process interruption, idempotent page upsert, and max three attempts.

  • Step 2: Run tests and verify RED

Run both focused test classes. Expected: failures because the OCR orchestration API does not exist and the worker still emits PERSONAL_PARSE_EMPTY.

  • Step 3: Implement job lifecycle

Implement owner-scoped methods:

void enqueue(Item item, byte[] pdfBytes);
boolean processNextBatch();
OcrProgressResponse progress(PersonalOwner owner, long itemId);
OcrProgressResponse retryFailedPages(PersonalOwner owner, long itemId);

Use conditional SQL updates to claim one job. Render/recognize at most 20 pages, upsert each page result, recompute counters, and release the job to PENDING when pages remain. On terminal completion aggregate successful page text in page order, call the existing fragment publication path, and set PERSONAL_OCR_PARTIAL only when failed pages remain.

  • Step 4: Integrate with the worker

Change only the empty-PDF branch:

if (chunks.isEmpty() && isPdf(item)) {
    pdfOcrService.enqueue(item, stored.bytes());
    return true;
}

Schedule processNextBatch() on the existing personal ingestion scheduler. Ordinary PDFs and all non-PDF formats keep the current path.

  • Step 5: Run focused and full personal tests
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest='Personal*Test' test

Expected: all personal tests PASS.

  • Step 6: Commit
git add backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfOcrService.java \
  backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalIngestionWorker.java \
  backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalPdfOcrServiceTest.java \
  backend/ruoyi-modules/ruoyi-aihr/src/test/java/org/dromara/aihr/personal/PersonalIngestionWorkerTest.java
git commit -m "feat(personal): process scanned PDFs in OCR batches"

Task 5: Progress, retry and deletion contracts

Files:

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalSpaceService.java

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/controller/PersonalAssistantController.java

  • Modify: backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalCleanupService.java

  • Modify tests: PersonalSpaceServiceTest.java, PersonalAssistantControllerTest.java, PersonalCleanupServiceTest.java

  • Step 1: Write failing contract tests

Assert item detail includes OCR progress, POST /items/{id}/ocr/retry-failed is owner-scoped, a non-OCR item rejects OCR retry, and cleanup deletes page rows before job rows.

  • Step 2: Run focused tests and verify RED

Expected: DTO/controller/cleanup assertions fail.

  • Step 3: Implement progress/retry/cleanup

Join OCR progress into item detail without multiplying list rows. Add:

@PostMapping("/items/{id}/ocr/retry-failed")
public R<OcrProgressResponse> retryFailedOcrPages(@PathVariable long id) {
    return R.ok(pdfOcrService.retryFailedPages(owner(), id));
}

Delete aihr_personal_ocr_page then aihr_personal_ocr_job in the existing cleanup transaction.

  • Step 4: Run focused tests and verify GREEN

Expected: all three focused test classes PASS.

  • Step 5: Commit

Commit backend contract and cleanup files with message feat(personal): expose OCR progress and retry.

Task 6: Mobile progress UI

Files:

  • Modify: mobile-uni/src/services/personal-assistant.ts

  • Modify: mobile-uni/src/pages/user/assistant/item.vue

  • Modify/Create matching Vitest tests under mobile-uni/src/**/*.spec.ts

  • Step 1: Write failing TypeScript tests

Assert ocrProgressText() returns:

正在识别扫描 PDF:20/86 页
已收录,2 页识别失败
文件超过 200 页,请拆分后重新上传

and that failed-page retry calls /items/{id}/ocr/retry-failed.

  • Step 2: Run and verify RED
npm --prefix mobile-uni run test:unit

Expected: tests fail because OCR fields/helpers are absent.

  • Step 3: Implement minimal UI

Extend PersonalItem with optional OCR progress, show a progress bar/copy in item.vue, poll only while item/OCR status is active, and show “重试失败页” only when failedPages > 0.

  • Step 4: Run tests, typecheck and H5 build
npm --prefix mobile-uni run test:unit
npm --prefix mobile-uni run typecheck
npm --prefix mobile-uni run build:h5

Expected: all commands PASS.

  • Step 5: Commit

Commit the service, page and test files with message feat(mobile): show scanned PDF OCR progress.

Task 7: Migration, real PDF smoke and documentation

Files:

  • Modify: scripts/personal-assistant-smoke.sh

  • Modify: docs/个人AI助理阶段二开发推进计划.md

  • Modify: docs/个人AI助理阶段二专项TechSpec.md

  • Step 1: Add smoke assertions before production verification

Extend the smoke script to assert OCR tables exist and, when AIHR_PERSONAL_SCANNED_PDF is set, upload that file, wait for OCR terminal state, require READY, run a personal-domain search against extracted text, then delete and verify OCR/OSS cleanup.

  • Step 2: Import the migration without resetting other data
docker exec -i wygj-mysql mysql -uroot -proot --default-character-set=utf8mb4 ry-vue \
  < backend/script/sql/aihr_personal_knowledge_mysql8.sql
  • Step 3: Run backend and ordinary smoke regression
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest='Personal*Test' test
AIHR_PERSONAL_API_URL=https://personal-assistant-phase2.wygj-api.localhost \
  ./scripts/personal-assistant-smoke.sh

Expected: all personal tests and existing TEXT/PDF/URL smoke pass.

  • Step 4: Run the real scanned PDF gate
AIHR_PERSONAL_API_URL=https://personal-assistant-phase2.wygj-api.localhost \
AIHR_PERSONAL_SCANNED_PDF='/Users/yuanjiantsui/workspace/项目-物业AI/补充资料/关于修订证书管理办法的通知.pdf' \
  ./scripts/personal-assistant-smoke.sh

Expected: 8 pages processed, item reaches READY, a query hits the item, and all temporary DB/vector/OSS rows are cleaned.

  • Step 5: Browser verification

Upload the same PDF from /h5/#/pages/user/assistant/capture, verify progress on item detail, final READY, searchable citation, then delete it and verify it disappears immediately.

  • Step 6: Update docs and commit

Document the 20-page batch, 200-page maximum, progress states, partial success and failed-page retry. Run git diff --check, then commit with message docs(personal): document scanned PDF OCR.

Task 8: Final verification

Files: none beyond prior tasks.

  • Step 1: Run complete backend module tests
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -am test
  • Step 2: Run mobile checks
npm --prefix mobile-uni run test:unit
npm --prefix mobile-uni run typecheck
npm --prefix mobile-uni run build:h5
  • Step 3: Verify repository hygiene
git diff --check
git status --short

Expected: no whitespace errors; only intentional uncommitted files, preferably none.

  • Step 4: Record remaining external-model boundary

If no enabled vision/chat model is configured locally, record the real OCR gate as unverified and retain the explicit PERSONAL_OCR_MODEL_UNAVAILABLE behavior. Do not substitute fake OCR text.