> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Make image-only PDFs in the personal assistant asynchronously OCR every page in bounded batches and become searchable without truncating pages.
**Architecture:** Keep Tika as the fast path. When a PDF has no text layer, create owner-scoped OCR job/page rows and let the existing scheduled ingestion worker process at most 20 pages per claim. A focused renderer converts PDF pages to bounded JPEG images; a focused vision gateway reuses the enabled OpenAI-compatible vision/chat model. Final fragments are published only after the job reaches a terminal result.
**Tech Stack:** Java 17, Spring Boot 3.5, JdbcTemplate, PDFBox (already transitively available through Tika; declare explicitly), JUnit 5/Mockito, MySQL 8, uni-app Vue 3/TypeScript.
---
## File map
- Create `backend/ruoyi-modules/ruoyi-aihr/src/main/java/org/dromara/aihr/personal/service/PersonalPdfPageRenderer.java`: PDF page counting and bounded JPEG rendering only.
assertTrue(sql.contains("`processed_pages` int NOT NULL DEFAULT 0"));
```
- [ ]**Step 2: Run the schema test and verify RED**
Run:
```bash
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalSchemaContractTest test
```
Expected: FAIL because the OCR table strings do not exist.
- [ ]**Step 3: Add the two tables and progress DTO**
Add `aihr_personal_ocr_job` with job status/counters/lease fields and `aihr_personal_ocr_page` with page status/text/attempt fields. Extend `ItemResponse` with an optional nested record:
Also assert every rendered image is `image/jpeg`, non-empty, and below the renderer byte limit.
- [ ]**Step 2: Run renderer tests and verify RED**
```bash
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalPdfPageRendererTest test
```
Expected: test compilation fails because `PersonalPdfPageRenderer` is absent.
- [ ]**Step 3: Implement the renderer**
Declare `org.apache.pdfbox:pdfbox` explicitly at the version resolved by Tika. Implement constants `MAX_PAGES=200`, `BATCH_SIZE=20`, render at bounded DPI, scale oversized pages down, JPEG encode with a fixed quality, and return:
Test runtime resolution order (`vision` before `chat`), disabled cost guard, missing runtime, HTTP failure, and normalized OCR text. Use an injected HTTP caller instead of a real provider.
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -Dtest=PersonalVisionOcrServiceTest test
```
Expected: compilation failure because the service is absent.
- [ ]**Step 3: Implement the gateway**
Query enabled model configuration with category order `vision`, then `chat`. Build the same OpenAI-compatible multimodal request used by the existing knowledge OCR, with temperature 0 and the exact extraction prompt:
Use conditional SQL updates to claim one job. Render/recognize at most 20 pages, upsert each page result, recompute counters, and release the job to `PENDING` when pages remain. On terminal completion aggregate successful page text in page order, call the existing fragment publication path, and set `PERSONAL_OCR_PARTIAL` only when failed pages remain.
- [ ]**Step 4: Integrate with the worker**
Change only the empty-PDF branch:
```java
if(chunks.isEmpty()&&isPdf(item)){
pdfOcrService.enqueue(item,stored.bytes());
returntrue;
}
```
Schedule `processNextBatch()` on the existing personal ingestion scheduler. Ordinary PDFs and all non-PDF formats keep the current path.
- [ ]**Step 5: Run focused and full personal tests**
Assert item detail includes OCR progress, `POST /items/{id}/ocr/retry-failed` is owner-scoped, a non-OCR item rejects OCR retry, and cleanup deletes page rows before job rows.
- [ ]**Step 2: Run focused tests and verify RED**
Expected: DTO/controller/cleanup assertions fail.
- [ ]**Step 3: Implement progress/retry/cleanup**
Join OCR progress into item detail without multiplying list rows. Add:
- Modify/Create matching Vitest tests under `mobile-uni/src/**/*.spec.ts`
- [ ]**Step 1: Write failing TypeScript tests**
Assert `ocrProgressText()` returns:
```text
正在识别扫描 PDF:20/86 页
已收录,2 页识别失败
文件超过 200 页,请拆分后重新上传
```
and that failed-page retry calls `/items/{id}/ocr/retry-failed`.
- [ ]**Step 2: Run and verify RED**
```bash
npm --prefix mobile-uni run test:unit
```
Expected: tests fail because OCR fields/helpers are absent.
- [ ]**Step 3: Implement minimal UI**
Extend `PersonalItem` with optional OCR progress, show a progress bar/copy in `item.vue`, poll only while item/OCR status is active, and show “重试失败页” only when `failedPages > 0`.
- [ ]**Step 4: Run tests, typecheck and H5 build**
```bash
npm --prefix mobile-uni run test:unit
npm --prefix mobile-uni run typecheck
npm --prefix mobile-uni run build:h5
```
Expected: all commands PASS.
- [ ]**Step 5: Commit**
Commit the service, page and test files with message `feat(mobile): show scanned PDF OCR progress`.
### Task 7: Migration, real PDF smoke and documentation
**Files:**
- Modify: `scripts/personal-assistant-smoke.sh`
- Modify: `docs/个人AI助理阶段二开发推进计划.md`
- Modify: `docs/个人AI助理阶段二专项TechSpec.md`
- [ ]**Step 1: Add smoke assertions before production verification**
Extend the smoke script to assert OCR tables exist and, when `AIHR_PERSONAL_SCANNED_PDF` is set, upload that file, wait for OCR terminal state, require `READY`, run a personal-domain search against extracted text, then delete and verify OCR/OSS cleanup.
- [ ]**Step 2: Import the migration without resetting other data**
```bash
docker exec -i wygj-mysql mysql -uroot -proot --default-character-set=utf8mb4 ry-vue \
Expected: 8 pages processed, item reaches `READY`, a query hits the item, and all temporary DB/vector/OSS rows are cleaned.
- [ ]**Step 5: Browser verification**
Upload the same PDF from `/h5/#/pages/user/assistant/capture`, verify progress on item detail, final `READY`, searchable citation, then delete it and verify it disappears immediately.
- [ ]**Step 6: Update docs and commit**
Document the 20-page batch, 200-page maximum, progress states, partial success and failed-page retry. Run `git diff --check`, then commit with message `docs(personal): document scanned PDF OCR`.
### Task 8: Final verification
**Files:** none beyond prior tasks.
- [ ]**Step 1: Run complete backend module tests**
```bash
mvn -f backend/pom.xml -pl ruoyi-modules/ruoyi-aihr -am test
```
- [ ]**Step 2: Run mobile checks**
```bash
npm --prefix mobile-uni run test:unit
npm --prefix mobile-uni run typecheck
npm --prefix mobile-uni run build:h5
```
- [ ]**Step 3: Verify repository hygiene**
```bash
git diff --check
git status --short
```
Expected: no whitespace errors; only intentional uncommitted files, preferably none.
- [ ]**Step 4: Record remaining external-model boundary**
If no enabled vision/chat model is configured locally, record the real OCR gate as unverified and retain the explicit `PERSONAL_OCR_MODEL_UNAVAILABLE` behavior. Do not substitute fake OCR text.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.