Intelligent document processing and data structuring for higher education
Mandacaru is an intelligent document-processing application developed in the context of the Information Technology Superintendency (STI) of the Federal University of Sergipe (UFS). It was designed to transform institutional PDF documents into structured, reusable information through a reusable processing core, a web interface, assisted extraction, validation workflows, and engineering documentation.
The product story is best understood as Motivation -> Documentation -> Product -> Impact. The repository shows a real product-development and engineering effort centered on intelligent document extraction and structured data delivery.
Mandacaru addresses the friction of working with institutional documents created primarily for human reading. Resolutions, ordinances, normative instructions, and undergraduate course pedagogical projects can contain important identifiers, dates, issuing bodies, decisions, signatures, and full text, but those fields are difficult to reuse when they remain trapped in PDF layouts.
The application was developed for the STI/UFS context. This establishes the institutional problem environment; it does not imply university-wide adoption, production scale, official endorsement, or measured operational results beyond the evidence documented here.
| Motivation | Documentation | Product | Impact |
|---|---|---|---|
| Make institutional PDFs easier to process and reuse. | Requirements, architecture, manuals, validation, issues, and benchmarks. | Upload PDFs, classify documents, extract fields, inspect tables, and export data. | Demonstrated structured-processing flows and technical evaluation; potential reduction of manual document work. |
The public portfolio uses Mandacaru's official visual identity: a green-led palette, restrained neutral surfaces, and a clear typographic hierarchy for technical reading.
| Token | Hex | Role |
|---|---|---|
| Mandacaru Green | #1F7A4F |
Primary brand and links |
| Deep Green | #0F3D2E |
Strong contrast and headings |
| Medium Green | #3CB371 |
Secondary emphasis |
| Near Black | #1F1F1F |
Primary text |
| Light Gray | #E5E7EB |
Dividers and borders |
| Background Gray | #F7F9F8 |
Soft surfaces |
| Institutional Yellow | #F2C94C |
Focus and calls to action |
| Mandacaru Orange | #FF751F |
Alerts and emphasis |
Typography follows the media kit's pairing: IBM Plex Sans for headings and Inter for body copy and interface-oriented text.
Mandacaru was developed as an application in the context of STI - Information Technology Superintendency, Federal University of Sergipe (UFS). The project materials frame the need around institutional information that must be organized, processed, and made more accessible to people and systems.
The public case study uses this context carefully: it explains why the problem mattered without claiming that Mandacaru was deployed across the university or that it produced unmeasured institutional gains.
Administrative and academic teams may need to read official documents and manually transfer relevant fields into spreadsheets or other systems. That workflow is repetitive, difficult to scale, and vulnerable to transcription inconsistencies.
Mandacaru was designed to reduce this friction by combining file validation, text and metadata extraction, document classification, schema-guided field extraction, and structured export.
The project was developed through a sequence of product and engineering activities:
flowchart LR
A[Institutional problem] --> B[Requirements and product framing]
B --> C[Architecture and technical specifications]
C --> D[Reusable processing core]
D --> E[Product interface]
E --> F[Validation and issue iteration]
F --> G[Benchmarking and deployment documentation]
G --> H[Documented product case study]
The repository includes evidence of problem framing, value proposition work, requirements, UML views, architecture, coding standards, a user manual, deployment material, validation scripts, issue records, integration reports, release demonstrations, and model benchmarking.
The user-facing workflow is intentionally direct:
- Upload one or more PDF documents.
- Let the system validate the input and extract available text and metadata.
- Review the detected document type and processing status.
- Inspect structured results in a table grouped by type.
- Download CSV, XLSX, or JSONL output for further use.
The project documentation describes support for resolutions, ordinances, normative instructions, undergraduate course pedagogical projects, and a generic metadata fallback. The interface also exposes progress information and processing logs for multi-file workflows.
The following sanitized screenshots show the documented product flow with generic sample data: upload documents, process a batch, inspect structured records, and export the result.
1. Upload documents![]() |
2. Process a batch![]() |
3. Review structured data![]() |
4. Export a spreadsheet![]() |
flowchart LR
A[PDF upload] --> B[PDF signature validation]
B --> C[Text extraction]
B --> D[Metadata extraction]
C --> E[Hybrid document classification]
D --> F[Metadata record]
E --> G[Type-specific extraction schema]
C --> G
G --> H[Assisted field extraction]
H --> I[Structured record]
F --> I
I --> J[Review table and logs]
I --> K[CSV / XLSX / JSONL]
- Ingestion: accepts PDF paths or in-memory file bytes.
- Validation: checks the PDF signature and skips invalid inputs.
- Text and metadata: extracts readable document content and available file metadata.
- Classification: combines document text, file characteristics, and rule-based signals to identify the document type.
- Extraction: selects document-specific fields and examples for assisted extraction.
- Structuring: combines extracted fields and metadata into structured records.
- Delivery: displays grouped tables and enables CSV, Excel, and line-delimited JSON downloads.
The documented engineering model separates a reusable processing core from the application layer. The architecture material also describes an API-oriented integration surface and optional external persistence.
flowchart TB
U[User or external client] --> W[Product interface]
W --> L[Reusable Mandacaru processing core]
A[API integration boundary] --> L
L --> M[File management]
M --> P[Preprocessing]
P --> C[Document classification]
C --> X[Schema-guided extraction]
X --> O[Structured output]
O --> W
O --> E[CSV / XLSX / JSONL]
O -. optional documented extension .-> DB[External persistence]
The architecture follows a modular data-flow approach. It separates preprocessing, classification, extraction, structuring, and presentation so that the processing core can be reused independently of the interface.
The public product surface focuses on intelligent document extraction and structured data delivery. Its core value comes from turning institutional PDF content into organized records that can be inspected, downloaded, and consumed by other workflows.
The public documentation is organized around the most useful evidence rather than exposing the raw delivery archive.
| Resource | What it demonstrates |
|---|---|
| User Manual | User journey, supported document types, outputs, and practical limitations. |
| Architecture Overview | Components, data flow, boundaries, and architectural decisions. |
| API and Deployment Guide | Integration model, deployment assumptions, configuration, and operational concerns. |
| Model Benchmarking | Evaluation criteria, recorded comparison, interpretation, and engineering value. |
| Product Development Notes | Evidence that the work progressed from problem framing to product development. |
| Engineering Decisions | Why the project used modular processing, type-specific schemas, and source-grounded extraction. |
The underlying implementation and original technical records are kept outside this curated public portfolio layer. This repository focuses on the product flow, documented decisions, sanitized evidence, and user-facing behavior.
The repository contains a benchmark of local language models for PDF-to-CSV extraction. It evaluates:
- CSV validity.
- Header and row conformity.
- Recorded p50 and p95 processing time.
- Trade-offs between response quality, latency, and implementation complexity.
| Model | Recorded p50 / p95 | Valid CSV | Header and row checks |
|---|---|---|---|
| Alternative A | 176.12 s / 176.12 s | Yes | Passed |
| Alternative B | 445.30 s / 445.30 s | Yes | Passed |
| Alternative C | 187.50 s / 187.50 s | No | Failed |
The engineering problem was how to obtain structured fields from extracted PDF text while preserving a strict output contract. Alternative extraction approaches were evaluated against structural validity and latency. The recorded snapshot did not establish a production-ready winner: one alternative produced the fastest valid output, another produced valid formatting at higher latency, and another did not satisfy the output contract. The result supports treating model-assisted integration as an evaluated engineering path rather than a guaranteed production capability.
These numbers describe a repository benchmark snapshot, not a production SLA or statistically representative evaluation. Raw documents, reference spreadsheets, prompts, and credentials remain private.
The documented design reflects several practical decisions:
- Reusable core: keep document-processing logic independent from the user interface.
- Type-specific schemas: use different extraction fields and examples for different institutional document types.
- Source-grounded extraction: instruct the extraction layer to preserve source wording, avoid invented values, and return structured fields.
- Interoperable outputs: produce CSV, XLSX, and JSONL so downstream consumers are not tied to one interface.
- Defensive input handling: validate PDF signatures and report file-level processing errors.
- Operational feedback: expose progress and logs during multi-file processing.
- Conservative capability claims: distinguish implemented extraction from unimplemented semantic retrieval and RAG features.
Mandacaru demonstrates a broader development process than an isolated model call:
Problem definition
-> requirements and product framing
-> research and model benchmarking
-> architecture and software specifications
-> reusable processing core
-> user interface and export workflow
-> validation, issues, and integration records
-> product documentation and deployment guidance
The repository evidence includes a value proposition, project planning artifacts, requirements and use cases, UML diagrams, coding standards, architecture, an MVP summary, user and deployment manuals, validation scripts, issue records, integration reports, release material, and technical benchmarking.
This public portfolio describes what the implementation does without exposing private stack choices.
Application surface
A web-based interface supports multi-file upload, progress feedback, tabular review, and download actions.
Document and data processing
The processing layer validates files, extracts readable content and metadata, classifies document types, applies document-specific schemas, and prepares structured records.
Assisted extraction
The extraction layer uses configured model-assisted processing with document-specific examples and structured output fields.
Engineering practices
Modular boundaries, documented interfaces, coding standards, static-analysis configuration, structured logs, protected secret configuration, and concurrent multi-file processing in the interface.
.
├── README.md
├── assets/
│ ├── dcomp-ufs.png
│ ├── mandacaru-header.png
│ ├── sti-ufs-logo.png
│ ├── ufs-logo.png
│ └── demo/
│ ├── 01-upload-empty.png
│ ├── 02-processing.png
│ ├── 03-structured-data.png
│ └── 04-exported-spreadsheet.png
└── docs/
├── README.md
├── api-and-deployment.md
├── architecture.md
├── engineering-decisions.md
├── model-benchmarking.md
├── product-development.md
└── user-manual.md
This is a curated public portfolio repository. The implementation source remains in the original public project reference, while this repository keeps only English documentation and a safe brand asset.
I was one of the people responsible for developing Mandacaru as a product. Repository evidence supports involvement across:
- Model benchmarking and technical evaluation.
- Product and engineering documentation.
- Architecture, requirements, and development standards as part of the team effort.
- Validation and issue-driven iteration.
- AI-assisted document processing and product communication.
These statements describe documented collaboration. They do not assign sole authorship, leadership, or ownership of components without direct evidence.
- A documented PDF-processing workflow exists from input validation to structured export.
- The product supports multiple institutional document categories and a generic fallback path.
- A user-facing workflow supports multi-file upload, processing feedback, tabular inspection, and downloads.
- A reusable processing core separates document processing from the interface layer.
- Model alternatives were evaluated using output validity, schema checks, and recorded latency.
- Deployment, API integration, architecture, validation, and user documentation were produced as part of the development effort.
The implemented workflow was designed to reduce the friction of locating, copying, organizing, and reusing information found in institutional PDF collections. Potential benefits include more accessible structured records, less repetitive manual document work, and easier integration with spreadsheets, dashboards, databases, or public-information workflows.
These are logical product benefits, not measured STI/UFS outcomes. The repository does not provide evidence for adoption numbers, time savings, cost savings, or university-wide deployment.
This public repository intentionally excludes:
- Videos, audio, and raw media.
- Raw benchmark inputs and outputs.
- Spreadsheets containing test or issue data.
- Personal information and contributor-specific delivery exports.
- Internal URLs, server addresses, credentials, tokens, secret configuration, and stack-specific implementation details.
- Untranslated source documents and private institutional material.
Configuration examples use conceptual descriptions only. Secret values must be supplied through protected environment configuration and must never be committed.
This repository is a curated documentation and portfolio surface. Runtime setup details, private deployment addresses, and implementation-specific configuration are intentionally omitted from the public version.
Access to any runnable deployment depends on the project owner and the target environment.







