Descrizione del lavoro:
Work Flexibility: Hybrid
What You will do:
* Analyze incoming order documents and map out their formats, fields, and edge cases.
* Evaluate OCR and document-extraction tools (e.g. Tesseract, PaddleOCR, layout-aware models like LayoutLM/LiLT) and help choose the right approach.
* Build a pipeline that extracts structured order data (customer, part numbers, quantities, prices, dates, PO references) from documents.
* Implement validation against master data and a review step that flags uncertain results for a human to check.
* Integrate the output with existing data infrastructure for downstream reporting and order entry.
* Document your work and support handover to the team.
What you will need:
Required Qualification:
* Working knowledge of Python (data processing, scripting).
* Masters degree in Engineering passing out in 2025 or 2026
* Interest in (or some exposure to) machine learning, computer vision, or NLP.
* Basic SQL and comfort with structured data.
* An analytical mindset and strong attention to data quality
* Basic knowledge of UI-based automations (eg. UiPath) and/or API-based automations (eg. PowerAutomate)
Preferred Qualification:
* Experience with OCR libraries or document-AI models.
* Familiarity with cloud data platforms and pipeline tooling.
* Version control (Git) and collaborative development experience.
Travel Percentage: 0
| Provenienza: | Web dell'azienda |
|---|---|
| Pubblicato il: | 25 Set 2026 (verificato il 29 Set 2026) |
| Tipo di impiego: | Stage |
| Settore: | Salute |
| Lingue: | Inglese |