Descripción del puesto:
Work Flexibility: Hybrid
What You will do:
* Analyze incoming order documents and map out their formats, fields, and edge cases.
* Evaluate OCR and document-extraction tools (e.g. Tesseract, PaddleOCR, layout-aware models like LayoutLM/LiLT) and help choose the right approach.
* Build a pipeline that extracts structured order data (customer, part numbers, quantities, prices, dates, PO references) from documents.
* Implement validation against master data and a review step that flags uncertain results for a human to check.
* Integrate the output with existing data infrastructure for downstream reporting and order entry.
* Document your work and support handover to the team.
What you will need:
Required Qualification:
* Working knowledge of Python (data processing, scripting).
* Masters degree in Engineering passing out in 2025 or 2026
* Interest in (or some exposure to) machine learning, computer vision, or NLP.
* Basic SQL and comfort with structured data.
* An analytical mindset and strong attention to data quality
* Basic knowledge of UI-based automations (eg. UiPath) and/or API-based automations (eg. PowerAutomate)
Preferred Qualification:
* Experience with OCR libraries or document-AI models.
* Familiarity with cloud data platforms and pipeline tooling.
* Version control (Git) and collaborative development experience.
Travel Percentage: 0
| Origen: | Web de la compañía |
|---|---|
| Publicado: | 25 Sep 2026 (comprobado el 29 Sep 2026) |
| Tipo de oferta: | Prácticas |
| Sector: | Salud |
| Idiomas: | Inglés |