المهام الوظيفية الأساسية: تصميم وتطوير وصيانة أنابيب استيعاب ومعالجة المستندات القابلة للتوسع لدعم إدارة المعرفة في المؤسسات والتطبيقات المدعومة بالذكاء الاصطناعي. بناء وتحسين مسارات عمل البيانات متعددة المراحل لاستخراج المحتوى وتحويله وإثرائه وفهرسته من صيغ مستندات متنوعة تشمل PDF وPPTX وDOCX وCSV والمصادر القائمة على الصور. تطبيق قدرات معالجة المستندات المتقدمة، بما في ذلك استخراج البيانات الوصفية، وتجزئة المحتوى، والتصنيف، والتلخيص، ومسارات عمل استرجاع المعلومات. إدارة استراتيجيات الفهرسة لقواعد البيانات المتجهية وتحسين البحث لضمان أداء استرجاع عالي الجودة وملاءمة دقيقة. تطوير وصيانة خوادم أدوات متوافقة مع MCP توفر إمكانية الوصول الآمن لقواعد بيانات المؤسسات، وواجهات برمجة التطبيقات (APIs)، والمنصات السحابية، والخدمات الخارجية لوكلاء الذكاء الاصطناعي. تصميم وتطبيق البنيات الموجهة بالأحداث بالاستفادة من خدمات المراسلة والتكامل لتمكين حلول معالجة بيانات قابلة للتوسع وسريعة الاستجابة. إنشاء ونشر وكلاء تحليل البيانات، وذكاء المستندات، واسترجاع المعرفة باستخدام قوالب تطوير قابلة لإعادة الاستخدام وأدوات المنصة. التعاون مع فرق متعددة الوظائف لتقديم حلول هندسة بيانات وحلول قائمة على الذكاء الاصطناعي موثوقة وقابلة للتوسع ومبتكرة.
مؤهلات المرشح المطلوبة
- درجة البكالوريوس في علوم الحاسوب، أو هندسة البيانات، أو تقنية المعلومات، أو الذكاء الاصطناعي، أو ما يعادلها من التعليم والخبرة.
- عادةً، 3 سنوات أو أكثر من الخبرة المهنية في تطوير Python وهندسة أنابيب البيانات.
- خبرة عملية في تصميم وصيانة مسارات عمل ETL/ELT باستخدام منصات التنسيق مثل Dagster أو Airflow أو التقنيات المكافئة.
- كفاءة عالية في أطر عمل ومكتبات معالجة المستندات لاستخراج وتحويل المحتوى من مصادر البيانات المنظمة وغير المنظمة.
- خبرة في العمل مع قواعد البيانات المتجهية، ونماذج التضمين، وتقنيات البحث الدلالي، وأساليب تحسين الاسترجاع.
- الإلمام بخدمات Azure للبيانات والذكاء الاصطناعي، بما في ذلك Azure AI Search وBlob Storage وEvent Hub وDocument Intelligence، أو المنصات السحابية الأصلية المماثلة.
- خبرة مثبتة في تطوير التطبيقات ومكونات أنابيب البيانات الحاوية باستخدام Docker وممارسات النشر الحديثة.
- فهم قوي لتصميم واجهات برمجة التطبيقات REST API، وأنماط التكامل، والبنيات الموجهة للخدمات.
- خبرة في بناء أنظمة استيعاب وفهرسة واسترجاع قابلة للتوسع للتطبيقات كثيفة البيانات.
- تُعد الخبرة في بنيات التوليد المعزز بالاسترجاع (RAG) ومنصات إدارة المعرفة في المؤسسات مزية إضافية.
- يُفضل الإلمام بتصميم الأنظمة الموجهة بالأحداث وأنماط معالجة البيانات الموزعة.
- يُفضل بشدة معرفة منظومات وكلاء الذكاء الاصطناعي، وتكاملات MCP، وبنيات تحويل البيانات إلى وكلاء الحديثة.
- مهارات قوية في التحليل، وتتبع الأعطال وإصلاحها، وحل المشكلات.
- قدرات ممتازة على التواصل والتعاون مع الجهات المعنية الفنية والتجارية.
- تُعد الشهادات ذات الصلة في Azure، أو هندسة البيانات، أو الذكاء الاصطناعي، أو التقنيات السحابية مزية إضافية.
Essential Job Functions: Design, develop, and maintain scalable document ingestion and processing pipelines to support enterprise knowledge management and AI-powered applications. Build and optimize multi-stage data workflows for extracting, transforming, enriching, and indexing content from diverse document formats including PDF, PPTX, DOCX, CSV, and image-based sources. Implement advanced document processing capabilities, including metadata extraction, content chunking, classification, summarization, and information retrieval workflows. Manage vector database indexing strategies and search optimization to ensure high-quality retrieval performance and relevance. Develop and maintain MCP-compatible tool servers that securely expose enterprise databases, APIs, cloud platforms, and external services to AI agents. Design and implement event-driven architectures leveraging messaging and integration services to enable scalable and responsive data processing solutions. Create and deploy data analysis, document intelligence, and knowledge retrieval agents using reusable development templates and platform tooling. Collaborate with cross-functional teams to deliver reliable, scalable, and innovative data engineering and AI-driven solutions.
Desired Candidate Profile
- Bachelor's degree in Computer Science, Data Engineering, Information Technology, Artificial Intelligence, or equivalent combination of education and experience.
- Typically, 3+ years of professional experience in Python development and data pipeline engineering.
- Hands-on experience designing and maintaining ETL/ELT workflows using orchestration platforms such as Dagster, Airflow, or equivalent technologies.
- Strong proficiency with document processing frameworks and libraries for extracting and transforming content from structured and unstructured data sources.
- Experience working with vector databases, embedding models, semantic search technologies, and retrieval optimization techniques.
- Familiarity with Azure data and AI services, including Azure AI Search, Blob Storage, Event Hub, Document Intelligence, or comparable cloud-native platforms.
- Proven experience developing containerized applications and pipeline components using Docker and modern deployment practices.
- Strong understanding of REST API design, integration patterns, and service-oriented architectures.
- Experience building scalable ingestion, indexing, and retrieval systems for data-intensive applications.
- Experience with Retrieval-Augmented Generation (RAG) architectures and enterprise knowledge management platforms is a plus.
- Familiarity with event-driven system design and distributed data processing patterns is preferred.
- Knowledge of AI agent ecosystems, MCP integrations, and modern data-to-agent architectures is highly desirable.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication and collaboration abilities across technical and business stakeholders.
- Relevant certifications in Azure, Data Engineering, Artificial Intelligence, or Cloud Technologies are a plus.