المسؤوليات
- إنشاء بحيرة بيانات مركزية على خدمات بيانات GCP من خلال دمج مصادر بيانات متنوعة عبر المؤسسة.
- تطوير وصيانة وتحسين أنابيب معالجة البيانات الدُفُعية والتدفقية المعتمدة على Spark.
- استخدام خدمات بيانات GCP لمهام هندسة البيانات المعقدة وضمان التكامل السلس مع عناصر المنصة الأخرى.
- تصميم وتنفيذ تحقق من صحة البيانات وضمان الجودة لضمان الدقة والكمال والاتساق عبر سير عمل البيانات لدينا.
- التعاون الوثيق مع فريق الذكاء الاصطناعي / تعلم الآلة لتمكين حالات استخدام الذكاء الاصطناعي وبشكل خاص تلبية احتياجات البيانات لتطوير نماذج التعلم الآلي ومراقبتها.
- التعاون مع فرق متعددة التخصصات، بما في ذلك محللي البيانات ومستخدمي الأعمال من أقسام العمليات والتسويق والتجاري، لاستخراج رؤى قيمة وتمكين اتخاذ قرارات معتمدة على البيانات.
- التعاون مع فرق المنتج لتصميم وتنفيذ وصيانة نماذج البيانات للاستخدامات التحليلية.
- المشاركة في استكشاف التكنولوجيا والبحث، وتطوير PoCs، وإجراء تحقيقيات معمقة.
- تصميم وإدارة عمليات ETL/ELT، وضمان تكامل البيانات وتوفرها وأدائها.
- استكشاف مشكلات البيانات وإجراء تحليل الجذر عند وجود تقارير حول البيانات.
- العمل على مبادرات متعلقة بالذكاء الاصطناعي التوليدي المرتبطة بأهداف الشركة.
المهارات التقنية المطلوبة
- PySpark - الدُفعي والتدفق
- GCP - Dataproc, Dataflow, DataStream, Dataplex, Pub/Sub, BigQuery وCloud Storage
- NoSQL (يفضل MongoDB)
- لغات البرمجة: Scala/Python
- Great Expectations، أو إطار DQ مشابه
- الإلمام بأدوات إدارة سير العمل مثل: Airflow, Prefect أو Luigi
- فهم حوكمة البيانات، تخزين البيانات ونمذجة البيانات
- معرفة جيدة بـ SQL
الأعمال
- القدرة على التواصل بفعالية، واستخلاص المعرفة التقنية في رسائل سهلة الهضم بشكل موجز/مرئي
- التحديد الاستباقي للمبادرات التطويرية مع الفريق، ودعم الأعضاء junior
المهارات المرغوبة
- البنية التحتية كرمز، ويفضل Terraform
- Docker وKubernetes
- Looker
- معرفة في هندسة AI/ML
- السَلَك، أو أدوات ذات صلة مثل Atlan
- DBT
الملف المرشح المطلوب
في ياسير، نؤمن بقوة التنوع وأهمية الثقافة الشمولية. إذا كنت مستعداً لتقديم منظورك وخبراتك الفريدة، فنحن متحمسون لسماعك. لا تتقدم لعمل فحسب، بل انضم إلى رحلتنا. لنصنع غدًا أفضل معًا. نتطلع لتلقي طلبك! بالتوفيق، فريق TA في ياسير
قد نستخدم أدوات الذكاء الاصطناعي (AI) لدعم أجزاء من عملية التوظيف، مثل مراجعة الطلبات، تحليل السير الذاتية، أو تقييم الردود وتحديد الإشارات أو التناقضات المحتملة في مواد التقديم بناءً على المعلومات المتاحة. هذه الأدوات تدعم فريق التوظيف لدينا ولكنها لا تستبدل الحكم البشري. قرارات التوظيف النهائية تتخذها الأشخاص في النهاية. إذا كنت ترغب في مزيد من المعلومات حول كيفية معالجة بياناتك، يرجى الاتصال بنا."
Responsibilities
- Build a centralized data lake on GCP data services by integrating diverse data sources throughout the enterprise.
- Develop, maintain, and optimize Spark-powered batch and streaming data processing pipelines.
- Leverage GCP data services for complex data engineering tasks and ensure smooth integration with other platform components.
- Design and implement data validation and quality checks to ensure accuracy, completeness and consistency across our data workflows.
- Work closely with the AI / ML team to enable AI use cases and in particular address data needs for Machine Learning model development and monitoring.
- Collaborate with cross-functional teams, including Data Analysts and other business users from Operations, Marketing and Commercial departments, to extract valuable insights and enable data-driven decision making.
- Collaborate with the Product teams to design, implement, and maintain the data models for analytical use cases.
- Engage in technology exploration and research, the development of PoC s, conducting deep investigations.
- Design and manage ETL/ELT processes, ensuring data integrity, availability, and performance.
- Troubleshoot data issues and conduct root cause analysis when reporting data is in question.
- Work on Gen AI related initiatives aligned with the company goals.
Required Technical Skills
- PySpark - Batch and Streaming
- GCP - Dataproc, Dataflow, DataStream, Dataplex, Pub/Sub, BigQuery and Cloud Storage
- NoSQL (preferably MongoDB)
- Programming languages: Scala/Python
- Great Expectation, or similar DQ framework
- Familiarity with workflow management tools like: Airflow, Prefect or Luigi
- Understanding of Data Governance, Data Warehousing and Data Modelling
- Good SQL knowledge
Business
- Able to communicate effectively, distill technical knowledge into digestible messages in a succinct / visual way
- Proactively identify and contribute with team development initiatives, and supporting junior members
Good to have skills
- Infrastructure-as-Code, preferably Terraform
- Docker and Kubernetes
- Looker
- AI / ML engineering knowledge
- Lineage, or relevant tools e.g. Atlan
- DBT
Desired Candidate Profile
At Yassir, we believe in the power of diversity and the importance of an inclusive culture. So, if you're ready to bring your unique perspective and experiences to the table, then we're excited to listen. Don't just apply for a job, come and be a part of our journey. Let's create a better tomorrow together. We look forward to receiving your application! Best of luck, Your Yassir TA Team
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.