Recruitment background

Senior AI Data Pipeline Engineer

42dot|2026. 7. 30. 게시|
38

공고 원문

경력3년~12년
채용 유형정규직
학력무관
지역경기
마감일마감
출처원티드

소개

About the Team & Mission 42dot의 AI 데이터 파이프라인 엔지니어는 전 세계에서 수집되는 데이터를 처리하고 관리하는 글로벌 데이터 파이프라인을 설계하고 확장합니다. 페타바이트(PB)급 데이터를 대규모 GPU 인프라에 안정적으로 전달하여, 핵심적인 AI 워크로드를 가동하는 고처리량 시스템을 구축하고 운영하게 됩니다. At 42dot, our AI Data Pipeline Engineer architect and scale global data pipelines that ingest and process data from worldwide sources. You will design and operate high-throughput systems to reliably deliver petabyte-scale data to our large-scale GPU infrastructure, powering mission-critical AI workloads.

주요업무

• 다양한 AI 및 머신러닝 프로젝트를 지원하기 위한 고성능·고확장성 데이터 파이프라인 설계 및 구축 • 글로벌 데이터 가용성 및 원활한 동기화를 위한 멀티 리전(Multi-region) 데이터 인프라 아키텍처 설계 및 구현 • 여러 AI 프로젝트를 동시 지원할 수 있도록 복잡한 브랜칭 및 로직 격리가 가능한 유연한 파이프라인 아키텍처 개발 • Databricks 및 Spark를 활용한 대규모 데이터 처리 워크로드 최적화(처리량 극대화 및 비용 최소화) • Kubernetes 기반 컨테이너 데이터 환경 유지 보수 및 고도화로 데이터 워크로드의 안정적 실행 보장 • AI 리서처 및 플랫폼 팀과 협업하여 고품질 데이터를 학습 및 평가 파이프라인으로 효율적으로 공급 • Design and build high-performance, scalable data pipelines to support diverse AI and Machine Learning initiatives across the organization. • Architect and implement multi-region data infrastructure to ensure global data availability and seamless synchronization. • Develop flexible pipeline architectures that allow for complex branching and logic isolation to support multiple concurrent AI projects. • Optimize large-scale data processing workloads using Databricks and Spark to maximize throughput and minimize processing costs. • Maintain and evolve the containerized data environment on Kubernetes, ensuring robust and reliable execution of data workloads. • Collaborate with AI researchers and platform teams to streamline the flow of high-quality data into training and evaluation pipelines.

자격요건

• 대규모 AI/ML 데이터셋을 위한 프로덕션급 데이터 파이프라인 구축 및 운영 경험 • Apache Spark 및 Databricks 생태계 등 분산 처리 프레임워크에 대한 높은 숙련도 • Apache Airflow 등 워크플로우 오케스트레이션 도구를 활용한 복잡한 의존성 관리 및 실무 경험 • Kubernetes 및 컨테이너 기술을 활용한 데이터 처리 컴포넌트 배포 및 확장 능력 • Apache Kafka 등 분산 메시징 시스템을 활용한 고처리량 데이터 수집 및 이벤트 기반 아키텍처 이해 • Python을 활용한 시스템 레벨 최적화 및 수준 높은 프로그래밍 역량 • 보안과 확장성을 고려한 클라우드 네이티브 서비스 및 인프라 구축 best practices에 대한 이해 • 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력 • 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량 • Extensive professional experience in building and operating production-grade data pipelines for massive-scale AI/ML datasets. • Strong proficiency in distributed processing frameworks, particularly Apache Spark and the Databricks ecosystem. • Deep hands-on experience with workflow orchestration tools like Apache Airflow for managing complex dependency graphs. • Solid understanding of Kubernetes and containerization for deploying and scaling data processing components. • Proficiency in distributed messaging systems such as Apache Kafka for high-throughput data ingestion and event-driven architectures. • Expert-level programming skills in Python for system-level optimizations. • Strong knowledge of cloud-native services and best practices for building secure and scalable data infrastructure. • Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems. • Strong communication skills to effectively collaborate with cross-functional teams and external partners.

우대사항

• 글로벌 멀티 리전 파이프라인 설계 및 국가 간 데이터 전송/지연 시간(Latency) 이슈 해결 경험 • Ray 등 AI 워크로드를 위한 분산 컴퓨팅 프레임워크 구현 경험 또는 깊은 관심 • Spark Streaming 또는 Flink를 이용한 실시간/준실시간(Near real-time) 파이프라인 구축 경험 • Terraform 등 Infrastructure as Code(IaC) 도구를 활용한 복잡한 데이터 환경 관리 경험 • 전체 ML 생애주기(MLOps) 및 데이터 인프라가 모델 실험과 배포를 지원하는 메커니즘에 대한 이해 • Experience in architecting global, multi-region data pipelines and solving challenges related to cross-border data transfer and latency. • Practical experience or a strong interest in implementing distributed computing frameworks like Ray for AI workloads. • Experience in building real-time or near-real-time pipelines using Spark Streaming or Flink. • Familiarity with Infrastructure as Code (IaC) tools such as Terraform to manage complex data environments. • Understanding of the end-to-end ML lifecycle (MLOps) and how data infrastructure supports model experimentation and deployment.

혜택 및 복지

[42dot만의 업무 몰입 프로그램] https://42dot.ai/ko/careers/employee-engagement-program

이런 공고는 어때요?

워트인텔리전스

워트인텔리전스

7년~20년 · 정규직 · 서울

Data Engineer: Ontology 전문가 (7년 이상)#그래프설계  #지식그래프
상시
그룹바이그룹바이
13
울트라 텐던시 코리아

울트라 텐던시 코리아

20년 이하 · 정규직 · 해외

[100%재택]데이터브릭스 레지던트 솔루션 아키텍트#데이터브릭스  #아키텍처설계
상시
그룹바이그룹바이
46
어쎈드(ASCEND)

어쎈드(ASCEND)

10년 이하 · 정규직 · 서울

Quantitative Research Engineer#퀀트트레이딩  #트레이딩개발  #재택가능
상시
그룹바이그룹바이
792
페이스웹

페이스웹

20년 이하 · 정규직 · 서울

SLM Modeling & Data Engineering#데이터자산화  #파인튜닝  #의료AI
상시
그룹바이그룹바이
76
미리비트

미리비트

10년 이하 · 정규직 · 경기

빅데이터 엔지니어,서버 백엔드 엔지니어 (인메모리 기반)#파이프라인  #실시간데이터  #빅데이터
상시
그룹바이그룹바이
653
워트인텔리전스

워트인텔리전스

10년~15년 · 정규직 · 서울

Senior Data Engineer (10년 이상) | 시니어 데이터 플랫폼 엔지니어#시용기간  #아키텍처설계  #데이터플랫폼
상시
그룹바이그룹바이
4
앰플랩

앰플랩

3년~20년 · 정규직 · 대전

Field AI Solutions Engineer#데이터설계  #AI자동화  #스마트팩토리
상시
그룹바이그룹바이
7
스텝에이아이

스텝에이아이

20년 이하 · 정규직 · 서울

[팀 스케일업] AI 풀스택 개발자 (LLM 파이프라인 · 대규모 데이터 수집)#B2BAI  #대용량데이터  #LLM파이프라인
상시
그룹바이그룹바이
152
미리디

미리디

3년~6년 · 정규직 · 서울

[미리디] Data Engineer#레이크하우스  #데이터엔지니어  #디자인플랫폼
상시
그룹바이그룹바이
41
콕스웨이브

콕스웨이브

3년~7년 · 정규직 · 서울

[AX AgentX] 데이터 엔지니어(RAG/LLM Pipeline)#LLMOps  #ETL파이프라인  #AI솔루션
상시
그룹바이그룹바이
42
미리비트

미리비트

3년~10년 · 정규직 · 경기

데이터 플랫폼 엔지니어 (3년차~PL) Spark · Airflow · K8s#스트리밍  #빅데이터
상시
그룹바이그룹바이
62
그로잉랩

그로잉랩

3년~9년 · 정규직 · 서울

개인투자자의 블룸버그, '버틀러'의 데이터 엔진을 구축할 시니어 엔지니어#연봉1억이상  #핀테크  #ETL파이프라인
상시
그룹바이그룹바이
201
페이스웹

페이스웹

20년 이하 · 정규직 · 서울

AI 학습데이터 · SLM 운영 담당#파인튜닝  #데이터정제  #의료AI
상시
그룹바이그룹바이
43
피에프씨테크놀로지스

피에프씨테크놀로지스

7년 이상 · 정규직 · 서울

B2B Senior Data Engineer(시니어 데이터 엔지니어)-Databricks#핀테크  #데이터플랫폼  #시차출퇴근
상시
그룹바이그룹바이
1
미리비트

미리비트

3년~10년 · 정규직 · 경기

데이터 플랫폼 엔지니어#숙소지원  #마이그레이션  #데이터플랫폼
상시
그룹바이그룹바이
71
위시켓

위시켓

3년~10년 · 정규직 · 서울

AIDP FDE(Forward Deployed Engineer) (3년 ~ 12년)#BI구축  #데이터솔루션  #원격근무
상시
그룹바이그룹바이
44
피에프씨테크놀로지스

피에프씨테크놀로지스

3년 이상 · 정규직 · 서울

B2B Data Engineer(B2B 데이터 엔지니어) - Databricks#레이크하우스  #핀테크  #시차출퇴근
상시
그룹바이그룹바이
1
레브잇

레브잇

20년 이하 · 정규직 · 서울

[쇼포트] Product Engineer (Web Scraping)#이커머스  #대용량트래픽  #웹스크래핑
상시
그룹바이그룹바이
162
페이스웹

페이스웹

9년 이하 · 정규직 · 서울

[FACEWEB] 의료 SLM/LLM 엔지니어: Fine-tuning·STT·RAG•경량화#파인튜닝  #인공지능
상시
그룹바이그룹바이
98
포트로직스

포트로직스

5년 이상 · 정규직 · 경기

[FDE] Forward Deployed Engineer - 플랫폼셀#AI에이전트  #포워딩  #데이터통합
상시
그룹바이그룹바이
3