Description de posteAs a Senior AI/ML Engineer at vector8, you will design, implement, and deploy AI solutions that bridge the gap between research and production. Your work will focus on integrating and fine‑tuning AI models, optimizing model performance, and ensuring enterprise‑grade reliability, security, and scalability.
Hands‑on Engineering Role
Develop and optimize LLM and VLM‑powered solutions for enterprise use cases
Develop and optimize TTS, STT, and ML models
Apply software engineering best practices (testing, CI/CD, modular design, documentation)
Collaborate with cross‑functional teams (data engineers, MLOps, cloud architects, and business stakeholders)
Solve real‑world enterprise challenges (security, compliance, legacy system integration)
Own the full lifecycle of AI models, from data exploration to production monitoring
The role is primarily based in Paris, with occasional travel to client sites and collaboration with teams across Europe.
Job Requirements
5+ years of experience in AI/ML engineering, software development, or a related field
Expertise in LLM architectures and training methodologies:
Transformers, attention mechanisms, fine‑tuning, RAG, quantization
Prompt engineering, model evaluation, bias detection
Strong knowledge of machine learning architectures: fully connected, CNN, LSTM, transformers, and classical ML models
Strong software engineering skills:
Proficient in Python (FastAPI, Pydantic, asyncio, type hints)
Experience with API development
Familiarity with modern toolchains (Docker, Kubernetes, Terraform)
Hands‑on experience with LLM integrations:
LLM providers
Vector databases (Pinecone, Weaviate, Milvus)
Model serving (vLLM, TGI, KServe)
Experience with MLOps and production deployments
Understanding of enterprise challenges:
Security, compliance, scalability, cost optimization
Experience with relational and non‑relational databases
Strong problem‑solving and debugging skills
Excellent communication and collaboration skills (fluent in English; German is a strong plus)
Bachelor’s or Master’s degree in Computer Science, Mathematics, Physics, or a related field
Experience with multi‑cloud environments (AWS, Azure, GCP)
Experience with code optimization (e.g., model quantization, parallelization)
Job Responsibilities
End‑to‑end model development
Design, implement, and deploy distributed, high-volume, high‑performance, low‑latency machine‑learning solutions, focusing on GenAI models, especially LLM integrations and API‑driven architectures
Take ownership of models throughout their lifecycle: data exploration and cleaning, reproducible, versioned datasets, state‑of‑the‑art research to identify best architectures, implementation, training, optimization, deployment, monitoring, maintenance, and optimization for performance, latency, and cost efficiency in LLM serving and inference
Write clean, modular, well‑documented Python code (FastAPI, Pydantic, asyncio)
Apply best practices: testing (unit, integration, end‑to‑end), CI/CD (GitHub Actions, GitLab CI, ArgoCD), observability (logging, monitoring, tracing)
Ensure security and compliance: data protection, access controls, encryption
Integrate models and code into CI/CD pipelines for seamless deployment
Design and implement AI‑powered solutions that integrate with APIs, microservices, and event‑driven architectures
Develop and optimize AI pipelines for dataset cleaning, preprocessing, and model training; fine‑tuning; retrieval‑augmented generation; prompt engineering
Model evaluation (benchmarking, bias detection, drift analysis)
Build scalable, secure, cost‑efficient serving infrastructure (FastAPI, vLLM)
Debug and optimize performance (latency, throughput, token efficiency for transformer‑based architectures)
Deploy and monitor AI models in production
Design and implement MLOps pipelines: training, fine‑tuning, evaluation, versioning and lineage tracking, A/B testing, canary deployments
Ensure scalability and reliability (auto‑scaling, fault tolerance, disaster recovery)
Collaborate with data engineers to build data pipelines (batch, streaming, real‑time)
Work closely with product owners, DevOps, QA in an agile cross‑functional team
Mentor junior engineers and promote best practices in AI/ML and software engineering
Translate product requirements into technical solutions and architectural decisions
Document architectures, decisions, and best practices for internal and client‑facing use
Develop relationships with internal and external stakeholders, including clients and partners
Stay ahead of the latest AI and ML architectures (transformers, Mixture of Experts, sparse attention)
Experiment with cutting‑edge techniques (quantization, distillation, speculative decoding)
Evaluate and benchmark open‑source and proprietary models (Llama, Mistral, Mixtral, GPT‑4, Claude)
Bring your own ideas through vector8’s ideation process
Contribute to vector8’s AI accelerators (reusable components for common industry problems)
Embrace a strategic and continuous improvement mentality to drive innovation
Job Benefits
A competitive compensation package with benefits
Flexible working hours, including remote work options (hybrid model)
25 days of paid vacation per year, plus additional flex days
Private health, life insurance and a pension plan for long‑term security
Home office allowance and lunch vouchers
Discounted fitness memberships
50% reimbursement of public transport costs
Free coffee, fruit, and snacks to keep you fueled
Access to the latest technologies (LangDock, Claude Code for developers)
Grants for training, coaching, and conferences
Opportunities to attend industry events and represent vector8 as a thought leader
A less‑formal work environment where authenticity and collaboration thrive
A diverse and inclusive team that values curiosity, ownership, and innovation
Why This Role is Unique
You will work at the intersection of AI research and enterprise software engineering with a strong focus on AI‑driven solutions
You will contribute to shaping the future of AI adoption in France’s most complex organizations
You will bridge the gap between cutting‑edge AI and real‑world enterprise constraints (security, compliance, legacy systems)
You will grow your skills, collaborate with talented engineers, and have real ownership over your work from day one
#J-18808-Ljbffr