English summary for screening — check the original posting before applying.
Frictio is developing an AI-Native CRM to eliminate friction in sales activities. This role focuses on designing, implementing, and operating the core LLM call infrastructure, inference routing, evaluation, and job orchestration from scratch. You will be responsible for the inference economics and quality of an AI-Native product.
Must-haves
- 3+ years of experience in distributed systems or data infrastructure design/operation
- Production system development experience with Python and practical SQL usage
- Production operation experience with containers/orchestration (Docker, Kubernetes)
- Experience with development flows assuming Infrastructure as Code (Terraform, Bicep) and CI/CD
- Infrastructure design experience on cloud platforms (Azure/AWS/GCP)
- Experience defining performance, cost, and quality numerically and verifying improvements based on measurements
- Experience designing/operating development environments using AI coding agents (Claude Code, Codex)
- Willingness to work full-stack to solve problems without being confined by job role boundaries
- Ability to formulate questions and propose designs independently in areas where specifications are not fixed
- Practical depth in at least one of the following: GPU inference/learning infrastructure, multi-backend routing/traffic control, ML/LLM application evaluation pipelines, or workflow orchestration platforms
Nice-to-haves
- Production operation experience with LLM applications (RAG, agents, tool calls)
- Practical experience with model distillation, quantization, or fine-tuning
- Experience making decisions based on inference economics (token cost, GPU throughput)
- Experience designing stream processing platforms (Kafka, Pub-Sub) and event-driven architectures
- Experience introducing/evaluating MLOps/LLMOps toolchains
- Experience designing large-scale log infrastructure (data lakes, columnar formats, query engines)
- Experience designing/operating LLM tracing/observability infrastructure using OpenTelemetry/Langfuse
- Practical experience with cost optimization (FinOps)
- Experience with infrastructure for speech recognition/NLP services
- Experience with OSS contributions, technical articles, or presentations
- Ability to catch up on research papers/technical trends and assess applicability to company architecture
- Ability to read technical documents and communicate in English
Tech stack
TypeScriptNext.jsReactNode.jsPostgreSQLLLMRAGAWSVercelvLLMTensorRT-LLMSGLangCUDADockerKubernetesTerraformBicepAzureGCPAirflowDagsterTemporalArgoOpenTelemetryLangfuseKafkaPub-Sub
Work style
Hybrid, Tokyo (Shinagawa-ku). Flexible working hours with core hours from 10:00 to 15:00.
Other notes
Annual salary: ¥9,900,000 - ¥20,000,000. Employment type: Full-time. Probationary period: 3 months. Social insurance included. Non-smoking indoors.