Building the core engine of TheROAD — a Korean long-form AI video generation engine for commercial film and drama. Twenty-five years spent where AI, video and large-scale infrastructure meet: 20M-user streaming platforms, 8,000-server CDNs, and AI systems running on real factory floors. TheROAD의 코어 엔진 개발 중 — 상업 영화·드라마를 위한 한국형 롱폼 AI 영상 생성 엔진. 25년간 AI·영상·대규모 인프라가 만나는 지점에서 일해왔음. 2천만 가입자 스트리밍 플랫폼, 8천 대 규모 CDN, 실제 공장에서 돌아가는 AI 시스템.
Today's generative video models produce beautiful eight-second clips and then fall apart. Faces drift between cuts, the same room regenerates as a different place, and the director loses control of the shot. That makes them useful for previz and useless for a finished film.
TheROAD attacks that gap. I design and build the core engine and Canvas node pipeline — a 100,000+ line proprietary engine whose central problem is cut-to-cut consistency: holding a character's identity, a location's geometry and a scene's lighting stable across a full-length narrative.
The architecture is co-directed and self-improving. A director's corrections are captured as structured judgment data; the engine generates multiple model / prompt / node-graph variants, and a VLM scoring layer picks the winner while rejected branches are kept as evidence for the next iteration. It learns the craft rather than guessing at it.
The training data is the part no big tech lab can copy: rights-clean original artwork from 100+ Korean commercial productions — Train to Busan, JUNG_E, Moving — owned by the studio itself.
지금의 생성형 영상 모델은 8초짜리 클립은 훌륭하게 만들지만 그 다음부터 무너집니다. 컷이 바뀌면 인물의 얼굴이 달라지고, 같은 공간이 다른 장소로 재생성되며, 감독은 연출 통제권을 잃습니다. 프리비즈에는 쓸 수 있어도 본편에는 쓸 수 없는 이유.
TheROAD는 그 지점을 공략합니다. 코어 엔진과 Canvas 노드 파이프라인을 직접 설계·구축 중 — 10만 라인 이상의 독자 엔진이며, 핵심 과제는 컷 간 일관성 제어입니다. 장편 서사 전체에 걸쳐 인물의 동일성, 공간의 구조, 씬의 조명 톤을 유지하는 아키텍처.
구조는 협업 연출형(Co-Directed) 자가 개선 엔진입니다. 감독의 수정 지시가 정형화된 판단 데이터로 축적되고, 엔진이 모델·프롬프트·노드 그래프를 여러 갈래로 변주해 생성하면 VLM 시각 판단 계층이 최적안을 채택합니다. 탈락한 후보도 다음 개선의 근거로 남습니다. 감(感)에 의존하지 않고 제작 노하우 자체를 학습하는 구조.
학습 데이터는 빅테크가 복제할 수 없는 자산입니다 — 〈부산행〉, 〈정이〉, 〈무빙〉 등 100여 편의 한국 상업 작품에서 축적된, 스튜디오가 직접 보유한 저작권 청정 오리지널 아트웍.
Four domains, each carried from prototype to production and real users. 네 개 분야 모두 프로토타입에서 실제 서비스·실제 사용자까지 끌고 간 경험.
Long-form generative video engines, agentic pipelines with human-in-the-loop feedback, computer-vision tracking, ML anomaly detection on sensor streams, and semantic vector search. Systems that decide, not demos. 롱폼 생성형 영상 엔진, 전문가 피드백이 순환하는 에이전틱 파이프라인, 컴퓨터 비전 기반 추적, 센서 스트림 위의 ML 이상 탐지, 의미 기반 벡터 검색. 데모가 아니라 실제로 판단하는 시스템.
Transcoding pipelines, players, CMS architecture and OTT delivery — from Korea's largest UGC video site to a US streaming service, plus embedded playback on IPTV set-top boxes. 트랜스코딩 파이프라인, 플레이어, 대용량 CMS 설계, OTT 전송. 국내 최대 UGC 동영상 사이트부터 미국 OTT 서비스, IPTV 셋톱박스 임베디드 재생기까지.
Next-generation CDN architecture, hybrid P2P delivery, WAF, TCP stack tuning and multi-CDN cost engineering across global points of presence. 차세대 CDN 아키텍처, P2P 하이브리드 전송, WAF, TCP 스택 최적화, 글로벌 거점 기반 멀티 CDN 비용 설계.
Sensor-to-cloud pipelines, BLE and ultrasonic device integration, firmware-level hardware control, and real-time analytics on industrial data. 센서-투-클라우드 파이프라인, BLE·초음파 디바이스 연동, 펌웨어 수준의 하드웨어 제어, 산업 데이터 실시간 분석.
Architected the video CMS and FFmpeg transcoding system behind Korea's #1 video platform — 1.5M DAU, 20M+ users, 3,000 uploads a day.국내 1위 동영상 플랫폼의 영상 CMS와 FFmpeg 트랜스코딩 시스템 설계. DAU 150만, 가입자 2천만, 하루 3천 건 업로드 처리.
Built the world's first video player with playback speed control — now a default feature on every major platform.세계 최초로 배속 재생 기능을 구현한 동영상 플레이어 개발. 현재 모든 주요 플랫폼의 기본 기능.
Technical lead in raising over $16M from Silicon Valley VCs for global expansion.글로벌 확장을 위한 실리콘밸리 VC 1,600만 달러 이상 투자 유치의 기술 리드.
Led R&D for a top-4 global CDN — 8,000 servers, over 1 Tbps, and a next-gen edge architecture designed for 15K servers across 110 PoPs.세계 4위 CDN의 연구개발 총괄. 서버 8천 대, 1Tbps 이상 트래픽, 서버 1만5천 대·110개 거점을 상정한 차세대 엣지 아키텍처 설계.
BLE hardware, mobile app and on-device AI tracking shipped as one product. CES 2022 Innovation Award, sold worldwide.BLE 하드웨어·모바일 앱·온디바이스 AI 추적을 하나의 제품으로 출시. CES 2022 혁신상 수상, 전 세계 판매.
3i closed a $24M (₩28B) Series A while I led software engineering for both of its products — Beamo.ai, an enterprise digital twin adopted by NTT and other global enterprises, and Pivo.ai, a CES award-winning AI tracking device.소프트웨어 개발을 총괄하던 시기에 쓰리아이가 시리즈A 2,400만 달러(280억 원)를 유치. 두 제품 모두 담당 — NTT 등 글로벌 기업이 도입한 엔터프라이즈 디지털 트윈 Beamo.ai, CES 혁신상을 받은 AI 추적 디바이스 Pivo.ai.
Deep-learning O-ring inspection across 50+ part types at 97.3% accuracy, 4 parts per second — built in 3 months, still running on the factory floor.50종 이상 고무링을 97.3% 정확도, 초당 4개 속도로 판별하는 딥러닝 검사 장비. 3개월 만에 완성, 현재도 공장에서 가동 중.
Founded iToVi and built a video search engine with a patented UI/UX — US20100036878.아이토비 창업, 동영상 검색 엔진 개발 및 UI/UX 미국 특허 출원 — US20100036878.
A local-first semantic memory search for AI agents — automatic document indexing, meaning-based retrieval, no server required. 286 tests plus benchmarks.AI 에이전트를 위한 로컬 우선 의미 검색 시스템 — 문서 자동 색인, 의미 기반 검색, 서버 불필요. 286개 테스트와 벤치마크 포함.
A personal project integrating LLM-based scenario parsing, entity extraction and automated illustration insertion via text-to-image models — the prototype that became TheROAD.LLM 기반 시나리오 파싱, 개체 추출, 텍스트-투-이미지 모델을 이용한 삽화 자동 삽입을 결합한 개인 프로젝트. TheROAD의 원형이 된 프로토타입.
LLM agents, diffusion, vector search, computer vision, TensorFlow
FFmpeg, transcoding, HLS / RTMP, OTT clients, embedded playback
Apache Traffic Server, Nginx, OpenResty, P2P hybrid delivery, TCP tuning
Python, Golang, Node.js, C / C++, Java
AWS, Azure, Kubernetes, Terraform, Kafka, NiFi, PostgreSQL, MongoDB
BLE, embedded Linux, ultrasonic sensing, firmware-level control
Open to conversations about AI in production, video systems at scale, and infrastructure that has to actually work. 프로덕션에 들어가는 AI, 대규모 영상 시스템, 실제로 동작해야 하는 인프라에 대한 이야기는 언제든 환영합니다.