Apr – Jul 2026 · Beijing

Tencent · Hunyuan

AI Product Engineer

Guided by a product-for-the-model approach, I built Tencent Hy Translate from 0 to 1 on top of hy-mt2.0 and helped launch it across Mini Program, iOS, and Android. The product supports online and on-device translation for text, voice, camera, and simultaneous interpretation. Alongside product design, I owned the model-quality track: building the evaluation set and methodology from scratch, then creating a case collection → product evaluation → algorithm iteration loop that fed real-world data back into model development. Camera translation was my deepest focus. Through competitive benchmarking and user research in China and abroad, I identified offline access, zero cost, and scenario-aware foundation-model translation as the differentiators, and drove the feature to P0 delivery.

0→1Multi-platform translation product launch26%→60%Reported bad cases resolvedBuilt model evaluation framework from scratch

My role

Independently owned the 0-to-1 product design of Hy Translate, its model evaluation framework, and bad-case-driven model improvement, while driving engineering delivery.

Product design (0→1)

  • Researched translation and AI products systematically, then used vibe coding to prototype text, voice, camera, simultaneous interpretation, and browser-extension experiences.
  • Benchmarked camera translation products and conducted user research on Xiaohongshu and Reddit, identifying offline, free, scenario-aware AI as the differentiated position and driving it to P0.
  • Drove launches across Mini Program, iOS, and Android, covering both online and offline use, so the model could reach real users.

Model evaluation framework (0→1)

  • Built a layered translation evaluation framework across language pairs, input length, usage scenarios, translation preferences, and instruction following, balancing broad coverage with priority use cases.
  • Used DCG to score old and new models case by case; graded errors as P0 / P1 / P2 and separated optimization need from root-cause attribution. Preference failures were further classified by style, terminology, and formatting to keep multi-rater standards consistent.

Model improvement (bad-case driven)

  • Mapped root causes to targeted solutions: a custom terminology library and terminology intervention for frequent terms and internet slang, source-language injection for dialects, and ground-truth examples for algorithm teams to use in targeted distillation on complex issues.
  • Increased the resolution rate of reported bad cases from roughly 26% to 60% across iterations.

What I learned

This experience taught me where model capability ends and product design begins. The best AI products translate model capability into value users can actually feel.