HyeongWook Lee

HyeongWook Lee

AI Solutions Engineer
Ansan, Gyeonggi · LG Innotek
Penn State University · BS Computer Science
PortfolioBlog

About

AI engineer at LG Innotek. I build the harness around LLMs — agentic loops with tool orchestration, human-in-the-loop approvals, and multi-agent coordination — and run them in daily production use.

Most of the job is working out the real problem with non-technical users first, then staying with it through deployment until they rely on it day to day.

Experience

LG Innotek

Gen AI Engineer

Jan 2023 – Present

I build and operate one on-premise AI platform (React + FastAPI) used across the organization — four production modules over Korean-language enterprise documents. Unattended ingestion, GPU-hosted embeddings (BGE-M3, ColQwen2.5) and Qdrant vector search underneath.

Document Intelligence over Microsoft Teamsfrom full re-scans to delta sync

20+ executives and team leads rely on it weekly.

Reports auto-ingest from Microsoft Teams via the Graph API; delta tokens replaced an ~18-minute full re-scan per cycle with an incremental change feed. A dual index — page-image embeddings (ColQwen2.5) fused with BM25 sparse text — resolves chart-heavy slides and exact keyword lookups in one search. Answers stream from a 13-tool agentic loop with resumable human-in-the-loop approvals.

MS Graph APIDelta SyncHybrid RetrievalColQwen2.5Agentic Loop

Failure-Analysis Automationfrom a searchable knowledge base to generated drafts

Search had a ceiling, so I generated the documents instead.

A two-pass vision-LLM pipeline segments inconsistent, multi-case decks (every author formats them differently) into individually embedded cases for cited, agentic Q&A. Then the generator drafts standardized reports straight from log data, building the charts, commenting on the data, and flagging likely causes. It's a draft, not a verdict: the engineer concludes, and every number is computed in code, never by the model. The drafts feed back into a cleaner knowledge base.

Vision-LLMQdrantReport Generationmatplotlib

Governed Text-to-SQL Analyticsfrom spreadsheet-locked data to self-service

Non-technical users ask spreadsheet-locked data questions in plain language.

One document set is mostly spreadsheets, so instead of retrieval this module runs Text-to-SQL on DuckDB. Every generated query passes a defense-in-depth check — sqlglot AST validation, SELECT-only whitelists, forced row limits, read-only connections, and one self-correcting retry — and a governed semantic layer exposes only admin-verified tables to the model. Users can also compose personal agents from a whitelisted tool palette, no code required.

DuckDBText-to-SQLsqlglotSemantic Layer

AWS Migrationfrom on-premise to AWS

Designed it with LG CNS, and I'm executing it now.

Migrating the production platform to AWS, making the cost and architecture calls that keep it viable.

AWSArchitectureCost ModelingMigration

LG Innotek

Software Engineer Intern

Jul 2022

Developed a CNN-based defect image detection system for manufacturing quality control.

Projects

Crude Compass

Crude Compass

AI Copilot for Crude-Oil Procurement Decisions

Crude procurement still rides on a few experts' intuition. Crude Compass turns public geopolitical and market signals into AI-written daily briefs a manager just approves or rejects, before prices move.

  • Agent Bricks multi-agent supervisor routes natural-language questions across Genie, a RAG knowledge assistant, and a recommendation agent — streaming the live sub-agent trace
  • Unity Catalog Bronze → Silver → Gold pipeline fuses 6 open data sources into a daily 0–100 procurement-pressure score
  • Deployed on Databricks Apps with Lakebase (Postgres) OLTP and Lakeflow scheduled jobs
Databricks AppsAgent BricksGenieLakebaseUnity CatalogFastAPIReactDatabricks Hackathon 2026
Gom Score

Gom Score

Aria Diction Library for Vocal Performers

A vocalist handed a foreign-language score (Italian, German, Latin) spends hours on diction and meaning research before rehearsal can even begin. Built for a professional vocalist's real workflow, Gom Score does that research up front and keeps it in one searchable place.

  • Model-agnostic via OpenRouter: after comparing several LLMs, the pipeline uses Gemini to extract original lyrics, vocal diction (IPA), word meanings, literal/figurative translations, and work commentary from score content
  • Searchable library of published arias with automated content validation and moderation
  • Full-stack Next.js app backed by Supabase with server-side rendering and end-to-end test coverage
Next.jsSupabaseOpenRouterGeminiTypeScriptTailwind CSS

Awards

Customer Excellence Award · 1st Place

DX Division, LG Innotek — enterprise LLM platform

LG Bootcamp Innovation Award · 2nd Place

LG — emotion-detecting AI assistant

Tech Stack

Languages
PythonSQLJavaScript / TypeScript
Databricks
Unity CatalogDelta LakeDLTGenieAgent BricksLakebase
LLM & AI
Agentic LoopsMulti-Agent SystemsProduction RAGText-to-SQLLLM Evaluation
Frameworks & Cloud
ReactNext.jsFastAPIAWSAzure OpenAISnowflake / Snowpark
Databases
PostgreSQLMongoDBQdrantDuckDB

Certifications

Databricks Certified Data Engineer Associate

Databricks

Feb 2026

Databricks Certified Generative AI Engineer Associate

Databricks

Feb 2026

AWS Certified Solutions Architect – Associate

Amazon Web Services

Nov 2025

Education

Penn State University

BS Computer Science · Minor in Mathematics

University Park, PA, USA

2017 – 2022