Victor Kwong
Data Scientist & Software Engineer
I build agentic AI, concurrent C++ systems and financial data tools. From enterprise AI at Capital One to audio-editing research with Dolby Laboratories, my work connects evaluation, reliable software and practical applications.
Hong Kong Permanent Resident · Available January 2027
Professional Experience
Capital One
Built a GPT-OSS-120B text-to-SQL agent achieving 90%+ query accuracy on PostgreSQL/Snowflake compliance metrics across Cyber, Technology Risk and Resilience; used dynamic schema selection and few-shot examples to reduce prompt context.
Integrated SQL, retrieval and direct-response workers into a LangGraph supervisor workflow for investigating non-compliant metrics in an enterprise risk platform.
Engineered the agent execution harness with SQL validation, execution-based verification and automated retries, handling generated-query failures within a LangGraph supervisor-worker workflow.
Created metric-level evaluation datasets and a Weave benchmark suite covering RAG source selection, SQL AST structure, executable validity and LLM-as-judge semantic equivalence.
AS Watson Group
Analyzed 2M+ clickstream events with PySpark on Databricks, reconstructing user sessions and joining behavioral features; findings supported a 10% improvement in usability and a 4% reduction in bounce rate.
Enhanced search classification with Naive Bayes, word embeddings and statistical scoring, achieving 91% classification accuracy and increasing search-related purchase conversion by 13%.
Developed a new categorization method for the e-commerce recommendation engine, increasing user engagement by 15% and sales by 10% across key product categories.
Built dashboards for session and search KPIs, helping product and operations teams investigate user journeys and recommendation performance.
Futu Holdings
Built scheduled Hive SQL pipelines and ClickHouse fact and snapshot tables for fund-flow and promotional-coupon analytics across Hong Kong and Singapore markets.
Designed transaction-level fact tables and daily, weekly and monthly snapshots, using partitioning and incremental updates to support recurring analytics.
Maintained promotional-coupon conversion dashboards tracking redemption, campaign performance and product adoption; partnered with operations on data accuracy and insights, contributing to a 14% increase in Cash Plus penetration.
Automated large-transaction monitoring with Bash scripts and Slack alerts, replacing manual checks with repeatable notifications.
Shanghai Commercial Bank
Validated 50+ batch reports by reconciling IBM DB2/AS400 data and legacy/new outputs, supporting 100% on-time delivery and identifying 15+ potential production defects; executed 300+ Postman tests for transfer services.
Authored Bash scripts for database preparation and output checks, and integrated batch validation into a GitLab CI/CD workflow.
Coordinated Jira test-status dashboards, delegated tasks and tracked intern progress, improving task visibility and reducing project delays by 20%.
Code Free Soft
Built Playwright UI tests for forms, drag-and-drop workflows and navigation; contributed 15+ reusable automation functions and classes for common test scenarios.
Automated checks for multilingual UI rendering to verify consistency across supported languages.
Research & Engineering
Dolby Laboratories: Agentic Audio Editing
Research in Progress
Researching natural-language audio editing through task planning and tool scheduling, with baseline comparisons on short-audio benchmark subsets.
Developed an agentic audio-editing framework with improved task planning and tool scheduling, orchestrating editing operations from natural-language instructions.
Benchmarked against SmartDJ and WavCraft on under-10-second subsets of MMAE and SpeechEditBench, achieving approximately 7% improvement in evaluated benchmark performance.
Agentic Audio Editing
Instruction
Plan & schedule
Audio tools
Evaluate
MMAE + SpeechEditBench · under-10-second subsets
Dopamine: Cross-platform Screen-time Tracker
Public Release v0.0.3
A local-first screen-time tracker for macOS and Windows. Native agents record the foreground app and window, and a bundled dashboard turns the day into a timeline, per-app breakdown and categories.
Built native tracking agents in Swift (NSWorkspace + Accessibility) and C# (.NET 8 Native AOT, GetForegroundWindow + GetLastInputInfo) that write window activity to local SQLite and pause on lock, sleep and idle.
Served a Next.js static-export dashboard from each agent on a localhost API protected by a pairing code, with a timeline, sessions, categories and a monthly calendar that holds 60 fps without charting libraries.
Designed an evidence-ordered categorisation engine (user override, browser site, app name, community votes, OS metadata, window title) validated on 115 real process names and a held-out set.
Shipped per-window forgetting, timed tracking pauses and daily update notices, and designed opt-in community categories on Supabase that require 5 installs with 70% agreement.
Automated CI builds and tests for the Swift, C# and TypeScript components, attaching macOS and Windows bundles to tagged releases.

Public Release v0.0.3
Download for macOS & Windows
Smart Badminton: Rally Detection & Highlights
Public Release v0.1.0
Turns fixed-camera badminton recordings into rally highlights: detect candidate rallies, refine their boundaries in an editor, and export the moments worth rewatching, with footage kept on the user's device.
Built a rally-detection pipeline that fuses audio, motion and pose features into a gradient-boosting rally-state model, then a trajectory-aware segmenter that owns the final cut boundaries.
Combined YOLO11 shuttle detection, TrackNetV3 sequence tracking and InpaintNet gap repair into hybrid shuttle tracking that supplies boundary and score evidence.
Shipped two editions: Browser Quick analyses and exports MP4 entirely in the browser with FFmpeg WebAssembly, while Native Pro runs the full Python pipeline locally with source-quality rendering.
Built a bilingual Studio editor with frame nudging, clip cutting, undo/redo and atomic timeline saves that never overwrite human-reviewed ground truth.
Public Release v0.1.0
Open the Web App
BusTub: Concurrent Database Engine
Status: Private Coursework
Database systems coursework spanning concurrent indexes, buffer management, background disk I/O and transaction isolation.
Implemented a concurrent C++ B+ tree with optimistic read-latch traversal and leaf-level write locking, falling back to ancestor write latches for node splits and structural changes.
Built buffer-pool management with atomic pin counts, shared/exclusive page latches and move-only RAII guards; implemented ARC using recency/frequency lists and adaptive ghost-history tracking.
Implemented queued disk I/O on a background worker with promise/future completion, coordinating page reads, dirty-page writeback and frame replacement.
Implemented snapshot reads through MVCC undo-log reconstruction, transaction commit/abort handling, scan-predicate validation for serializable transactions and watermark-based garbage collection.
Status: Private Coursework
View Upstream Project
CMU Computer Systems: 18-613 / CS:APP
Status: Completed
Systems coursework covering multithreaded networking, explicit free-list allocation, Unix process control and machine-level security.
Engineered a concurrent, multithreaded caching HTTP proxy in C, combining network request forwarding with shared response caching for simultaneous client requests.
Developed a custom dynamic memory allocator using explicit free lists and boundary-tag coalescing, implementing malloc, free and realloc to manage heap blocks and reuse freed memory.
Built a Unix shell with job control, supporting foreground/background execution and process management; implemented a cache simulator to study memory-access behavior.
Completed Attack Lab exercises in x86-64 code injection and return-oriented programming (ROP), analyzing stack layout and control flow with GDB.

Status: Completed
View Course Syllabus
Hermes Engine: Quantitative Trading System
Private Research / Paper Trading
A quantitative research and paper-trading system connecting market data, factor screening, validation and broker-adapter workflows.
Built multi-source financial-data ingestion with PostgreSQL, DuckDB and Parquet storage, including validation before downstream research.
Implemented six factor families spanning risk, momentum, cointegration, crowding, entropy and formula alphas, with VectorBT parameter sweeps for candidate screening.
Implemented broker adapters for order management and position reconciliation in a paper-trading environment.
Designed a pre-deployment validation layer to separate strategy research from execution and capital-allocation decisions.
Private Research / Paper Trading
Internal Research Tool
FinanceCLI: Financial Research Toolkit
Public Research Toolkit
Composable Python tools for financial research and AI agents, with SEC filing evidence, structured outputs and reproducible calculations.
Built a Python CLI exposing composable financial-research tools for SEC filings, XBRL statements and calculations, with structured JSON and readable Markdown outputs.
Preserved source identifiers, calculation inputs and methods in tool outputs so users and AI agents can inspect the evidence behind financial analysis.
Packaged research commands with installation guides and optional OCR, table-extraction and VectorBT backtesting dependencies for different workflows.
$ finance filings.read AAPL --form 10-K
sec.edgar / xbrl / statement sections
$ finance market.ohlcv MSFT --interval 1d
json output ready for agent workflows
Public Research Toolkit
View Documentation
Finspresso: Financial Insights Assistant
Financial AI Application
A scalable AI-powered platform integrating LLM with RAG for real-time financial data synthesis and contextual intelligence.
Built a LangGraph plan-retrieve-generate workflow with semantic output checks and a bounded repair step for financial question answering.
Developed a React/TypeScript interface and FastAPI backend, integrating Supabase/PostgreSQL with Chroma for relational and vector retrieval.
Validated p95 first-token latency of 2 seconds or less and approximately 20-second agentic workflows; combined relational and vector retrieval with query latency below 200 ms.
Implemented Stripe webhooks with idempotency and retry handling, supporting 99.9% reliable real-time subscription updates and integrated payment workflows.

Financial AI Application
Visit Finspresso
FlinkNewsArena: Streaming Data Pipeline
Project in Development
News ingestion and enrichment with stateful stream processing, content deduplication and recoverable data delivery.
Built a Kafka/Flink pipeline for news ingestion, deduplication and storage, with checkpoint recovery and idempotent sinks to support failure handling.
Applied SimHash to identify similar articles across sources and automated content enrichment with SQL and JavaScript workflows.
Streaming Data Pipeline
News sources
Kafka / Flink
Deduplicate
Store & enrich
Checkpoint recovery · idempotent sinks
CMU Cloud Computing
Cloud Coursework
Building and managing large-scale distributed systems using cloud-native architectures.
Automated multi-tier infrastructure deployment with Terraform on AWS and Azure, managing repeatable cloud provisioning through infrastructure as code.
Implemented horizontal scaling and load balancing for web applications, with CloudWatch monitoring and automated fault-recovery strategies.

Cloud Coursework
View Course Syllabus
Academic Archives

Carnegie Mellon University
Business Intelligence & Data Analytics
The Hong Kong University of Science and Technology
Finance & Information System
Tech Stack
Systems & Data Engineering
Concurrent database systems, streaming pipelines and financial data platforms.
Agentic AI & Evaluation
Planning, tool orchestration and evaluation for enterprise and research workflows.
Machine Learning
Model development, feature engineering and benchmark-driven experimentation.
Software & Delivery
Backend APIs, interactive applications and repeatable testing and delivery.
Cloud & Analytics
Cloud provisioning, observability and analytics for operational decisions.
Global Recognition
Home Credit - Credit Risk Model Stability
Manipulated Polars dataframe, concatenated 9 CSV files by Personal ID
Reduced memory usage by 70%, selected 465 features for faster processing
Used five-fold StratifiedGroupKFold cross-validation for grouped samples
Developed ensemble voting model leveraging LightGBM and CatBoost
Optiver - Trading at the Close
Engineered statistical/financial features like imbalance, momentum, and price pressure
Accelerated feature computation with Numba JIT compilation
Ensembled 5-fold cross-validated CatBoost, LightGBM, and XGBoost models
Achieved MAE score of 5.3458 with post-processing strategies
// End of Competitive Log // Data Science Excellence

