I am a fourth-year Ph.D. student in Computer Science at the University of Chicago and a MongoDB PhD Fellow, advised by Prof. Aaron J. Elmore in close collaboration with Dr. Goetz Graefe (Google). I build the next generation of data systems for cloud workloads, advancing memory and compute efficiency through paged query execution with fine-grained spills, fast address translation via address hints that narrows the disk-vs.-memory gap, and memory-efficient, skew-resilient sorting. My research has appeared at VLDB and CIDR.
Prior to my doctoral studies at UChicago, I completed my Bachelor’s degree in Aerospace Engineering at the University of Tokyo.
| Google Scholar | GitHub |
Jul 2026 — Paper accepted to VLDB 2026: “CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort”. See you in Boston!
Jun 2026 — Joining Google Spanner team as a Software Engineer Intern for Summer 2026.
Apr 2026 — Selected as a MongoDB PhD Fellow. [UChicago news]
University of Chicago, Sep 2022 - Present
PhD in Computer Science
Advisor: Prof. Aaron Elmore
Close Collaborator: Dr. Goetz Graefe (Google)
Transitional MS Degree was awarded in Mar 2025.
University of Tokyo, Mar 2017 - Mar 2022
Bachelor in Aerospace Engineering
Advisor: Prof. Takehisa Yairi
Uppsala University, Aug 2019 - Jun 2020
Exchange Student
Riki Otaki, Charles Benello, Fuheng Zhao, Aaron J. Elmore, and Goetz Graefe
CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort
To appear at Very Large Data Bases (VLDB), 2026
Riki Otaki, Jun Hyuk Chang, Aaron J. Elmore, and Goetz Graefe
Enhancing Transaction Processing through Indirection Skipping
Very Large Data Bases (VLDB), 2025
Riki Otaki, Jun Hyuk Chang, Charles Benello, Aaron J. Elmore, and Goetz Graefe
Resource-Adaptive Query Execution with Paged Memory Management
Conference on Innovative Data Systems Research (CIDR), 2025
Rui Liu, Jun Hyuk Chang, Riki Otaki, Zhe Heng Eng, Aaron J. Elmore, Michael J. Franklin, and Sanjay Krishnan
Towards Resource-adaptive Query Execution in Cloud Native Databases
Conference on Innovative Data Systems Research (CIDR), 2024
CrocSort: Memory-Efficient, Skew-Resilient Parallel External Sort (2024–2025)
Sorts 200 GB stably with only 2 GB of RAM; Postgres, DuckDB, and ClickHouse take >2× longer, time out, or abort. Driven by a configuration-first sort planner that derives minimum memory and thread allocation via analytical modeling and experiments. 2× faster merge under skewed keys and payloads via novel sparse-index range partitioning that balances both key counts and I/O volume across workers (vs. key-range partitioning). ~30% lower merge cost for multi-pass sorts via Offset-Value Coding + Tree-of-Losers for faster per-record comparisons.
LIPAH: Bridging Disk and In-Memory Transaction Processing (2023–2025)
Up to 19.7× speedup on TPC-C-like workloads with 40 threads via combined index and buffer-pool skipping (index alone: 1.3× over BP skipping)—substantially narrowing the disk-vs.-memory throughput gap. LIPAH (Logical ID with Physical Address Hinting) is a generalized fast-path skipping technique derived from pointer-swizzling that uses stale hints with cheap validation at lookup, instantiated as index skipping and BP skipping with a concurrent Foster B-Tree as the index target.
Query Execution with Paged Memory (2022–2024)
A pipelined execution engine where intermediate results live in the buffer pool as paged memory, enabling fine-grained spills, query suspension/resumption, and agile resource reallocation across operators. 15 of 22 TPC-H queries run within 1.5× of the non-paged baseline (slowdowns isolated to LIKE/regex-bound queries). Includes a logical optimizer with correlated-subquery unnesting (O(n²) → O(n)) and filter/projection pushdown.
Software Engineer Intern, Google — Sunnyvale, CA (Summer 2026)
Part-time Engineer, Preferred Networks — Tokyo, Japan (2022)
Developed a storage engine from scratch for Optuna, a hyperparameter optimization framework, enabling use without access to RDBMS—particularly beneficial on supercomputers. Supported distributed access to the storage via Network File System (NFS), allowing parallel tuning jobs across multiple nodes.
Last updated: Jul 5, 2026