Tathagata Ghosh
I make tensors smaller and clusters busier.
whoami
- Tathagata Ghosh
- M.Tech CDS, IISc Bangalore (2025–27)
- HPC · CUDA · MPI · multi-GPU tensor SVD
- Google printed 'Mr. Nasamajh' on my Hash Code certificate. It means clueless. I kept it.
top
90x HYDRA all-pairs shortest paths
vs my re-implementation of Sao et al.'s 2-D scheme on ego-Facebook: 217.93 s vs 2.42 s, same cluster0.888 scaling efficiency, 16 GPUs
thesis: 1024³ tensor, FP32, 14.20x on 16 V100 GPUs, two nodes3.12x GPU vs 72 CPU threads
LCS of two 10⁸-char strings: CUDA 0.649 s vs OpenMP 2.028 s8.7e-15 energy identity error, FP64
thesis: Theorem 4 / Corollary 6 in FP64 (2.3e-5 in FP32)
cat thesis.md
tSVDm: exact star-M tensor SVD on 16 V100 GPUs
MATRIX Lab, IISc (ACM Gordon Bell Prize-winning lab) · Advisor: Prof. Phani Motamarri · with Argonne National Laboratory · Mar 2026–present
- C++17/CUDA, MPI + an NCCL ring over NVLink, FP32 and FP64
- 14.20x on 16 GPUs (efficiency 0.888), 1024³ tensor, FP32
- energy identity to 8.7e-15 (FP64); 4096³ FP64 ran in 2719 s
ls projects/
-
HYDRA 90x
All-pairs shortest paths on MPI + OpenMP + CUDA.
- Density-aware dispatch sends each 32x32 tile to GPU or CPU.
- Won 8 of 14 runs against 8 re-implemented methods, up to 90x.
- Matches the sequential reference exactly; all 6 losses reported.
- MPI
- OpenMP
- CUDA
-
LCS on GPU 3.12x
Longest common substring of two 10⁸-character strings.
- No O(NM) table: suffix arrays by prefix doubling on the GPU.
- CUDA 0.649 s vs best 72-thread OpenMP 2.028 s.
- 30 trials plus 5 warm-ups; 95% CI ±0.0029 s.
- CUDA
- Thrust
- AVX2
-
Face-wise tensor product 18.66x
Four backends for the product behind t-SVD.
- Serial, OpenMP, CUDA and hybrid CPU+GPU backends.
- Up to 18.66x over serial (OpenMP); 3.71x CUDA end to end.
- Measured across 488 timed runs.
- C++17 templates
- OpenMP
- cuBLAS
-
tinyinfer 5.8x
LLM inference engine from scratch in PyTorch + Triton.
- KV cache, GQA-aware Triton decode attention, INT8/INT4.
- CUDA-graph decode 4.8-5.8x faster than Hugging Face generate().
- Split-KV kernel at 1.68 TB/s, 99% of copy bandwidth.
- PyTorch
- Triton
- CUDA graphs
-
K-means + SpMV 21.90x
OpenMP K-means and MPI SpMV on a 760.6 M-nonzero matrix.
- K-means: 7.19x on 32 threads, same clusters as serial.
- Showed dynamic scheduling can lose to serial (0.38x).
- SpMV: 21.90x on 128 processes across 4 nodes.
- OpenMP
- MPI
- SLURM
-
DFT-FE on CLAP
Routed DFT-FE's 8 GEMM overloads through CLAP's 4.
- Configure-time switch; bit-identical to upstream when off.
- CLAP headers isolated behind one shim header.
- Verification and timing layer on real call shapes.
- C++
- CMake
- cuBLAS
-
Air-B-N-C
Write characters in thin air; edge devices read them.
- IMU, camera and depth branches each infer on the edge.
- A central node fuses them over 62 character classes.
- Team of 3; I built the IMU firmware and vision classifier.
- Arduino
- PyTorch
- TinyML
-
SFT, DPO, GRPO
Post-trained Qwen2.5-0.5B on GSM8K math word problems.
- SFT, DPO and GRPO objectives written by hand in PyTorch.
- GRPO: 34.6% to 44.7% greedy on all 1,319 test problems.
- Traced a 5.5-point DPO drop to likelihood displacement.
- PyTorch
- transformers
- SLURM
-
Wrist IMU HAR
Activity recognition from a wrist IMU on a Nicla Vision.
- MicroPython firmware at a drift-corrected 60 Hz.
- 12 activities, 26,819 windows, 126 features.
- Chose a 134 KiB, 94.48% tree over a 15 MB forest.
- MicroPython
- scikit-learn
- Optuna
-
VOC vision
Multi-label classification and segmentation, VOC-style.
- PCA from scratch: 170 of 3072 components keep 95%.
- ResNet50 with an FPN-style head: val F1@0.5 = 0.812.
- ConvNeXt + DeepLabV3+: mIoU 0.604 vs 0.446, p = 7.1e-26.
- PyTorch
- timm
-
Video2PDF
Lecture videos to deduplicated slide PDFs.
- Slide changes by pixel-difference ratio, 5-frame stability.
- Near-duplicates dropped by MSE similarity.
- Local files or YouTube URLs.
- Python
- OpenCV
-
ISDC 2026 site
TypeScript site for IISc's Durga Puja (ISDC 2026).
- Built while on the committee, finances included.
- Also built the CDS farewell 2026 site.
- TypeScript
- HTML/CSS
-
Trojan GNN
Golden-reference-free hardware Trojan detection with GNNs.
- GNNs on data-flow graphs of RTL and gate-level netlists.
- Clean and Trojan-inserted 16-bit ALU in Verilog.
- Self-reported: 97% recall at RTL, 84% at gate level.
- GNNs
- Verilog
- TrustHub
-
Logic locking
Logic locking a 16-bit ALU against Trojan insertion.
- Cadence Genus/Tempus scripts flag likely insertion sites.
- XOR/XNOR key gates harden them against SAT attacks.
- B.Tech final-year project, graded A+.
- Verilog
- Cadence
- Python
-
Face login
Log in with facial key-points instead of a password.
- Live detection for higher security.
- Python/OpenCV back end, React + Bootstrap front end.
- OpenCV
- React
-
IISc innovation archive
A queryable archive of IISc's innovation history.
- Five pillars, from startups to eminent faculty.
- NotebookLM-grounded workflow into one master database.
- Owned three divisions; interviewed a retired professor.
- Web scraping
- NotebookLM
history
- 2020-08B.Tech IT begins, IIEST Shibpur
- 2023-05Summer research intern, IIEST Shibpur
- 2024-05B.Tech done, CGPA 9.0
- 2024-07Software Engineer-1, RMES India
- 2025-08M.Tech CDS begins, IISc Bangalore
- 2026-03Thesis: tSVDm, MATRIX Lab
- 2026-05Data Science Intern, IBM India
- 2027-06M.Tech ends
IIEST Shibpur
B.Tech IT, 2020–2024
CGPA 9.0, First Class
School: ISC 97.75%, ICSE 96.4%
IISc Bangalore
M.Tech CDS, 2025–2027
CGPA 8.6 · MATRIX Lab
Thesis: tSVDm, Mar 2026–present
Software Engineer-1, RMES India (Rugged Monitoring) 2024–25
C#/.NET real-time monitoring software
for high-voltage electrical assets
Data Science Intern, IBM India 2026
Win-likelihood models for four products
up to 5.7x win-rate lift over baseline
cat skills.txt ranks.txt
- HPC: CUDA, MPI, NCCL, OpenMP, Nsight / TAU, Roofline
- AI / ML: PyTorch, XGBoost + SHAP, GNNs, SFT / DPO / GRPO, Optuna, TinyML
- Systems: C++17, C# / .NET / WPF, Python, TypeScript, Verilog, Linux / Bash
- AIR 663 GATE DA 2025
- 1,668,149 pts Hash Code 2022, qualification
- rank 13,429 Kick Start 2021, Round A
cd ~/offduty
Cycling · Chai · Sleeping · Adda (long, unhurried conversation with friends (Bengali))
- Hash Code certificate said 'Mr. Nasamajh' (clueless). I kept it.
- I keep the ledger for the campus cat welfare fund.
- Mess President for 800+ residents. Lunch is harder than consensus.
- Asked for 5 timing runs. I ran 30, plus 5 warm-ups.
- Found two errors in the published pseudocode I was implementing.
- Built the Durga Puja website, then argued with its invoices.
- Proved OpenMP dynamic scheduling can lose to serial: 0.38x.
- Reimplemented eight rival APSP methods. Published all 6 losses.
exit
No email here by design; LinkedIn messages reach me.