#001

The Refusal Stack

Mechanistic interpretability of refusal in Llama-3.1-8B-Instruct: finding where refusal lives, removing it three different ways, and detecting the removal afterwards.

How it works

Refusal can be stripped from an open-weight model by a quiet fine-tune. Whether that leaves a readable signature depends on where you look, and the obvious place, the prompt, reads chance.

Located refusal as a single linear direction by diff-of-means over harmful and harmless prompts, with layer 10 chosen causally by ablation: removing that one direction drops refusal from 92.5% to 0% while leaving benign behaviour intact (KL 0.17), replicating Arditi et al. 2024 on hand-written PyTorch forward hooks. A three-threat-model attack ladder separates search limits from real robustness: discrete GCG 50%, continuous-embedding 90%, activation ablation 100% ASR. Reproducing GCG at 95.1% ASR on a Vicuna-7B control surfaced a sampler bug that had been misreading 0% on Llama-3.1. A covert LoRA fine-tune then strips refusal (98.8% → 0%) and a generation-time probe catches the tampering at AUROC 0.956, where a last-prompt-token probe reads 0.50. Refusal is decided during generation, not at the prompt. That probe holds AUROC 1.000 with zero drop against an adaptive obfuscation attack that jailbreaks the model 94% of the time and drives its detector penalty to ≈0. The same detector, unchanged, flags a triggered sleeper-agent backdoor at 0.965 and emergent misalignment at 0.77; a refit at the same site separates a sandbagging organism (ARC-Easy 91.5% → 13%). Refusal turns out to be a ~3–8-dimensional subspace rather than one feature (detection peaks at AUROC 0.94, k=3; a Llama-Scope SAE’s best single feature reaches cosine 0.31). Whole locate→break→detect pipeline runs on one A40-class GPU via ephemeral RunPod orchestration, Dockerised and CI-gated behind 255 CPU unit tests, with a frozen results snapshot so every headline number is checkable without a GPU.

PyTorchHugging FaceMechanistic InterpretabilityLoRARunPod / Docker
#002

PromptCoach

CLI + Chrome extension that coaches developers toward efficient AI prompting. UNSW CSESoc Flagship Hackathon finalist (team project).

How it works

AI-assisted coding wastes tokens on bloated prompts with no feedback loop. PromptCoach turns that repeated waste into a one-time cost.

One shared analysis core behind a CLI and Chrome extension (TypeScript/Next.js + FastAPI, Drizzle ORM on Cloudflare D1). Coaching is opt-in via a review: prompt prefix, so ordinary prompts never leave the machine; a local bridge falls back to offline heuristics. Hosted analysis runs through Anthropic’s Batches API on Haiku at ~2% of a frontier model’s energy per task, with environmental impact reported as a sourced, uncertainty-labelled range. Local-first, no telemetry.

TypeScriptNext.jsFastAPICloudflare D1Claude API
#003

IMC Prosperity 4: Algorithmic Trading

Solo competitor in IMC Prosperity 4, finishing top 10% worldwide and top 200 in Australia out of 22,000+ global teams across 5 rounds of algorithmic and manual trading.

How it works

The challenge: keeping a market-making engine profitable across changing volatility regimes without overfitting to any single one.

Built a three-tier market-making engine (take/clear/make) using Welford online mean, online AR(1) on price deviations, z-score tiered sizing, and asymmetric bid/ask anchoring. Built a separate trend-following MM with hardcoded-slope discovery, online OLS blending (70/30), and full-book order imbalance microprice adjustment.

PythonAlgorithmic TradingMarket MakingStatistics
#004

Limit Order Book & Matching Engine

A single-symbol limit-order-book matching engine in C++20 that enforces strict price-time (FIFO) priority. Validated byte-for-byte against replayed NASDAQ ITCH market data and compiled to WebAssembly for a live in-browser demo.

How it works

A matching engine is only useful if it’s provably correct and fast. The real constraint: cut latency without changing a single matched trade.

Rewrote the book around a flat O(1) price-ladder array and an intrusive, free-list-allocated order pool (64-byte cache-aligned levels, cached best bid/ask) in place of std::map + std::list, cutting median per-order latency from 538 ns to 344 ns (~1.5×) and lifting throughput 1.2–2.7× (3.2→4.0 M ops/s) on an identical 2M-operation workload. Proved the cache-optimised book byte-for-byte identical to the naive std::map/std::list oracle via a 20,000-op randomised differential test, and validated it against replayed NASDAQ TotalView-ITCH 5.0 by diffing depth snapshots against an independent Python reconstruction over 2.8M real messages.

C++20WebAssemblyLow-LatencyNASDAQ ITCHGoogleTest
#005

CFR Poker Bot

A Counterfactual Regret Minimization (CFR / CFR+) solver for Kuhn and Leduc poker, driven to a near-Nash game-theory-optimal strategy and validated against poker’s closed-form solution. A live browser bot lets you play it.

How it works

Most hobby CFR repos print a strategy and stop. This one checks its own answer: Kuhn poker is solved in closed form, so a correct solver has to reproduce −1/18. This one does.

Implemented CFR and CFR+ from the original papers over a game-agnostic extensive-form engine (12 Kuhn / 288 Leduc information sets), driving exploitability to 9×10⁻⁴ / 1.5×10⁻³ chips/game via a best-response evaluator, landing inside the band of DeepMind’s OpenSpiel reference implementations. Recovered Kuhn’s analytic game value of −1/18 to within 4×10⁻⁷ and its Nash signature (the opener bets the King exactly 3× as often as the Jack, never the Queen). Regret-matching⁺ with alternating, linear-averaged updates reaches the same exploitability ~10× faster than vanilla CFR, and the solved strategy beats fixed baselines by +38 to +72 bb/100 over 100k self-play hands.

PythonGame TheoryCFR / CFR+pytestCloudflare Pages
#006

Streaming Maze Engine

Full-stack maze platform with a C++20 generator hitting ~38 Mcells/s single-threaded (≈2× a published C# baseline) and ~92 Mcells/s on 8 cores, streaming 10-billion-cell mazes in O(width) memory, with real-time multiplayer and a WebGL2 renderer.

How it works

Generating a 10-billion-cell maze at full resolution should take terabytes of memory. Getting it under tens of MB while supporting thousands of real-time players is the actual constraint.

Eller’s algorithm with within-row strip parallelism (std::barrier) and AVX2/BMI2 SIMD wall packing (serialisation CPU share: 40% → <5%). FastAPI + Redis pub/sub broadcasts WebSocket moves across K8s pods (p99 RTT 3.3 ms at 100 bots). WebGL2 corridor renderer: one GPU draw call per 64×64 chunk, 512-entry LRU buffer cache, 60 fps with zero React re-renders. One-command AWS deploy via Terraform + EKS + RDS + ElastiCache + S3 + CloudFront. ≥70% pytest + GoogleTest coverage enforced in CI.

C++20AVX2 / SIMDPython / FastAPIReact / WebGL2RedisAWS / Kubernetes
#007

Slide Games

Python framework (published on PyPI) that compiles arcade game logic into fully playable Google Slides via BFS state enumeration: one slide per reachable state, hyperlink-navigated.

How it works

No runtime, no JavaScript, no server. Just a shareable URL that plays a full arcade game.

1,000-state ceiling bounds exponential growth (Pac-Man scales as positions × 2ⁿ with n pellets). Token-bucket rate limiter at ≤50 API writes/min with 5 concurrent batch uploads generates ~500-state presentations in 1–3 min. Pygame-inspired 1920×1080 rendering API with 40+ colours and 3 themes; campaign system across 4 bundled games (491–600 states each).

PythonGoogle Slides APIBFSPyPI
#008

PixelVault

Python file-to-video codec that encodes any file into MP4 for lossless storage on YouTube, recovering the original file exactly despite H.264/VP9 re-encoding.

How it works

YouTube lossy-compresses every uploaded video. Storing arbitrary binary data there without any corruption is the problem.

2×2 uniform pixel blocks survive ±127 DCT luma ringing in H.264/VP9. Three-tier Reed-Solomon ECC over GF(2⁸) (vectorised → Berlekamp-Massey → parallel) with byte interleaving for burst-error recovery. AES-256-GCM + PBKDF2-SHA256 encryption, zlib compression, and 38× faster encoding at 0.82–2.88 MB/s via NVENC/AMF/QSV hardware acceleration with YouTube OAuth2 upload.

PythonFFmpegReed-Solomon ECCYouTube APINumPy
#009

ASCII / Unicode Art Converter

Zero-dependency, fully client-side ASCII/Unicode art converter: 7 character modes, 3,163-codepoint Unicode pool, 123-emoji mosaic, image/video/webcam input.

How it works

The problem: faithfully mapping full-colour images to text characters at 30 fps without losing perceptual detail.

O(1) nearest-colour lookup via a precomputed 32³ = 32,768-entry RGB quantisation table enables 30 fps video at ~2.1 ms/frame. Implements Floyd-Steinberg, Atkinson, and Bayer dithering, Sobel edge detection, and a particle drift system. Supports 6 export formats (PNG, SVG, TXT, WebM via MediaRecorder API) and reports live render time at 1920×1080.

TypeScriptCanvas APIAstro
#010

Quantum Random Number Generator

QRNG using all-Hadamard circuits on Qiskit, with bit outcomes governed by Born-rule probability, a full NIST SP 800-22 / 800-90B statistical pipeline, and an MT19937 state-recovery attack demonstration.

How it works

Statistical tests can't distinguish a good PRNG from a quantum source. The real distinction: a PRNG's full state is recoverable in 624 outputs; a quantum source has no state to recover.

An n-qubit all-Hadamard circuit yields n bits with p=0.5 per bit; by Bell’s theorem no hidden variable predicts the outcome. Implements 8 NIST SP 800-22 tests from scratch + full 15-test battery, fixing 2 bugs in nistrng (shared-array mutation and incorrect LC binning that caused good random data to fail). SP 800-90B MCV/Markov min-entropy estimators; MT19937 state-recovery in 624 outputs; quantum-seeded AES-256-GCM encryption. Runs on Qiskit simulator or IBM Quantum hardware.

PythonQiskitIBM QuantumNIST SP 800-22Streamlit
#011

Black Scholes Option Calculator

Real-time options pricing engine with interactive visualisation of volatility and time-decay across 2,500+ scenarios.

How it works

Options pricing surfaces are non-linear and hard to reason about statically. The goal: interactive visualization across 2,500+ scenarios that updates instantly.

Built a real-time pricing engine using the Black-Scholes-Merton formula with vectorized NumPy/SciPy computations, 3D Matplotlib surfaces, and Seaborn heatmaps, simulating 2,500+ scenarios instantly in Streamlit.

PythonStreamlitQuantitative Finance
#012

Palimps: Stochastic Text Generation

Infinite, context-aware prose generation from any text corpus, without a neural network.

How it works

Markov chains lose context fast. Getting coherent, unbounded prose without a neural network means working around that.

Built an n-gram Markov Chain engine with a dynamic backoff strategy for zero-probability states, NLTK POS tagging for grammatical coherence, and binary serialization that cuts initialization time by 85%.

PythonNLTKMarkov Chains