How close does Jev get to task-specific models?
An empirical comparison of general-purpose zero-shot inference with task-specific supervised models.
Projects, experiments, writing, and notes from over the years. Dates mark when each artifact was first published or made public.
An empirical comparison of general-purpose zero-shot inference with task-specific supervised models.
Reproducible comparisons of Jev against supervised and zero-shot baselines across five NLP benchmarks.
A push-based breadth-first search solver and visualizer for Sokoban puzzles.
Extending a supervised fine-tuning experiment with reinforcement learning for simpler technical answers.
A minimal LangGraph ReAct agent evaluated on 250 natural-language SQL questions.
Lessons from using a personal AI agent for learning, household workflows, and small everyday problems.
Complex language-model concepts explained in ten words or fewer.
A quick guide to decoding GPU names and comparing the specifications that matter for a workload.
Fine-tuning a small language model to produce answers that are both correct and simple.
An experiment in post-training Qwen3.5-4B to write clear, correct technical answers in simple English.
Small, focused PyTorch exercises for learning how large language models work.
A growing, linked knowledge base built from AI newsletters and personal research.
A walkthrough of building an end-to-end production machine-learning pipeline with TensorFlow Extended.
The practical data-science reading list I wish I had from day one.
How I think, collaborate, communicate, and what I value when working with a team.
Using agent-based models and GANs to examine how competition and cooperation can coexist.
What rapid prototyping and continuous feedback at hackathons reveal about building products.
A reproducible setup for carrying a personal development environment across Linux and macOS.
A bot that monitors scarce driving-test appointments and sends timely Slack notifications.