Hi internet friends, I recorded a workshop about building your own LLM without any math / ML prerequisites. It covers everything from machine learning fundamentals, deep neural networks, transformer architecture, and pre/post-training. The only prerequisite is being comfortable with learning through code & excel examples. Sampling Large Language Models Reverse Engineering Large Language Model Perceptrons: wx+b Activation Functions: ReLU, GELU, SwiGLU GPU Coding: PyTorch, torch.compile(), fused kernels, CUDA, Triton MLPs/FFNs: Multi-input, Multi-Layer Perceptrons, Feed-Forward Networks Loss Functions: Residual errors, RMSE, Cross Entropy, Loss Landscapes Backpropagation: Training loops, Optimizers, Learning Rate, Batch Size Saving & Loading Models Initialization: Kaiming, Glorot Residuals: Addition, Scaling, Gated, Concatenation Normalization: Pre-norm vs. Post-norm, RMSNorm, BatchNorm, LayerNorm Regularization: Dropout, Gradient Clipping, Weight Decay SoftMax Tokenizers: By Character, By Word, BPE, SentencePiece Embeddings: Absolute vs. Learned, Sinusoidal vs. RoPE Attention: MHA, GQA, MQA, MLA Transformers Pre-training: Data Sources, Datasets, HTML Cleaning, Quality Filtering, Sharding Evaluation: Leaderboards, Benchmarks, Verifiers vs LLM-as-Judge Instruction Tuning: Alpaca & Other Formats, Self Instruct, Capabilities Reinforcement Learning: Policy Optimization, SimPO What We Didn't Cover: Scaling Each section has slides teaching the concepts, followed by excel-by-hand developing intuition for the math, and then coding examples. The goal is able to grok all parts of modern LLM development. We did this workshop in-person in San Francisco last month and hopefully the spaciousness of watching online works for everyone. If don't like watching videos, you can get the slides and exercises and work self-paced. submitted by /u/JustinAngel [link] [Kommentare]
Need help with my servos. Using DS3240 MG servo motors. Getting weird jitter sometimes, sometimes not. Sounds very scratchy. Power supply is definitely strong enough. Happens with and without load. Signal wires seem far enough from noise sources. The motors have stalled once when the arms of the two collided, and I'm thinking that the gears got damaged because of that. Although I think it's unlikely because these motors are designed to be able to hold a max load of 40kg submitted by /u/YengaJaf [link] [Kommentare]
as per the title, how to access books3 dataset for research purposes? submitted by /u/xolmnyc [link] [Kommentare]
Practice real interview problems on Multi-Agent Systems, RAG, Vector Databases, and production AI architectures.
📺 View the presentation slides When television was first invented, they would film people reading radio dramas in front of microphones. People only know how to use a new medium fr...
Start free with Kalshi Agent. Foundation + three tracks: Operator, Quant, and Builder — from Python Kalshi bot to live algorithmic prediction-market strategies.
Download PDF - Hidden Order: How Adaptation Builds Complexity (helix Books) [PDF] [7fb3c54faic0]. The book begins with a bunch of statistical formulas, but don't let that throw you. This is an extremely readable book o...
Home Biography The Complete Works Slideshow Sitemap Links Contact Boulevard Des Capucines Order a Hand-Painted Reproduction of this Painting San Giorgio Maggiore At Dusk Order a Hand-Painted Reproduction of this Painting The Garden At Argenteuil Aka The Dahlias Order a Hand-Painted Reproduction of this Painting The Luncheon (Monet's Garden At Argenteuil) Order a Hand-Painted Reproduction of this Painting The Red Boats, Argenteuil Order a Hand-Painted Reproduction of this Painting Self Portrait Oscar-Claude Monet (1840-1926) Oscar-Claude Monet (1840-1926) is a famous French painter and one of the founders of the Impressionism movement along with his friends Renoir, Sisley and Bazille. Monet rejected the traditional approach to landscape painting and instead of copying old masters he had been learning from his friends and the nature itself. Monet observed variations of color and light caused by the daily or seasonal changes. Claude Monet was born on November 14, 1840 on the fifth floor of 45 rue Laffitte,in the ninth arrondissement of Paris. He was the second son Claude Adolphe Monet and Louise-Justine Aubree. On the first of April 1851, Monet entered the Le Havre secondary school of the arts. He became known locally for this charcoal caricatures, which he would sell for ten to twenty francs. Monet also undertook his first drawing lessons from Jacques-Francois Ochard, a former student of Jacques-Louis David. On the beaches of Normandy in about 1856/1857 he meet fellow artist Eugéne Boudin who became his mentor and taught him to use oil paints. Boudin taught Monet "en plein air" (outdoor) techniques for painting. Click here for more 12 pictures per page 24 pictures per page 48 pictures per page 64 pictures per page 96 pictures per page Popularity Alphabetical Page 1 of 170 | Paintings: 2,029 Previous 123456789 Next Last A Cart On The Snow Covered Road With Saint Simeon Farm Pathway In Monets Garden At Giverny Impression Sunrise A Haystack The Water Lily Pond Aka Japanese Bridge Argenteuil (Red Boats) Irises In Monets Garden San Giorgio Maggiore At Dusk The Walk Woman With A Parasol A Farmyard In Normandy A Field At Gennevilliers A Woman Reading 12 pictures per page 24 pictures per page 48 pictures per page 64 pictures per page 96 pictures per page Popularity Alphabetical Page 1 of 170 | Paintings: 2,029 Previous 123456789 Next Last View all 2,029 Works
Knowledge Workers ≠ Developers The AI industry optimizes for developers. Frontier models are benchmarked on code generation, competitive math, and multi-step agentic reasoning — tasks where raw capability is the bottleneck and cost is secondary. That makes sense for developers: they write novel code, debug complex systems, and need the model to think as hard as possible. But knowledge workers — the hundreds of millions of people in spreadsheets, email, and documents every day — have structured, domain-specific tasks where speed and cost matter more than ceiling capability. They draft reports, build trackers, write formulas. The ceiling on most of these tasks is not model intelligence; it's context, speed, and reliability. This distinction has massive economic implications. If 80% of knowledge-worker requests can be served by a model that costs 10× less and responds 2× faster, defaulting every request to a frontier model isn't a quality strategy — it's a waste strategy. Core Thesis Most knowledge-worker tasks sit well within the capability of small, domain-tuned models. The right architecture is not "always use the best model" — it's "always use the right model", selected automatically by a lightweight router. The Proof: #2 on GDPVal With a Nano Router GDPVal is OpenAI's benchmark for real-world knowledge work — 220 tasks across 44 occupations (accountants, financial managers, engineers, clerks), each graded by human experts against professional deliverables. The GDPval-AA leaderboard by Artificial Analysis ranks 368 model configurations on these tasks. We built a nano-model-based router that classifies each task with a sub-cent nano-class model and dispatches to either GPT-5.5 (for hard tasks) or GPT-5.4 Mini (for everything else). It reaches #2 overall: #ModelELOClass 1GPT-5.5 (xhigh)1769Frontier 2Nano-Routed (GPT-5.5 + GPT-5.4 Mini)1759Router 3Claude Opus 4.7 (max)1753Frontier 4Claude Sonnet 4.6 (max)1676Frontier 5GPT-5.4 (xhigh)1674Frontier 6MiMo-V2.5-Pro1571Mid-tier 7DeepSeek V4 Pro (Max)1554Mid-tier 14GPT-5.4 mini (xhigh)1417Small 19Gemini Flash1197Small GDPval-AA ELO Leaderboard (selected, June 2026). Source: Artificial Analysis. GPT-5.4 Mini alone scores 1417. GPT-5.5 alone scores 1769. The nano-routed combination lands at 1759 — within 10 points of pure frontier — by using the cheap model wherever it's good enough and the expensive one only where it matters. It beats Claude Opus 4.7 and every other single-model entry. The cost difference between GPT-5.5 and GPT-5.4 Mini is over 10×, but the routed quality loss is just 10 ELO points. The architecture is simple: 📝 Task User request arrives → Nano Classifier