dpaste.com is a pastebin site for easily sharing and storing code snippets. Syntax highlighting, clean interface, markup preview, quick sharing options.
It’s incredibly hard to deliver drugs to the right organ, especially to reach the brain. Tiny gas-filled spheres that burst on command could change that.
The Thames may be cleaner than when it was declared biologically dead in 1957, but other rivers are close to ecological...
Google Drive kann privat mit einem Google-Konto oder geschäftlich mit einem Google Workspace-Konto verwendet werden.
I recorded all the "terms and conditions"1 that I've had to click "agree" to or otherwise claim to have read on a computer or phone for 10 years (starting April 21st 2016 and ending April 21st 2026). See the terms and conditions here(page may take several seconds to load) I started doing this because Ts and Cs feel like an aspect of modern life that doesn't work as it's meant to but which we ignore because the workaround is fine and doesn't cause too many problems. Most people regularly make the claim once every few days to have read and understood something they haven't read or understood, and that's probably fine. Most companies aren't slipping nefarious clauses into their Ts and Cs and if they did surely some consumer rights nerd would find it and flag it and it would get in the news and the company would be punished? I wanted to quantify the burden of our official obligation to read and comprehend all these Ts and Cs. ...this doesn't seem like an unreasonable ask when framed like this. A testament to how much we can achieve in the sum of 4 minutes a day over 10 years. Makes me think of all the books and wikipedia pages I could have read in the past 10 years... Is 238 words/min fair? The legalese would require some time (and possibly further research) to understand adequately enough that "agreeing" would be legally sensible. Also, the content is extremely dry making attention slip and requiring re-reading passages. Also the average word length is probably longer than what that 238 words/min number comes from. The average syllables-per-word is 1.74 for this Ts and Cs corpus vs. 1.62 for the Australian constitution, 1.56 for the US constitution, 1.38 for The Great Gatsby. I tried timing my reading and understanding of some passages: This seems consistent with the 200-300 or so words/min. Probably the biggest issue with estimating time taken to read/comprehend is that some of the Ts and Cs refer to other documents that in turn I must imply I've read (e.g. "I acknowledge that I have read the General Information sheet" and "[I acknowledge that...] my personal information may be used by police for general law enforcement purposes, including those purposes set out in the Australian Crime Commission Act 2002 (Cth)"). I didn't burrow into and include these but that could multiply this Ts and Cs corpus several-fold Here are the ten most common words across all 821 agreements (excluding simple stopwords like "the", "and", "to"): The five longest agreements in the corpus: Word counts from Project Gutenberg plain text editions. I'm sure there are Ts and Cs that I missed. Either i was time-pressured at work or in an airport or a clinic or something and didn't have time to copy and paste the Ts and Cs or record a link to them, or I lost them afterwards among other emails or notes. I would be suprised if the lost notes were more than 10% of the current corpus, so it wouldn't change the count much. Some terms and conditions I noted down but when looking later, I couldn't find them. For example, the policies for free wifi in Rome and Lisbon airports, Hotel Bellvedere. Some terms and conditions I couldn't find at the time (i.e. the link to read them was broken) but I still had to agree that I'd read them. I couldn't record these. I did not record any updates to Ts and Cs. E.g. my bank frequently notifies me (feels like weekly but is probably actually once every few months) of updated Ts and Cs and makes me agree to the updated ones. I didn't record these each time, only the first time. In 2018 when cookie banners on website became much more common (because some European GDPR rules came into effect or something), I developed an approach of avoiding having to agree/acknowledge these as much as possible. I would set display:none through the "inspect" function in the browser, or simply navigate away from the page and find another way of accessing the information. Sometimes though, I was forced to agree and in several cases I did not record these. I did not include agreements that I'd signed "in real life" E.g. signing work contracts or my mortgage agreement or conveyancing documents (It was only the agreements/Ts and Cs on software or the internet that I counted) Sometimes I only noted down which policy I'd agreed to rather than copy-pasting the whole thing at the time, and it may have been months to years later that I retrieved the contents of the policy, so that the date on the policy may apparently post-date when I actually agreed to it. There are a few hundred words that are my own additions that should be strictly subtracted from the word count, but it would be a negligable difference. E.g. my own annotations in: "Rome Airport Wifi Privacy Policy Rome airport wifi privacy policy [could not find]", which is a placeholder for the missing rome Ts and Cs 1 I use "Terms and conditions" or "Agreements" as a general term to cover all variants of this genre of legal document laying out an organisaiton's agreement/notice/terms with me. Synonyms as far as I'm concerned include: "Terms of Service", "Terms of Use", "End User License Agreement", "Software License Agreement", "Privacy Policy", "Privacy Notice", "Privacy Statement", "Privacy Pledge", "Cookie Policy", "Cookie Notice", "Acceptable Use Policy" and "Code of Conduct". ↩ 2Zya made the app Ditty.it, which I would have downloaded to make memes of a style that was popular in about 2016-2017. They're defunct now
Triage your inbox by priority, draft replies in your tone, track follow-ups automatically, and transcribe meetings. One AI layer across Gmail, Outlook, and the web.
Digital immortality is a myth, what can we do about it?
...also don't tell lies. But I'm getting ahead of myself already. I keep running into people online who openly say that they use AI to do their writing for t
VaultSort is the all-in-one solution for Mac users who want to take control of their digital files with military-grade security and intelligent organization.
Clear — An intent-first agentic development language. Write readable specifications. AI agents turn them into real software.
think again
How good are local LLMs at translation, and do you actually need the cloud? A reproducible benchmark of 24 on-device, self-hosted, and cloud models translating into English, with the low-resource case (Afrikaans) front and centre. The headline: on Afrikaans→English a local 18 GB model lands in a statistical tie with frontier cloud. Same blinded Tatoeba sentences, same prompt, greedy decoding, scored multi-reference with COMET (meaning) and chrF++ (surface). Built to pick a translation model for Lector. Open and reproducible: harness, every model's raw outputs, and the seeded test sets: github.com/heuwels/llm-lang-eval On Afrikaans the field is tightly bunched: 20 of 24 models fall within ~1.5 COMET (sampling noise) of the top score (≈95), a statistical tie, not a ranking. The self-hosted 18 GB gemma-4-12b-qat (95.0) sits in that band alongside frontier cloud, so for Afrikaans→English, you don't need the cloud or a big box. Leaderboard All 24 models across three languages on one zoomed COMET axis: each row a model (ranked by Afrikaans), one dot per language, the connector its cross-language spread. The zoom makes the tight differences legible; gaps under ~1.5 COMET are sampling noise (see Significance). Afrikaans German Spanish · COMET, zoomed axis 88 to 96; dot before each name = tier Deployment tier:on-device (laptop) self-hosted box (18 GB) cloud (OpenRouter) The numbers COMET (meaning, ×100) over chrF++ (surface), per language. chrF++ rewards character overlap with the reference, so it docks valid paraphrases ("scenery is magnificent" vs "landscape is breathtaking"); COMET scores meaning and credits them. Green = leading band (within ~1.5 COMET, a statistical tie, not a single winner). Rows share the chart's order (ranked by Afrikaans COMET), so near-ties can sit a place apart despite equal rounded scores. n=200 per language. ModelAfrikaansCOMET · chrF++GermanCOMET · chrF++SpanishCOMET · chrF++gpt-595.3chrF 83.593.1chrF 75.193.9chrF 78.1gemini-2.5-pro95.3chrF 83.093.5chrF 75.894.4chrF 80.3claude-opus-4.895.1chrF 82.193.3chrF 75.694.5chrF 80.5claude-sonnet-4.695.1chrF 81.293.1chrF 74.894.6chrF 81.3gemma-4-12b-qat95.0chrF 82.893.1chrF 74.593.6chrF 78.4gpt-4o-mini95.0chrF 81.392.9chrF 74.094.3chrF 79.3gpt-4o94.9chrF 81.093.4chrF 75.094.2chrF 79.3mistral-large94.8chrF 81.893.2chrF 75.394.0chrF 78.5llama-3.3-70b94.7chrF 82.392.9chrF 74.893.8chrF 78.1gemini-2.5-flash94.7chrF 82.592.4chrF 75.092.4chrF 77.4deepseek-v3.294.7chrF 80.992.6chrF 72.994.0chrF 78.3gemma-2-27b94.5chrF 79.791.2chrF 71.492.2chrF 76.6gemma-4-12b94.5chrF 81.593.0chrF 74.493.3chrF 77.5gemma-3-12b94.5chrF 80.292.2chrF 72.392.6chrF 75.2gemma-3-27b94.4chrF 79.692.6chrF 72.492.9chrF 75.7claude-haiku-4.594.3chrF 80.292.9chrF 74.894.4chrF 79.9ministral-3-14b94.3chrF 78.291.8chrF 71.593.5chrF 76.0gemma-4-e4b94.2chrF 80.392.2chrF 70.993.1chrF 77.8qwen3.5-9b94.1chrF 78.992.2chrF 71.893.0chrF 75.1mistral-small-3.2-cloud94.0chrF 79.892.8chrF 73.793.5chrF 77.5gemma-3n-e4b93.7chrF 77.092.3chrF 71.192.9chrF 75.0ollama-llama3.1-8b93.5chrF 79.391.6chrF 70.492.6chrF 75.7apfel-foundation92.1chrF 74.589.6chrF 64.190.9chrF 70.6qwen2.5-coder-14b91.8chrF 74.391.4chrF 68.693.0chrF 76.1 Cost, and what you can actually run Cloud models fill the top of that board, but this is a study of local models, and most of the cloud field can't run on the box at all. The frontier APIs (GPT-5, Claude, Gemini) are closed weights, so self-hosting them was never an option. The open models I could reach through OpenRouter are mostly too big for the hardware: Llama 3.3 is 70B, Mistral Large is larger again, and even the 24B and 27B open models (Mistral Small, Gemma 2 and 3 at 27B) sit at or past the ceiling of an 18 GB Mac once the OS and the KV cache take their share. What genuinely fits is the on-device and self-hosted-box tiers, so here is that field on its own. Afrikaans German Spanish · COMET, zoomed axis 88 to 96; dot before each name = tier The strongest model that actually fits, gemma-4-12b-qat at 7.5 GB, is the same one sitting in the frontier band up top. apfel-foundation is Apple's built-in Foundation model, the one that ships with macOS, run through the Apfel harness. It scores respectably on what it answers, but it refused or errored on 26% of the Afrikaans sentences (52 of 200), against 8% in Spanish, and Apple doesn't list Afrikaans among its supported languages, which is why it sits near the bottom. Cost: use what you've got Cost barely enters into it. A translation is tiny, roughly 80 tokens, so the entire cloud sweep (24 models across three languages, plus the holdout and the cloze probe) came to $13.62 on OpenRouter, well under a cent per translation even on the frontier models. Per-token pricing still spans an order or two of magnitude, the frontier APIs against the cheap tiers like Gemini Flash or GPT-4o-mini, but at this token count the absolute bill is small whichever way you go. What moves the decision is what you already own. A spare Mac is a sunk cost, so a local model is free per lookup beyond the electricity. An existing Claude plan is free at the margin too, within its limits, which is why I reach for the Anthropic OAuth route first. OpenRouter is the only one of the three that adds a real per-token bill, and it earns its place when you need a specific model you can't self-host or don't have a plan for. So the honest answer is usually to use what you've got: a spare box runs a local model, an existing plan already covers the lookups, and with neither, the cheap cloud tier is pennies per thousand. Since a 12B you can run at home already ties the frontier on Afrikaans, paying frontier rates per token buys very little for this particular job. Contamination check: does it survive on unseen data? The honest limitation. Tatoeba is almost certainly in every model's pretraining, so a high score can mean "translated well" or "regurgitated a memorised pair". The score alone can't tell us which. To bound it, each model is compared on two matched 150-sentence Afrikaans samples (same length filter): pre-2023 (added 2010 to 2022, almost certainly seen in training) versus 2025-26 (added after the training cutoff of the older-generation models here, so they cannot have memorised them). A large drop on the recent set is the fingerprint of memorisation; a stable score is evidence of genuine translation ability. Modelpre-2023COMET2025-26COMETΔgemma-4-12b-qat94.693.4-1.2gemma-3-12b94.093.2-0.8gemma-4-12b94.593.1-1.4gemma-3n-e4b93.892.5-1.4qwen3.5-9b93.392.1-1.2ministral-3-14b93.892.0-1.8gemma-4-e4b94.091.8-2.2qwen2.5-coder-14b90.388.9-1.4 Caveat on the caveat: exact training-cutoff dates aren't published for every model, and recently-added sentences may differ subtly in style or difficulty, so read a small Δ as "holds up", not as a precise measurement of contamination. Parroting probe: memorisation, measured directly The sharpest contamination test. We blank one informative word per sentence and ask each model to fill it. On unseen (2025-26) sentences it can only predict from context; if it recovers the exact original word much more often on seen (pre-2023) sentences, that gap is the model parroting memorised text rather than reasoning about the language. (It doubles as a cloze-ability score, Lector's own practice task.) Modelseenrecovery %unseenrecovery %gapclaude-opus-4.84231+11qwen3.5-9b107+4gemma-3-12b1715+2gemma-4-12b-qat1817+1gemma-3n-e4b67-1gemma-4-e4b67-1 Recovery = exact match of the blanked word. A large positive gap = memorisation; near-zero = genuine context prediction. n≈150 per cell, so gaps within ~±10 are noise. Side-by-side generations The numbers only say so much. Here are the actual translations where models disagree most. Green marks the highest per-sentence chrF++ for that sentence. Afrikaans: where the models splitSentences where models split into clear camps, several agreeing on one wording, several on another (one-off wordings collapsed to a tail). Count × tier-dots per camp; green = within ~6 chrF of the closest-to-reference camp.Why they differ: almost none of this is error. It's paraphrase choice. A contraction vs the full form, 'by the end' vs 'before the end', one valid synonym over another. Each camp diverges from the single crowd-sourced reference in its own way. The spread is widest on longer, structurally flexible sentences (more ways to order the English) and on the high-resource languages, which is exactly why the COMET (meaning) gaps are far smaller than the chrF (surface) gaps. Read the camps as equally-valid translations, not right-vs-wrong.afrNiemand het gekom nie.refNo one came. · Nobody came.12×100gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, llama-3.3-70b, gemini-2.5-flash, gemma-3-12b, gemma-3-27b, claude-haiku-4.5, gemma-4-e4b, gemma-3n-e4b, ollama-llama3.1-8bNobody came.11×100gpt-5, gpt-4o-mini, gpt-4o, mistral-large, deepseek-v3.2, gemma-2-27b, gemma-4-12b, ministral-3-14b, qwen3.5-9b, mistral-small-3.2-cloud, qwen2.5-coder-14bNo one came.+1 one-off wordings (apfel-foundation)afrDie meisie in die blou jas is my dogter.refThe girl in the blue coat is my daughter.11×82gpt-4o-mini, llama-3.3-70b, deepseek-v3.2, gemma-2-27b, gemma-3-27b, claude-haiku-4.5, ministral-3-14b, gemma-4-e4b, qwen3.5-9b, mistral-small-3.2-cloud, gemma-3n-e4bThe girl in the blue jacket is my daughter.10×100gpt-5, gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, gpt-4o, mistral-large, gemini-2.5-flash, gemma-4-12b, gemma-3-12bThe girl in the blue coat is my daughter.2×79ollama-llama3.1-8b, qwen2.5-coder-14bThe girl in the blue dress is my daughter.afrHierdie tomaties het geen smaak nie.refThese tomatoes don't have any taste.11×57gpt-5, gemini-2.5-pro, claude-opus-4.8, llama-3.3-70b, gemini-2.5-flash, claude-haiku-4.5, ministral-3-14b, gemma-4-e4b, mistral-small-3.2-cloud, ollama-llama3.1-8b, qwen2.5-coder-14bThese tomatoes have no taste.10×44gemma-4-12b-qat, gpt-4o-mini, gpt-4o, deepseek-v3.2, gemma-2-27b, gemma-4-12b, gemma-3-12b, gemma-3-27b, qwen3.5-9b, gemma-3n-e4bThese tomatoes have no flavor.+3 one-off wordings (claude-sonnet-4.6, mistral-large, apfel-foundation)afrOns het 'n groot tuin.refWe have a big garden. · We have a big yard.14×100gpt-5, gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, llama-3.3-70b, deepseek-v3.2, gemma-4-12b, gemma-3-12b, gemma-4-e4b, mistral-small-3.2-cloud, gemma-3n-e4b, ollama-llama3.1-8b, qwen2.5-coder-14bWe have a big garden.10×63gpt-4o-mini, gpt-4o, mistral-large, gemini-2.5-flash, gemma-2-27b, gemma-3-27b, claude-haiku-4.5, ministral-3-14b, qwen3.5-9b, apfel-foundationWe have a large garden.afrDit lyk asof ons sukkel met die selfde ou probleem.refWe seem to keep grappling with the same old problem.10×59gpt-5, gemini-2.5-pro, claude-sonnet-4.6, gemma-4-12b-qat, gemma-2-27b, gemma-4-12b, claude-haiku-4.5, qwen3.5-9b, gemma-3n-e4b, ollama-llama3.1-8bIt looks like we’re struggling with the same old problem.9×61claude-opus-4.8, gpt-4o, mistral-large, llama-3.3-70b, deepseek-v3.2, gemma-3-12b, ministral-3-14b, mistral-small-3.2-cloud, apfel-foundationIt seems like we're struggling with the same old problem.2×61gemini-2.5-flash, gemma-3-27bIt seems we're struggling with the same old problem.+3 one-off wordings (gpt-4o-mini, gemma-4-e4b, qwen2.5-coder-14b)German: where the models splitSentences where models split into clear camps, several agreeing on one wording, several on another (one-off wordings collapsed to a tail). Count × tier-dots per camp; green = within ~6 chrF of the closest-to-reference camp.Why they differ: almost none of this is error. It's paraphrase choice. A contraction vs the full form, 'by the end' vs 'before the end', one valid synonym over another. Each camp diverges from the single crowd-sourced reference in its own way. The spread is widest on longer, structurally flexible sentences (more ways to order the English) and on the high-resource languages, which is exactly why the COMET (meaning) gaps are far smaller than the chrF (surface) gaps. Read the camps as equally-valid translations, not right-vs-wrong.deuAlle meine Kinder wurden in Boston geboren.refAll of my children were born in Boston.12×87gpt-5, gemini-2.5-pro, claude-sonnet-4.6, gpt-4o-mini, gpt-4o, gemini-2.5-flash, gemma-2-27b, gemma-3-12b, gemma-4-e4b, mistral-small-3.2-cloud, ollama-llama3.1-8b, apfel-foundationAll my children were born in Boston.12×100claude-opus-4.8, gemma-4-12b-qat, mistral-large, llama-3.3-70b, deepseek-v3.2, gemma-4-12b, gemma-3-27b, claude-haiku-4.5, ministral-3-14b, qwen3.5-9b, gemma-3n-e4b, qwen2.5-coder-14bAll of my children were born in Boston.deuEr studiert an der Technischen Universität.refHe studies at the technical university. · He's studying at the technical university.12×65gpt-5, gemini-2.5-pro, gpt-4o-mini, gpt-4o, mistral-large, llama-3.3-70b, gemma-2-27b, gemma-3-12b, qwen3.5-9b, mistral-small-3.2-cloud, ollama-llama3.1-8b, apfel-foundationHe is studying at the Technical University.12×73claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, gemini-2.5-flash, deepseek-v3.2, gemma-4-12b, gemma-3-27b, claude-haiku-4.5, ministral-3-14b, gemma-4-e4b, gemma-3n-e4b, qwen2.5-coder-14bHe studies at the Technical University.deuDu darfst das Buch lesen.refYou may read this book.11×72gpt-5, claude-opus-4.8, gemma-4-12b-qat, mistral-large, llama-3.3-70b, gemini-2.5-flash, deepseek-v3.2, gemma-4-12b, gemma-3-27b, ministral-3-14b, gemma-4-e4bYou may read the book.11×38gemini-2.5-pro, claude-sonnet-4.6, gpt-4o-mini, gpt-4o, gemma-2-27b, gemma-3-12b, claude-haiku-4.5, mistral-small-3.2-cloud, gemma-3n-e4b, apfel-foundation, qwen2.5-coder-14bYou are allowed to read the book.2×37qwen3.5-9b, ollama-llama3.1-8bYou're allowed to read the book.deuNiemand hat Tom gerufen.refNobody called Tom.12×62gpt-5, gemma-4-12b-qat, gpt-4o-mini, gpt-4o, llama-3.3-70b, gemma-2-27b, gemma-4-12b, ministral-3-14b, gemma-4-e4b, qwen3.5-9b, mistral-small-3.2-cloud, qwen2.5-coder-14bNo one called Tom.11×100gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, mistral-large, gemini-2.5-flash, gemma-3-12b, gemma-3-27b, claude-haiku-4.5, gemma-3n-e4b, ollama-llama3.1-8b, apfel-foundationNobody called Tom.+1 one-off wordings (deepseek-v3.2)deuDie Polizisten spielten auf der Polizeistation Schach.refThe police officers were playing chess at the police station.13×100gpt-5, gemma-4-12b-qat, gpt-4o, mistral-large, llama-3.3-70b, gemma-4-12b, gemma-3-12b, gemma-3-27b, ministral-3-14b, qwen3.5-9b, mistral-small-3.2-cloud, ollama-llama3.1-8b, qwen2.5-coder-14bThe police officers were playing chess at the police station.11×78gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gpt-4o-mini, gemini-2.5-flash, deepseek-v3.2, gemma-2-27b, claude-haiku-4.5, gemma-4-e4b, gemma-3n-e4b, apfel-foundationThe police officers played chess at the police station.Spanish: where the models splitSentences where models split into clear camps, several agreeing on one wording, several on another (one-off wordings collapsed to a tail). Count × tier-dots per camp; green = within ~6 chrF of the closest-to-reference camp.Why they differ: almost none of this is error. It's paraphrase choice. A contraction vs the full form, 'by the end' vs 'before the end', one valid synonym over another. Each camp diverges from the single crowd-sourced reference in its own way. The spread is widest on longer, structurally flexible sentences (more ways to order the English) and on the high-resource languages, which is exactly why the COMET (meaning) gaps are far smaller than the chrF (surface) gaps. Read the camps as equally-valid translations, not right-vs-wrong.spaPodéis ver la televisión después de cenar.refYou can watch television after dinner.13×66gpt-5, gemma-4-12b-qat, gpt-4o, mistral-large, llama-3.3-70b, gemini-2.5-flash, gemma-4-12b, ministral-3-14b, qwen3.5-9b, mistral-small-3.2-cloud, gemma-3n-e4b, ollama-llama3.1-8b, apfel-foundationYou can watch TV after dinner.11×100gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gpt-4o-mini, deepseek-v3.2, gemma-2-27b, gemma-3-12b, gemma-3-27b, claude-haiku-4.5, gemma-4-e4b, qwen2.5-coder-14bYou can watch television after dinner.spaTenía 23 años de edad cuando pinté este cuadro.refWhen I painted this picture, I was 23 years old.13×83gemini-2.5-pro, claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, gpt-4o-mini, gpt-4o, mistral-large, llama-3.3-70b, gemini-2.5-flash, gemma-2-27b, gemma-3-12b, gemma-3-27b, gemma-3n-e4bI was 23 years old when I painted this picture.10×69gpt-5, deepseek-v3.2, gemma-4-12b, claude-haiku-4.5, ministral-3-14b, gemma-4-e4b, qwen3.5-9b, mistral-small-3.2-cloud, ollama-llama3.1-8b, qwen2.5-coder-14bI was 23 years old when I painted this painting.+1 one-off wordings (apfel-foundation)spaLa pelota de golf casi entró en el hoyo.refThe golf ball almost went in the hole.14×90claude-opus-4.8, claude-sonnet-4.6, gemma-4-12b-qat, gpt-4o-mini, gpt-4o, mistral-large, gemini-2.5-flash, deepseek-v3.2, gemma-4-12b, claude-haiku-4.5, ministral-3-14b, mistral-small-3.2-cloud, apfel-foundation, qwen2.5-coder-14bThe golf ball almost went into the hole.10×100gpt-5, gemini-2.5-pro, llama-3.3-70b, gemma-2-27b, gemma-3-12b, gemma-3-27b, gemma-4-e4b, qwen3.5-9b, gemma-3n-e4b, ollama-llama3.1-8bThe golf ball almost went in the hole.spaEstoy en casa de mis padres.refI'm at my parents' house.14×100gpt-5, gemini-2.5-pro, claude-opus-4.8, gpt-4o, mistral-large, gemini-2.5-flash, deepseek-v3.2, gemma-3-12b, gemma-3-27b, ministral-3-14b, qwen3.5-9b, mistral-small-3.2-cloud, gemma-3n-e4b, ollama-llama3.1-8bI'm at my parents' house.10×88claude-sonnet-4.6, gemma-4-12b-qat, gpt-4o-mini, llama-3.3-70b, gemma-2-27b, gemma-4-12b, claude-haiku-4.5, gemma-4-e4b, apfel-foundation, qwen2.5-coder-14bI am at my parents' house.spaSi es necesario, vendré mañana a las nueve.refIf necessary, I'll come at nine tomorrow.12×70claude-sonnet-4.6, gemma-4-12b-qat, gpt-4o-mini, deepseek-v3.2, gemma-2-27b, gemma-4-12b, claude-haiku-4.5, ministral-3-14b, gemma-4-e4b, mistral-small-3.2-cloud, apfel-foundation, qwen2.5-coder-14bIf necessary, I will come tomorrow at nine.10×83gpt-5, gemini-2.5-pro, claude-opus-4.8, gpt-4o, mistral-large, llama-3.3-70b, gemini-2.5-flash, gemma-3-27b, gemma-3n-e4b, ollama-llama3.1-8bIf necessary, I'll come tomorrow at nine.+2 one-off wordings (gemma-3-12b, qwen3.5-9b) Methodology Task. Blinded source → English. The model sees only the source sentence; references are held out for scoring. Data. Tatoeba sentence pairs (CC-BY 2.0 FR), seeded random sample, multi-reference where available, length-filtered. Prompt. One fixed user message, identical across models (no per-model tuning). Chain-of-thought disabled (reasoning_effort: none); translation needs none. Structured output. Every model is constrained to emit {"translation": "…"} via a json_schema response_format. This is the equaliser: small models otherwise "think out loud" in plain text and bury the answer in preamble. Constrained decoding makes that impossible, gives every model the identical constraint, and mirrors how Lector itself prompts. Decoding. temperature = 0 (greedy), one model resident at a time on an 18 GB host (JIT load/evict). Metrics. chrF++ and BLEU via sacreBLEU (signatures recorded). COMET planned. Caveats Significance: read bands, not ranks. n=200 per language (Afrikaans largely single-reference), so per-system COMET 95% confidence intervals are roughly 1 to 2 points either way. Differences below ~1.5 COMET are sampling noise: the green leading band is a statistical tie and the sort order within it is not meaningful. Per-segment bootstrap CIs are future work. This is a proxy. It measures general sentence MT, not Lector's actual word/phrase dictionary-lookup task, a strong signal for model choice, not "Lector's output graded." Contamination, the big one. Tatoeba is in these models' pretraining, so a high score can reflect memorising the pair rather than reasoning about the language, and the score alone can't separate the two. The contamination-check section above bounds this with a post-cutoff holdout; treat absolute scores with suspicion and weight the pre-vs-post deltas and relative gaps over the headline numbers. Into-English is the easy direction, and Afrikaans here is largely single-reference. Read accordingly.
This book has its genesis in the author’s wildly popular (and oft-downloaded) papers, “An Elementary View of Maxwell’s Displacement Current„ and “The Fundamentals of Electromagnetic Theory Revisited,” published in this Magazine in February 2008 and November 2009, respectively. In these works, Dr. Arthur discussed the many misunderstandings that arise in the standard presentation of electromagnetic theory, which uses Oliver Heaviside’s now-familiar notation for Maxwell’s equations based on Josiah Willard Gibbs’s vector notation. In particular, that work lamented the common depiction of magnetic quantities B and H, including their relationship with one another, their nature as fundamental or derived fields, and their importance for the theory of special relativity.
Biggest AI firms will likely recoil at Bernie Sanders' AI wealth fund.