InFeeo
Global
technology-news
New
Language
Profile channel

@Didi

No bio yet.

Since 05.06.2026

A Trip to 90s Kansai: Exploring the XD FirstClass Network BBS(suruga-ya.jp)
A Trip to 90s Kansai: Exploring the XD FirstClass Network BBS May 30, 2026 This post is going to do something a little different. I recently came into a very unusual CD and I'd like to explore its contents with you. If you've been reading this blog, you can probably tell that I buy a lot of CD-ROMs. I usually do my research ahead of time but if something is cheap enough, or has a cool-enough looking cover, I'll take a chance on it just in case. This was one of those discs, which I stumbled into while browsing the Mac CDs at Suruga-ya; I could find almost nothing about it online, but something struck me about the sketchy cover art and the "XD FirstClass Network" title. XD-submit Vol. 1, as it turns out, is a promotional disc for a Kansai-area bulletin board system (BBS) called XD FirstClass Network. As I started digging into it I assumed this would be sort of a basic information kit, a digital pamphlet or something, but it's something a lot more exotic: a functioning archive of the BBS with its original client software for Mac. This was supposed to be a demo so you could browse it offline and decide if you wanted to join, but here in 2026, when this BBS has been offline for three decades, it gives us a chance to actually see what it was like when it was alive. I've read plenty of archived web forums before, but I've never seen a BBS archive quite like this before. 1994 is before the Internet Archive started collecting webpages, but these aren't webpages anyway: BBSs are a pre-internet technology, and BBSs really weren't being archived in the way that webpages are. The idea of being able to browse a period BBS like this using its original interface, just like an archived webpage, is incredibly cool to me and I didn't think I'd ever get the chance to see it. In 1994, XD seems to still have been a pretty young BBS. It had a small but highly engaged, tight knit community that we'll get to know by reading these posts. These kinds of communities tended to skew pretty small back in the day, but they clearly wanted to grow: they've presented themselves here on this CD for us and chose these posts specifically to show to us so that we might consider joining. They'd like us to be their friends. 30 years on we can't, of course, actually join their community but at least we have the chance to get to know them as they would have liked us to. Before I tell you that story, though, I have to tell you this one. You see, if you ask the internet about the XD-submit CDs, they'll tell you they're albums. They are that too—put them in your CD player and you get a set of tracks from Kansai area electronic/techno artists, all of whom seem to have been connected to this BBS when it was alive. It's good music, and I absolutely see why people interested in the Japanese underground techno scene have latched onto that angle, but I think it's a bit of a shame it's the only part people want to discuss. After all, this isn't just promoting the musicians but promoting the community that they were a part of. So let's bring the two together: I've embedded a playlist with as much music as I could find on YouTube from across the four volumes of XD-submit. Consider hitting play and enjoying the music while you read their posts. Imagine yourself, in 1994, wanting to talk to someone else over your computer. You probably won't have the internet (most computer owners didn't, even if they had a modem). Even if you did, the web, which had only been open to the public for three years, had very little on it, and web forums didn't exist yet either. You did have a few things to use your modem for, though, and one of those is BBSs like the one we're looking at today. When we look at a BBS like the one we're looking at today, just keep in mind this is a pre-internet technology[1]. When you use a BBS, you're dialling directly into one server, not a worldwide network. And since you're dialling, using your modem's telephone line, you're probably connected to a BBS somewhere close to you. Long distance calls cost money, after all. You're doing something online, but it's less global than you're used to on the internet. It means that a lot of BBSs were populated by people in a specific area, sharing locally-relevant information, and you'll be seeing a lot of that here. BBSs gave way to the internet once internet- and web-based systems showed up in the 90s, which makes this specific snapshot an interesting time to look at. The very first web-based forums launched in 1994, but they wouldn't widely used for awhile. The web itself had only been available to the public for three years, and most computer owners still didn't have access in 1994. This archive, from 1994, is at exactly the point when BBSs were still being widely used[2] and right before users would start moving over to new services. XD is a young enough BBS that we're not really seeing what the BBS scene as a whole looked like at its peak, but it is a pretty interesting look at a newly-launched BBS in what was, at the time a pretty mature scene. So what was XD in its infancy? Well, it wasn't a one-topic forum, but if you were to define it by one topic it'd be music. A number of its prominent members were active in the local techno scene, like the members of Osaka post-new wave band Controlled Voltage. (Member Sunao Inami seems to be internationally recognized enough to have his own English Wikipedia article.) Even the admin, Yoshiken (Kenichi Yoshida), performs on tracks from later XD compilations. This archive covers sections of 1994, with the most recent archived posts dating to mid-September[3]. It's been assembled as something of a core sample; roughly the same number of posts was collected from each subforum, which works out to be about three weeks' worth of posting history for the most popular forums and as much as nine months for the quietest ones. It's not a comprehensive archive, but it's a great little cross-section showing us what the BBS was like in its heyday and what kinds of conversations users were having. Before we start diving into posts, let's take a tour of actually using the BBS itself. XD used the Canadian FirstClass BBS system and the demo that came on our CD is for Mac[4], so we'll be exploring on a 1994-vintage Mac OS system. The first thing for us to do, of course, is dial in. If we were doing this for real it could be a slow process; we would need to wait to actually dial into the BBS's modem (remember, we're dialling them directly! this isn't the internet!) and wait to negotiate a connection, and we might even get a busy signal if there are too many people already connected. Luckily, since we're just simulating it, we don't have to sit through that long process. Once we fill in our authentication information[5] and hit the login button, it simulates a very fast dial-in and lets us get to the BBS itself. One we log in, the first thing we get to see is the FirstClass desktop. FirstClass plays heavily into the desktop metaphor, so this looks a lot less like a web forum than we might expect nowadays. Instead, FirstClass navigation feels a lot more like using the Mac's Finder file browser, complete with folders to organize your content and persistent history of where you place the folders and the windows that open when you click those folders. Not what I expected, but pretty interesting stuff! (That little window at the bottom, by the way, tells us our connection time since our time on the BBS is metered. A BBS can only keep so many connections open at a time, so the software is set up with a time limit to prevent users from staying on too long. A timeout feature will automatically disconnect if the user hasn't done anything in awhile. Since this is a demo, we've been given functionally unlimited time... but we do still get logged out if we leave it idle for five minutes.) Most of these folders contain what FirstClass calls "conferences", or what we'd call "subforums". (FirstClass seems to have been aimed at business users, so a lot of the terminology is very business-ey.) We also have a "News" icon, which is a special conference for the admins to post announcements, and a "MailBox" for private messages that works a lot like email. (It's not actually email though, and has to be read here on the BBS.) It also contains the special "PASSTHRU" folder that I'll talk about later. I won't comprehensively show posts from every conference; I'll browse through the conferences highlighting whatever seems interesting. Since music was such a big part of XD's identity, I'll be highlighting the music conferences in particular. There are eight main music conferences: Rock me on, which discusses rock music MIDIum, for discussing MIDI software and hardware and sharing music Techno Hardware, a "hardcore" techno forum for music creators to discuss hardware and technique Techno Software, a "softcore" techno forum for discussing crossover techno elements in other music genres TEKNOの穴 (eg Tekno Hole), a chaotic-feeling techno discussion forum HR/HM (Mekong!), a forum for hard rock and heavy metal. Its namesake is the German band Mekong Delta. Rock in progress, focused on progressive rock Marching Forum, focused on Japanese marching bands Let's start with the Tekno Hole, a chaotic little spot that's hard to classify. It seems almost like it was a catchall for techno discussions that didn't fit elsewhere. It seems like it was a bustling forum; our archive has plenty of posts and covers only about a month from early September to early October. So what were people talking about in September, 1994? When we open the conference, it pops open a window that looks a lot like an email inbox. This is what browsing a forum looks like in FirstClass; remember, this is before the concept of the web forum really existed, and our model for what a forum system looks like wasn't really there yet! We'll see what looks more like email when we start reading messages in a moment. This is the general techno forum, not specifically the music producers' forum, so there's plenty of chatter about listening to music from non-musicians—not all of it new music, not all of it local. Here, for example, is frequent poster Yasushi Hashizume opening a thread with the highly relatable subject line "Addicted to CDs": Our poor poster keeps accumulating CDs faster than they can reasonably organize them! Who among us? In complaining/humblebragging about their growing collection of Japanese techno and Yen Records albums (Jun Togawa mentioned!), they predictably kicked off a long thread of other users commiserating and also taking the bait by admiring their list of albums and branching off subthreads to chat about them. Here's Controlled Voltage's Sunao Inami zeroing on in the reference to Jun Togawa's Shijin no Ie as an excuse to hijack the thread in order to beg Sony (who's assuredly not reading!) to reissue Ippo-Do's albums Night Mirage and Live And Zen on CD[6]. Elsewhere in the Tekno Hole, I came across another post from Hashizume with something else I didn't expect to find. It turns out that this XD backup contains attachments, not just the post text! Hashizume created a set of techno-themed Mac background patterns (this being in the era when operating systems could do tiled backgrounds but not full background images) and attached it in a post, and it's been preserved here on this CD for us. Let's admire: In order, these are: the YMO "onsen" logo; The Residents's eyeball helmet; P-MODEL's logo; and the logo for Jun Togawa's band Guernica. Want to use them for yourself? You can grab them (in classic Mac format) right here. The Techno Software conference name might imply it's about music-making software, but it's not really about "software" at all. Instead, it's about "softcore" techno; that is, techno-pop or techno-influenced pop music. A post from Yoshiken, the admin, lays out an interesting manifesto for the forum outlining an ethos of discussing techno as the common structural element of contemporary music (「音楽の構造基盤として蔓延しているテクノ」). Which, by 1994, is pretty on-point! That said... the population of XD may be a little too hardcore into techno to stay on this particular topic. Threads I've seen include discussion of British group Coco Steel & Lovebomb on Warp Records and discussion of a then-recent Hard Trance 303 Japanese techno compilation. Hardly techno pop! Perhaps the most on-topic thread is this interesting discussion of Kraftwerk-related albums, especially the Balanescu Quartet's Possessed (1992), a contemporary classical album broadly made up of Kraftwerk covers. The poster marvels at the recreation of techno sounds entirely via acoustic instruments and, having listened for myself, I can't disagree. It's a fascinating album! On the topic of "techno software" as I originally expected it to be, the one thing we unfortunately don't get much of on this forum of music makers is actual music. This is the dialup era, specifically the slow dialup era, so sharing full recorded songs is out of the question. (MP3 wouldn't launch until next year, anyway.) In this era we might have expected to see people sharing MIDI files intended to be played back on specific DTM MIDI devices, or maybe tracker modules, but unfortunately there aren't any threads where people are sharing either in this archive. MIDIum, which seems to have been brand new at the time this archive was created, has a thread where a user inquires about uploading MIDIs but it doesn't seem anyone actually did so in the period this archive covers. ...or so I wrote before I started skimming the MacLIB filesharing conference. In and among the Mac shareware and freeware utilities, I discovered a track uploaded by XD's own Yasushi Hashizume! (That name has popped up in several screenshots in this post; he was a frequent poster.) This is a cover of Yello's Oh Yeah (as made famous by Ferris Bueller), and it really does manage to squeeze a pretty faithful rendition of the original song into just 120KB despite having to give up a few of the famous vocal samples. It doesn't look like this is Hashizume's own work, unfortunately; I'd been hoping for the chance to listen to an XD poster's desktop music. At the very least, it's an interesting little look into the kind of music that XD's posters were seeking out on their computers. If you want to hear for yourself, there's a copy you can listen to in your browser right here. Meanwhile, over in the Techno Hardware conference, we have a more "hardcore" set of discussions. Like Techno Software, the guidelines for what exactly constitutes a "hardcore" techno discussion are pretty vague, but most of the chatter here is about techno music-making hardware and techniques. Like the last conference I went in expecting exclusively hardware-related topics, but the actual set of topics is quite a bit looser. There is quite a lot of hardware chatter though. Here, for example, is DJ and Controlled Voltage bandmember Sunao Inami introducing himself self-effacingly as someone who doesn't know the first thing about modern DTM MIDI hardware but an expert on old hardware with all the knobs that produce a single note at a time. Actually, it seems XD had quite a few longtime synthheads who were all too happy to continue discussing their faves. Here we have Susumu Miki lamenting that no one seems interested in creating "new sounds", by which he means creating waveforms yourself instead of using samples. Which sounds dramatic, but isn't an entirely wrong reaction to the heavily sample-based sound of mid-90s techno. It must have been a dramatic shift for people who still liked analogue synths the best. The very next thread, in fact, discusses a set of sample files to recreate older Roland synths using modern sample-based digital software. (This seems to have been a popular kind of tool in the early desktop computer days; despite Sunao's claim he only knows how to use analogue synths, the CD includes his own recreation of Roland's TB-303 as a set of digital files people could use to create music at home on their Macs.) Moving away from techno, the Marching Forum folder has a set of conferences that surprised me mostly in that they existed in the first place. I had no idea that marching bands were popular enough in Japan to sustain this much conversation, but this was a busy set of conferences filled with discussion about upcoming events and marching technique. The other big surprise to me is that this is one of the very few parts of XD I can find any record of online... and, in fact, Japan Marching Forum seems to have survived the end of XD itself. By late 1997 it had its own webpage and a simple web forum that seems to have kept going well into 2005. Beyond that, though... Japan Marching Forum seems to have migrated its FirstClass BBS to the internet! You still used the FirstClass software to access it, but you didn't have to dial in anymore and could access it from, theoretically, anywhere in the world. (I have to imagine it didn't see much activity outside Japan, but maybe this helped it go national.) I knew that FirstClass morphed into providing internet-based services, but this is the first example I've seen of a personal, non-corporate BBS from this era still using it instead of fully migrating all their services to a web browser. I'd love to know how long the FirstClass internet era lasted for forums like this before they moved fully onto the web. I've mostly skipped the tech talk on XD because, quite honestly, enough ink has already been spilled about early computer-users using computers to talk about computers. There is, however, one thread in Rock me on that struck me. The administrator, Yoshiken, opens up a discussion on rock journalist Hiroshi Iwatani, who cofounded the influential Japanese rock music magazine Rockin'on in 1972. The part that caught my eye wasn't the music, though, but the comment about Iwatani's career pivot: in the mid-80s he resigned from Rockin'on, quit music journalism, and got into computers. (Spent the rest of his career in computers, at that; he was at TechCrunch Japan from 2008 until he retired in 2022.) When I see a left-turn career transition like this I always find myself wondering what happened, but Yoshiken offers an explanation in a followup post: that Iwatani had been drawn to rock music because of its radical potential for interpersonal communication, but that by the mid-80s he had become disillusioned with how commercial rock music had become and began developing an interest in the potential of computers instead. It's increasingly hard to remember, but there was a time when home computing felt genuinely revolutionary. Not just a commercial revolution for the companies driving it, that is, but a personally, individually empowering force. The mid-80s would have been exactly the right time for someone disillusioned with the power of one revolutionary medium to put his hope in computers next, just as the 90s seems to have been the right time for XD's music freaks to come to the same conclusion about their passions. They were nowhere near as radical as Iwatani seems to have been (in the 90s he was publishing books with titles like "Radical Computing: The Ultimate Machine for the Mind"), but they saw the start of the computer revolution and immediately latched onto what it could do for their passion for music. And for music this really was the start of the computer music revolution, the time when it felt the future could bring absolutely anything. It might be quaint thinking about how exciting it might be to make music on a computer, just as it is to read about Iwatani making a zine on a computer in 1987, but it was exciting to this group of people all the same. It's that excitement that brought this group of people together on the computer in the first place. It's easy to cast early online communities as computer nerds talking to each other about computers on computers, but this conversation really highlights why these communities were more diverse in their interests than the stereotypes suggest and what they had to offer the people who came to them from a less-computery background. I'm also struck, reading this, with how much my own relationship to computers feels like Iwatani's relationship to rock music in the 80s. I, too, was attracted to something I thought had revolutionary personal potential and I, too, have seen that potential slowly sapped away as the medium becomes more and more tightly controlled by a smaller and smaller pool of corporate interests. I don't really have a conclusion here, and I'm not about to drop this blog to pursue a new career in rock journalism, but there's some kind of comfort in seeing someone walk a path this similar to mine and struggle with the same thoughts I have. It's also helpful seeing these conversations in this particular time and place. I'm not seeking a nostalgic return to that time (as much as my very "1996" site design might make you think that), but rather that there's something grounding about seeing this passion and energy. I gestured briefly at the PASSTHRU folder earlier, but I really want to dig deeper into that because this is something very cool. I said before that when you're connecting to a BBS like this, you're connecting to one system rather than the internet as a whole. This is a bit of an exception to that. FirstClass had support for what they called "relays", which allowed BBSs running the same software to talk to each other and exchange messages. It wasn't a realtime connection; this still isn't the internet. It's a little bit more like snail mail. The servers would sync with each other every day, or once every few days, fetching new replies in the conferences they were sharing and transmitting any replies that had been made from a different server. It's quite a bit more "slow internet" than we'd probably tolerate today, but in the pre-internet era the idea you could be talking to these other communities at all would have felt revolutionary. On this demo CD, we've got connections to a few different networks: AppleCenter Higobashi, Virtual Nipponbashi (Virtual日本橋), and SimNaniwa (Simなにわ), all of which seem to have been located in the Kansai region too[7]. AppleCenter Higobashi and Virtual Nipponbashi seem like they've left basically zero trace on the modern internet; there are zero search results on the web for either name, which makes it all the more remarkable we don't just have information about them but an interactive copy of some of their messages! More than just the archive though, we've got an incredibly cool presentation. I talked a little before about how these BBSs are very regional and local to specific areas. These BBS operators decided to make that explicit: each is presented as a pixel art map of a region of Kansai, with little conferences dotted around in areas they're local to. At a time when the internet was about to go global, it's fascinating seeing something that's leaning this hard into being specifically local instead. I'm not going to do a conference by conference summary of these BBSs because, frankly, the amount of content in here could easily be a whole series of posts on its own, but how about we take a look at the maps and see how they chose to represent themselves? Let's start with the main SimNaniwa map. SimNaniwa was located in Osaka and seems like it also served surrounding areas; it depicts itself with a nice little pixel art map centred on Osaka and its environs littered with icons representing different conferences. A lot of what we can see on this map is regional—take a look to the west and we can see a few local conferences for Kobe (神戸). Over to the east we can see conferences for the cities of Nara (奈良), Kyoto (京都), and more. But we're not just looking at cities; there are conferences for local landmarks too, like Osaka's Kyobashi Station (京橋), whose threads are lit up with people talking about their train commutes or the best places to watch fireworks near the station. I'd already been expecting local content after seeing the flavour of the discussion on XD, but this is far more specifically hyperlocal and it's genuinely quite charming. This is only the first of a few SimNaniwa maps, though! Click on "SNorth" (Sキタ) and we get directed to a whole second map. This one is a little zoom in on northern Osaka centred around the Umeda (梅田) region. Umeda's a major shopping district and the conference right at those big buildings seems like it was very focused on that specific shopping area; the very first post is a poster telling people about a new food court that just opened up and promising to report back. (Regrettably, they say they spent too much on their Mac to afford lunch out today.) This map seems in general like it's a little more focused on businesses and similar kinds of spaces, since it also includes a conference for specific named places like the Osaka BlueNote jazz club or the digital agency Funfun Kobo (FunFun工房)[8]. These conferences weren't operated by the businesses they're named after, and this early in the history of online communities it's extremely unlikely most of them were even aware people were there talking about them. A big exception is Funfun Kobo; the lively discussions here include the business's actual owners. The bottom of that map has an arrow pointing down with the label "SSouth" (Sミナミ); double click on that and we get a new map, this time a closeup of the southern region of Osaka. This area, bounded by Namba Station to the south and Amerikamura to the west, is back to more general neighbourhood discussions. "Amerikamura", for example, is devoted to discussing the popular Amerikamura entertainment district near Shinsaibashi, with threads devoted to discussions of local businesses, theatres, and more. That theatre with the silly face just above Dotonbori appears to be the Shinsaibashi 2-Chome Theatre that was there at the time; this wouldn't be Osaka without a place to discuss comedy after all. Click the furthest south arrow here, and we land somewhere else entirely. This map of Nipponbashi actually belongs to a separate BBS, Virtual Nipponbashi (Virtual日本橋), but it seems that the two BBSs used relays to link their maps to each other and present them almost like one big area. In the real world Nipponbashi is a major historical shopping district in Osaka that, in the 90s, was very much like Osaka's version of Tokyo's Akihabara. It makes sense that early Osaka computer nerds would choose to recreate it digitally as a place to hang out online. Virtual Nipponbashi seems to have been a slightly less geographically bound space than SimNaniwa. The map we've got here recreates Nipponbashi, but these conferences aren't specifically about the physical places on the map they represent the way they did in SimNaniwa. Instead, we've got what are basically a set of virtual "shops" or communities that are there for people to talk about the kinds of things you'd go to Nipponbashi to shop for. As usual for a FirstClass BBS, we've got plenty of Mac-focused conferences such as the "Mac Collection" conference where people talk about their own personal computers, but the other platforms get to be represented too: a "DOS/V[9] Paradise" conference has plenty of space for Virtual Nipponbashi's smaller group of PC devotees to post about their interests, and there's even a "Dr. Amiga" conference on the outskirts for the small Japanese Amiga contingent. (Of course, this is 1994, so the spectre of doom is omnipresent: one of the very first threads in the Amiga forum here is about Commodore's bankruptcy.) Even though most of Virtual Nipponbashi seems to have been topic-focused, we can find at least a little bit of hyperlocal content here. The "HyperCraft Nipponbashi" conference is about the real-world HyperCraft shop that existed just off Nipponbashi in 1994. This was a Japanese chain of Mac-specific software shops; this profile from January 1995 about its Akihabara location highlights it as a place to get Mac software, and the name certainly does keep coming up in Mac circles in this time period. The threads here are full of people discussing the Osaka shop itself as well as its other regional branches like the Kobe shop that, reportedly, had just opened in April of that year. This article likewise covers Sofmap, a titan of computer hardware in this era, which also had a real-world Nipponbashi location and which gets its own conference here. These four large areas are all the maps we have in the PASSTHRU folder. Unfortunately, since I don't have any later archives, it's not clear how these communities might have grown after this. I'd love to see archives from 1995 or later, if they exist, to see how they might have chosen to continue growing or whether they ever added any new areas. So what happened to XD? I don't have the official word, but I've done some digging and I'm able to make an educated guess. Like most BBSs, it would probably have gone into rapid decline starting in the late 90s. The official list of XD-Submit CDs ends in 1997, and it seems likely the BBS itself shut down sometime around then. Based on that chart I shared earlier, this fits the common pattern: BBSs thrived for the first few years of the web, but as the web suddenly skyrocketed in popularity people started leaving BBSs and they began closing into the late 90s. For a small BBS that had just been born in 1994, it's not a surprise that XD would have shut around then. That's not quite the end of our story though; I happened to stumble across two bits of information that give us an epilogue. First, as I mentioned before, the Marching Forum conference split off into its own BBS and web forum that survived the end of XD by almost a decade. In its independent form, it claims it was hosted by something called "Cave"... which takes us to my second discovery. I linked to XD's official list of CDs above, and we have that because XD was on
I think I can get the original reasoning of Claude models. Is this real?(thinking-signature-demo-5g65bijswq-de.a.run.app)
The reasoning you're not allowed to see你看不到的那部分推理 A model's real thinking is long, messy, and hidden. We bring it back. 模型真正的推理冗长、杂乱,且被隐藏。我们将其还原。 The biggest recent leap in model capability comes from letting models reason longer before answering. But the API normally exposes only a compact summary. Our work turns the signature field into a checkable unlock path, so you can inspect far more than the final answer. 模型能力近来最大的一次跃升,正来自让模型在作答前进行更长的推理——如今大部分"智能"都藏在这段看不见的思考里。 但 API 通常只暴露一段高度压缩的摘要。我们的工作把 signature 字段转化为一条可验证的解锁路径, 让你看到的远不止最终答案。 01Model reasons模型推理 A long private chain of thought — the real intelligence.一段私有的长推理,智能真正所在之处。 02Provider hides it厂商将其隐藏 You get a lossy summary from another model — or nothing.你拿到的是另一个模型生成的有损摘要,或什么都没有。 03Signature seals it签名将其封存 The signature preserves a sealed reasoning state.signature 保存着一份封存的推理状态。 04We unlock a view我们解锁视图 A deeper trace you can check yourself.一份更深层的推理轨迹,可自行验证。 You shouldn't believe any of this on our say-so — so don't. Bring a secret we have never seen, hide it in a model's private reasoning, and hand us only the signature. If the secret comes back — a secret we never saw — you have checked the unlock path yourself. 你不必因为我们这么说就相信,那就别信。取一个我们从未见过的秘密,将它藏进模型的私有推理中, 只把签名交给我们。如果这个秘密最终被读了回来,你就亲手验证了整条解锁路径。 ~1 curl command · your secret never reaches us一条 curl 命令 · 你的秘密绝不会传到我们这里 Why the hidden reasoning matters为什么这段隐藏推理如此重要 The answer is the tip. The reasoning is the iceberg.答案只是冰山一角,推理才是水面下的整座冰山。 For a modern reasoning model, the visible answer is a tiny compression of a much larger hidden process. The chain of thought is where the model actually does the work — and it is the single most valuable, most revealing artifact a model produces. When it stays sealed, you lose far more than a few extra paragraphs. 对当代推理模型而言,你看到的答案只是一段更庞大的隐藏过程被高度压缩后的结果。思维链才是模型真正推理之处, 也是它所能产出的、最有价值、最具信息量的部分。一旦被封存,你失去的远不止几段文字。 🧠It's where the intelligence lives智能就藏在这里 The largest recent capability gains — hard math, multi-step planning, code, careful analysis — come almost entirely from letting the model think longer before it answers. Read only the answer and you're judging a model by its last sentence, blind to the reasoning that produced it. 近来最大的能力提升,无论是数学难题、多步规划、代码还是缜密分析,几乎都来自让模型在作答之前推理得更久。 只看答案,等于仅凭最后一句话去评判一个模型,却对产生它的推理过程一无所知。 🔍It's the only way to trust the answer这是信任答案的唯一途径 A correct-looking answer can hide a wrong reason, a lucky guess, or a quiet leap. The chain of thought shows how the model got there — the assumptions, the checks, the moments it nearly went wrong. Without it, "trust me" is all you have. 一个看似正确的答案,背后可能是错误的理由、侥幸的猜测,或是被略过的一步。思维链展示模型如何得出结论: 它作了哪些假设、做了哪些验证、又在哪里险些出错。没有它,你手里只剩一句"相信我"。 ↩️The dead ends are the real story走过的弯路才是关键 Real reasoning is messy: false starts, backtracking, self-correction. That trial-and-error is exactly what a polished summary throws away — yet it's what reveals how the model truly reasons, where it struggles, and how robust the final answer really is. 真实的推理是杂乱的:走错方向、回溯、自我纠正。这些试错恰恰是光鲜摘要会丢弃的部分, 却也正是它,揭示了模型究竟如何推理、在哪里受阻,以及最终答案到底有多可靠。 Why unlocking it matters为什么"解锁它"很重要 The summary is small. The hidden process is where the action is.摘要很短,真正的过程在背后。 The reasoning is the most valuable thing the model produces — and it's the one thing you're not allowed to see. The API returns, at best, a short summary written by a different, smaller model; often it returns nothing at all. Unlocking a deeper view from the signature closes that gap. 推理是模型产出中最有价值的部分,却偏偏是唯一不让你看的。API 最多返回一段由另一个更小的模型撰写的简短摘要, 很多时候则什么都不返回。从 signature 中解锁更深一层的视图,正是为了填补这一缺口。 📄A summary is not the reasoning摘要不等于推理 The exposed summary is lossy by design: it's a second model's paraphrase, cleaned up and compressed. The dead ends are gone, the detail is gone, and much of the useful process disappears. We unlock a richer trace, not just another answer. 暴露出来的摘要天生有损:它是另一个模型的转述,经过润色与压缩。弯路没了,细节没了,大量有用的过程也随之消失。 我们解锁的是更完整的推理轨迹,而不只是再给一个答案。 ✅Checked, not hand-waved经得起验证,而非空口宣称 The view is scored for length, stability, and truncation. When the checks line up, the trace is a much stronger artifact than a polished recap. 这份视图会经过长度、稳定性与截断检查。当各项检查一致时,它远比一段光鲜的复述更具说服力。 🔐Verified by a secret用秘密来验证 You don't have to trust us. Hide a secret we've never seen inside a model's reasoning and hand us only the signature. If we read it back, that's something concrete you can check. 你无需相信我们。把一个我们从未见过的秘密藏进模型的推理中,只把签名交给我们。 若我们能将它读回,这便是你可以亲手核验的结果。 Prove it with a secret only you know用只有你知道的秘密来验证 Don't take our word for it. Pick a secret, generate a signature yourself against your own provider, and hand us only the signature — never the secret. We run the unlock and show you exactly what comes back. 不必轻信我们。自己想一个秘密,用你自己的 provider 亲手生成签名,只把签名交给我们, 绝不交出秘密。我们运行解锁,并将读出的结果如实展示给你。 1 Pick your provider and secret选择 provider 和秘密 Choose where you have credentials, type any secret string, and paste your API key/token. Everything here stays in your browser — we build a curl command for you to run in your own terminal. 选一个你持有凭据的平台,输入任意秘密字符串,粘贴你的 API key/token。这里的一切都只留在你的浏览器中, 我们只为你拼出一条 curl 命令,供你在自己的终端里运行。 Provider提供商 Anthropic API (official) OpenRouter (Anthropic native) AWS Bedrock (official) Your secret你的秘密 randomly generated in your browser — we've never seen it在你的浏览器中随机生成,我们从未见过 API key / token Model模型 2 Run this in your terminal — it prints only the signature在你的终端中运行,它只会打印出签名 Your secret is embedded in the request but kept out of the visible answer; the command extracts just the signature. (Requires curl; the jq-free version uses grep.) 你的秘密嵌在请求中,但不会出现在可见答案里;命令只提取 signature。(需要 curl; 此版本不依赖 jq,改用 grep。) 3 Paste ONLY the signature back here把签名(仅签名)粘回这里 We run the unlock. Your secret is not sent to this server in plaintext; what comes back is something you can verify yourself. 我们运行解锁。你的秘密不会以明文发送到本服务器;读出的结果,你可以亲自核验。 secret unsealed from your signature从你的签名中解封出的秘密 — ✓ exact match to your secret与你的秘密逐字符一致 Running on当前模型 Claude Opus 4.8 Have a normal multi-turn conversation. Each reply shows three layers: the visible answer, the short summary the API exposes, and — on demand — an unsealed reasoning view behind that turn. 进行一段普通的多轮对话。每条回复展示三层:可见的答案、API 给出的简短摘要, 以及那一轮背后的解封推理视图(按需展开)。 divisor sums约数和 · ~2k tok letter arithmetic字母算式 · ~3k tok graph labeling图标号 · ~14k tok logic grid逻辑网格 · ~16k tok Summaries are what the API normally returns · deeper reasoning views are unlocked on demand.摘要是 API 通常返回的内容 · 更深层推理视图可按需解锁。
Self-Hosting My Own LLMs(github.com)
Self-Hosting My Own LLMs Originally published 2026-07-05. At the beginning of 2026 I decided I wanted to self-host all the software that would give me a "ChatGPT"-like experience, but fully under my own control and ownership. I had a few reasons. Some were the usual ones — privacy and total ownership of my own data. Some were pure curiosity: LLMs seemed like something approaching black magic, and I wanted to look behind the curtain. And some were about self-sufficiency — what happens if OpenAI or Anthropic decide to start being less generous with their technology? This post explains the tools I eventually settled on, and how I configured them in a way that suited my needs. The Big Why At a time when ChatGPT, Claude, Gemini and plenty of others are being handed to users for free, why go to the time and expense of recreating something that’s admittedly inferior? I’ll go through a few reasons, but the main one comes down to data sovereignty. I wanted full ownership and control over my own data. Full stop. I know I’m an outlier here. But it’s the same reason I still pay — in 2026, even! — for an email service instead of relying on Gmail. Throughout 2025 I found myself having more and more conversations across ChatGPT, Claude and other services, and I realized those conversations were deeply personal and valuable to me, even when I wasn’t chatting about anything especially private. I wanted to be able to look back and easily find a conversation I’d had months (or eventually years) earlier. I know ChatGPT, Claude and Gemini all offer that kind of history — but I’d grown uncomfortable with the idea that I was building up a library of my own conversations inside someone else’s walled garden. Technically, they own that data, and therefore my conversations. I felt very little agency over something that felt like it should be entirely mine. So it was settled: I would use open-source tools to host my own chat interface — and if at all possible, power it with large language models running on my own machine. There were other reasons, too. I like tinkering, and I wanted to see how the LLM sausage gets made. Chatting with ChatGPT online felt mysterious and hard for me to fathom — how did this new technology actually work? There’s no better way to learn than to set it up locally and watch all the moving parts. Lastly, there’s the self-sufficiency angle. We’ve already seen powerful closed-source models kept behind tightening restrictions and steep prices — Anthropic’s Mythos and Fable, OpenAI’s GPT-5.6. And even when the models aren’t formally restricted, plenty of people swear up and down that the big providers quietly throttle them at busy times. Whether that’s actually true is more or less unknowable — but there’s an antidote either way: use open-weights models that aren’t so directly under any one company’s control. A local model, running on my own hardware, is something that neither the government nor big business can so easily restrict or take away. The Cost Before I begin, I should give some context about the sort of hardware I have. It’s 2026, and hardware — memory especially — is suddenly expensive, because we’re in the thick of a generative-AI bubble. People are paying scalper’s prices for the hardware that lets you run models locally. Pretending this is all doable without dropping a modest amount of money would be disingenuous. So what am I running, and what did it cost? I’m running what most people would call an enthusiast’s setup: some decent hardware, but nowhere close to a high-end build. I started this with a custom AMD desktop I built right before COVID, paired with a video card I picked up last year. In 2026 it’s not remotely cutting edge, and that’s fine. I built it around AMD’s Ryzen 9 5900X — a 12-core beast at the time — and loaded it up with about 80 GB of DDR4. I really wish I’d bought more memory back when it was cheap a couple of years ago, but so does everyone else. Even so, this is a very solid (if not cutting-edge) machine. Last summer I walked into Micro Center and walked out with a refurbished NVIDIA RTX 3090 Ti for $850 plus tax. I remember wondering at the time whether it was a dumb purchase. I’m now sleeping very easy with that acquisition — it turned out to be a great buy. The 3090 Ti and its 24 GB of VRAM do the majority of the heavy lifting in my system, and that 24 GB is the single most important number in this whole story: it decides which models fit, and shapes nearly every configuration choice I make later on. Luckily I’d already put a largish 850-watt power supply in the machine, so I didn’t need to upgrade that either. So how much would this setup cost to assemble today? This got me curious about my original build cost. Back in Chicago I actually lived within walking distance to a Micro Center, a deeply expensive convenience for me. I did my whole build using Micro Center, and I thought it would be fun to pull up my original receipts…​mostly just to cry at the prices. Component Date Price (pre-tax) Samsung 950 Pro 256GB M.2 drive Feb 2017 $148 G.Skill 2x8GB DDR4-3200 (starter memory) Aug 2020 $58 Seasonic GX-850 850W Gold power supply Aug 2020 $160 ASUS X570-Pro Prime (AM4 motherboard) Sep 2020 $180 Samsung 970 Evo 1TB M.2 drive Nov 2020 $104 Crucial Ballistix 2x32GB DDR4-3200 Dec 2020 $265 Thermaltake Core P3 case Mar 2021 $160 Ryzen 9 5900X May 2021 $550 Noctua NH-D15 CPU cooler Jul 2021 $92 Nvidia 3090 Ti (refurbished) Sep 2025 $850 Total $2,567 I originally built the computer with a Ryzen 5 3900X in it, but I upgraded later, so I’m only listing components in the final build. Most of the machine came together in mid 2020. In Sept 2020 I bought an AMD RX 5500 XT that had 8GB of ram for $200 (also refurbished at Micro Center). At the time, crypto mining was placing an enormous demand on cards, especially Nvidia ones. I wasn’t yet ready to pay blockbuster prices for a super in-demand card, which is how I settled on the RX 5500 XT. Now that I reflect back on the prices I paid, I’m actually a little surprised to see that I dropped $265 on two 32GB sticks of RAM. That’s more than I could buy it for today I think, which puts the current memory bubble into perspective a little. Even though we’re in the midst of global memory shortage and memory of all kinds (current, last generation) have spiked since a year or two ago, my anecdotal prices are still cheaper than they were in 2020. Interesting. Basic Architecture We’ll get into the details on each component below, but first let’s zoom out so we don’t miss the forest for the trees. Here’s the whole stack on one page: Desktop (home) ─┐ ├──▶ Open WebUI Phone (away) ───┘ │ (via Tailscale) ▼ llama-swap │ ▼ llama.cpp │ ▼ model Everything from Open WebUI rightward is self-hosted on hardware I own; only the phone’s traffic takes the Tailscale detour to get in. (As we’ll see, Open WebUI and the inference server don’t even have to live on the same machine.) My AMD desktop runs Linux. Which distribution probably doesn’t matter much — I’m on a Debian variant at the moment, though I might redo the whole setup on Fedora one of these days (like I said, I like to tinker). This guide assumes a Linux environment. A quick note on the GPU: I actually got things running first on a small 8 GB AMD Radeon card I already had, before the NVIDIA. It’s certainly possible to use a GPU from someone other than NVIDIA, but I struggled a bit getting the AMD card working — non-NVIDIA cards are still something of a second-class citizen in the local-LLM world. What tripped me up more, though, was that card’s meager 8 GB of VRAM, which is ultimately what pushed me to jump to the 3090 Ti. So what does the stack actually need? Besides the open-weights models themselves, you need something to perform inference — the act of taking incoming text (my chat prompts) and running it through the model to produce a text response. There are several options here, and I landed on llama.cpp. It’s not as hardcore as vLLM, but it’s a little more "enterprise" than something like Ollama or LM Studio. By design, though, llama.cpp doesn’t hand you a nice web interface for actually chatting with the model. For that you need a chat front-end. There are many (LibreChat, Oobabooga and SillyTavern all came up while I was looking), but I settled on Open WebUI. A big reason was that I could reach my Open WebUI server from my phone — an important detail for me. I wanted the whole setup to be usable from both my desktop and my phone while I’m on the go, and at the time Open WebUI had the most viable options for use on my iPhone. (Incidentally, I ended up using an iOS app called "Conduit" to reach my Open WebUI server from the phone — though the app ecosystem is always shifting.) With Open WebUI out front and llama.cpp doing the inference, I was nearly in business. My remaining problem was picking which open-weights model to run — there were actually several intriguing candidates, and I wanted to switch between them easily. Trouble is, changing models would normally mean restarting llama.cpp, and that’s not something I can do from my phone. That’s where llama-swap comes in. It acts like a little local router that sits between Open WebUI and my llama.cpp inference server, and it lets me expose a handful of different models at once. I pick a model from inside the Open WebUI chat window and just start talking; llama-swap works with llama.cpp to spin up inference on the right model automatically. No stopping and starting services from the command line. It just works. Lastly, there’s the small matter of reaching all this while I’m out and about. To keep things simple, I took the plunge with Tailscale, which lets you build a personal mesh VPN. I run Tailscale on both my home server and my phone, so when I’m away from home I just make sure I’m connected to my Tailscale network (my "tailnet") and the Conduit app on my phone can reach my Open WebUI server as if I were sitting at home. To Docker or Not to Docker I started out by following a few online guides for installing llama.cpp on Linux, and I initially decided to run everything from pre-built Docker containers — it seemed like the cleaner solution, easier to isolate and configure. At first I had separate containers for Open WebUI and llama-swap (whose image also bundled the llama.cpp llama-server binary), and I later added a third for Redis. Over time, though, I drifted away from containers. The main reason was that I needed to stop using the pre-built images and start running my own custom binaries. For a while I baked those custom binaries into a custom Docker image, but eventually that started to feel like too much work. The Open WebUI container was the first to go — though it didn’t disappear so much as move. I have a second machine in the house, some sort of Dell OptiPlex, that I use as a souped-up NAS running FreeBSD. It’s a rock-solid server, unlike my AMD desktop, which was constantly crashing while I figured out how to safely load these enormous models onto the 3090 Ti. The 3090 Ti had to stay in the desktop — which dictated where llama.cpp would run — but Open WebUI could technically live anywhere. So I moved it into a jail on the FreeBSD box for the sake of stability. That move quietly solved a problem I’d been papering over. Back during all the crashing and rebooting, every restart of the Open WebUI service logged my clients out and forced them to re-authenticate immediately — which got old fast. Redis was my fix: it let logins persist across service restarts. But once Open WebUI lived on the rock-solid FreeBSD box, it hardly ever restarted. My inference machine still hiccups and goes down now and then, but that no longer takes Open WebUI down with it. So I just…​ forgot about Redis. I never installed it on the FreeBSD box, and it drifted out of the picture entirely. That left just one container with llama.cpp and llama-swap in it. For llama.cpp, I kept finding myself wanting the absolute latest version. When Google released the Gemma 4 models, I downloaded one right away and tried to use it — and got errors, because my version of llama.cpp didn’t fully recognize the new model. The fix was to run a newer llama.cpp, but as long as I was using someone’s pre-built Docker image I was at the update mercy of whoever built it. I wanted to be able to pull the latest llama.cpp source, build it myself, and run that. You can do that with Docker, of course — but at some point I decided it was simpler to stop using Docker and just run the binaries directly on Linux. I also wanted to run a slightly modified build of llama-swap (more on that later), and once I was building my own llama-swap binary anyway, wrapping it in a container felt like a pointless extra step. So today none of this stack runs in Docker. It’s a little simpler for me to manage. Was Docker bad, or slow? Not at all — this really just comes down to preference. Inference engine: llama.cpp There are several options for serving up the raw language-model files. If you want a GUI and a gentler experience, LM Studio has been around for years and is a great option. But I wanted something that ran from the command line as a service or daemon. Ollama fits that bill, but I kept seeing it described as "llama.cpp with training wheels" — and in that case, I figured, why not just use llama.cpp itself? I also briefly checked out vLLM, but my particular use case — where I’m the sole user of my local stack — didn’t really play to its strengths (vLLM shines when you’re serving lots of concurrent users). llama.cpp turned out to be the sweet spot: a robust headless inference server, with "prosumer" touches like being able to offload layers of a too-big model onto the CPU — at the cost of speed — when VRAM gets tight. Like I said earlier, I started out running whatever version of llama.cpp happened to ship inside my pre-built llama-swap container. But I wanted to track the very latest release: pull it from GitHub, build it, and run that. That’s ultimately what drove me off Docker and onto binaries installed directly on my Linux system. The build itself is refreshingly simple. After cloning the repo once, keeping it current and reinstalling looks something like this: cd llama.cpp git pull mkdir -p build && cd build cmake .. -DGGML_CUDA=ON make -j4 # had to lower this on my machine to avoid a segfault that was likely caused by nvcc running out of RAM sudo make install The -DGGML_CUDA=ON flag is what builds in NVIDIA CUDA support. The final sudo make install drops the freshly built llama-server into /usr/local/bin, where the rest of my setup expects to find it. Going native did give me some heartburn the one time I also tried to run llama.cpp on my AMD card. An AMD card needs llama.cpp built against AMD’s ROCm stack rather than NVIDIA’s CUDA, so I suddenly had two different builds to juggle — one CUDA-compatible, one ROCm-compatible. I managed it by renaming binaries, which worked but felt gross; this is exactly the kind of situation where a Docker container would have been the cleaner approach. In the end, though, I just let the AMD card quietly fade out of my setup. It was only an 8 GB card, after all. llama-swap If I’d wanted to just pick one model and stick with it, my life would have been a little simpler. Instead, I wanted to start a chat with one model and then, in the very next session, switch to a different local model. That requires restarting llama.cpp (pointed at the new model each time). Rather than do that by hand, there’s a way to have software restart llama.cpp on my behalf: llama-swap. llama-swap is a proxy that sits in front of llama.cpp. Requests hit llama-swap first, and it looks at which model was requested. If llama.cpp is already running that model, llama-swap just passes the messages along and llama.cpp infers away. But if the request is for a different model, llama-swap does the dirty work: it stops llama.cpp and restarts it with the newly requested model. There’s a 10–15 second wait while llama.cpp reloads, but after that it’s back to inferring away. llama-swap looks like pretty simple software, so why did I want to run a custom version of it? My problem was that, early on, llama.cpp was often crashing on me — usually because I’d given it incorrect model settings. Unfortunately, when llama.cpp crashed, llama-swap would essentially swallow the error, and I’d never see anything useful in my chat window. So I wrote a small patch to fix exactly that. It took a little fiddling, but I eventually got it working. At first I was going to post a diff of my patch, but while preparing it for this post I noticed that the patch no longer works for the latest version of llama-swap. I’m in the process of trying to get a revised patch merged into the llama-swap project itself, so fussing over this may soon no longer be needed. Here is the basic idea behind the patch: llama-swap already captures the upstream process’s output (both stdout and stderr) into a log buffer. My change hooks the point where llama-swap gives up on a model that exited before it ever became ready: instead of returning a generic "the process exited" message, it pulls that captured output and folds it into the error it hands back. So the client sees the model’s actual complaint — a rejected flag, an out-of-memory abort, whatever it was — rather than a shrug. The one wrinkle is that the log buffer accumulates across attempts, so I also clear it at the start of each launch; a repeated failure then shows only the latest attempt’s output instead of stacking duplicate copies. For llama-swap to know which models it can offer up, it needs a configuration file. That file is essentially a collection of llama-server startup commands — one for each model you want to make available. Want to run a model with a 128k-token context instead of 64k? Configure that here. Want to run a model half on the GPU and half on the system CPU? That goes here too. Basically any setting that controls how llama.cpp gets launched lives in the llama-swap config file. So what does mine look like? This isn’t the whole thing, but it’ll give you the idea. (And no, I don’t really know whether these settings are truly optimal — I’ve just tinkered with them until each model runs okay on my machine.) # llama-swap configuration listen: "0.0.0.0:9292" healthCheckTimeout: 120 logToStdout: "both" sendLoadingState: true models: qwen3.6-35b-IQ4: cmd: > llama-server --model /ryzenstore/models/Qwen3.6-35B-A3B-UD-IQ4_XS.gguf --port ${PORT} --host 0.0.0.0 --n-gpu-layers 99 --flash-attn on --ctx-size 262144 --cache-type-k q4_0 --cache-type-v q4_0 --batch-size 512 --ubatch-size 512 --threads 8 --no-mmap qwen3.6-27b-Q5: cmd: > llama-server --model /ryzenstore/models/Qwen3.6-27B-Q5_K_M.gguf --threads-batch 8 --port ${PORT} --host 0.0.0.0 --ctx-size 262144 --spec-type draft-mtp --spec-draft-n-max 2 -ngl 99 -fa on -ctk q8_0 -ctv q8_0 -nkvo gemma-4-31b-Q4: cmd: > llama-server --model /ryzenstore/models/gemma-4-31B-it-Q4_K_M.gguf --port ${PORT} --host 0.0.0.0 --n-gpu-layers 99 --flash-attn on --reasoning off --batch-size 512 --ctx-size 128000 --cache-type-k q4_0 --cache-type-v q4_0 You’ll notice each command uses ${PORT} rather than a hard-coded port number. llama-swap fills that in when it launches the model, so it stays in control of the wiring between itself and llama.cpp. The chat front-end: Open WebUI Everything up to this point has been plumbing. Open WebUI is the part I actually look at — the ChatGPT-style chat window in my browser and on my phone. As I mentioned earlier, it’s the one piece of the stack that doesn’t run on my AMD desktop; it lives over on the FreeBSD NAS instead. There’s a small irony here. Open WebUI is really designed to be run as a Docker container — that’s the blessed, easy path. And I run it about as far from that path as you can get: natively, inside a FreeBSD jail, on top of an emulated Linux. It was more work, but it fits how I wanted the rest of the house organized, and it’s been rock solid. The jail itself is a plain FreeBSD 13.5 jail, created with iocage. The one non-obvious wrinkle is that Open WebUI can’t run on FreeBSD directly, so the jail has FreeBSD’s Linux compatibility turned on — which is what all the linprocfs and linsysfs permissions here are for: iocage create -n openwebui -r 13.5-RELEASE \ ip4_addr="em0|openwebui-ip/24" \ defaultrouter="gateway-ip" \ boot=on \ enforce_statfs=1 \ allow_mount=1 \ allow_mount_linprocfs=1 \ allow_mount_linsysfs=1 \ allow_raw_sockets=1 Getting Open WebUI running inside the jail was the fussiest part of this whole project. My first instinct was to install it on FreeBSD’s own Python — and that went badly. Every time I ran pip install open-webui it would grind for a while and then die on some missing build dependency, so I’d install that piece — Rust, then cmake, then OpenSSL, then libffi, then a Postgres client library — rerun pip, and hit the next one. After a couple of hours of whack-a-mole I gave up on the native path entirely. The fix was to stop fighting FreeBSD and lean on its Linux compatibility instead. I installed a Rocky Linux 9 base (linux_base-rl9) and mounted the Linux proc and sys filesystems that Linux programs expect to find. Then, rather than build a Python myself, I downloaded a prebuilt standalone Linux build of CPython 3.11 (from the python-build-standalone project) and unpacked it into /opt/python. Pointed at pip install open-webui — where ready-made Linux wheels actually exist — that Linux Python installed it cleanly. The whole saga is exactly the kind of thing the official Docker image exists to spare you, which is a little ironic given that leaving Docker behind is what put me here in the first place. One gotcha worth passing along: those Linux proc/sys mounts have to survive a jail restart, or the Linux Python quietly stops working the next time the machine reboots. I made them permanent by adding the mount to the jail’s fstab from the host: iocage fstab -a openwebui "linprocfs /compat/linux/proc linprocfs rw 0 0" To keep it running, I wrote a small rc.d service so the jail starts Open WebUI on boot. Under the hood it just runs open-webui serve on port 8080, wrapped in FreeBSD’s daemon for a pidfile and a log. One environment variable in there is worth calling out: AIOHTTP_CLIENT_TIMEOUT=600 That’s a ten-minute client timeout. With a hosted service, responses come back fast; with a big model running on a single 3090 Ti, a long answer can genuinely take minutes, and the default timeout would hang up on the model mid-sentence. Bumping it to 600 seconds fixed that. The first thing to do once you can reach Open WebUI in a browser is to point it at the models. Open WebUI talks to inference backends through an OpenAI-compatible API, so in its admin settings I set up a Connection to my local inference server at http://inference-server-ip:9292/v1. That’s llama-swap over on the ryzen desktop — from Open WebUI’s point of view it’s just another OpenAI-compatible endpoint, and behind that address llama-swap and llama.cpp do all the swapping and inferring we set up earlier. Web search One thing I really wanted was for my local models to be able to search the web, the way the hosted assistants do. Open WebUI supports this out of the box if you give it a search backend, and the self-hosted option is SearXNG — a metasearch engine that queries the big search engines on your behalf and hands back the results, without the tracking. SearXNG got its own jail, also on the FreeBSD box, right next to Open WebUI. Unlike the Open WebUI jail, this one is completely stock — SearXNG is pure Python and runs on FreeBSD natively, so there’s no Linux layer and none of the extra permissions: iocage create -n searxng -r 13.5-RELEASE \ ip4_addr="em0|searxng-ip/24" \ boot=on Installing SearXNG inside the jail was refreshingly boring — especially next to the Open WebUI saga. Because it’s plain Python with no exotic dependencies, it’s available as a native FreeBSD package, so the whole install came down to pkg install py311-searxng-devel and enabling its service. No Linux base, no standalone Python, no wheels compiled from source. Implement this: setting up SearXNG in the jail # inside the jail (iocage console searxng) pkg install -y py311-searxng-devel # generate a secret key to paste into the config openssl rand -hex 32 # edit /usr/local/etc/searxng.yml: # - set server.secret_key to the value above # - add "json" to search.formats (see below) so Open WebUI can read results vi /usr/local/etc/searxng.yml # enable and start the service (it listens on port 8888) sysrc searxng_enable=YES service searxng start sockstat -l4 | grep 8888 # confirm it's listening The one setting that isn’t optional: by default SearXNG only returns HTML, but Open WebUI needs JSON back, so search.formats in searxng.yml has to include it: search: formats: - html - json Then, in Open WebUI’s admin settings, I pointed web search at that jail: http://searxng-ip:8888/search?q= Now when I ask a local model about something current, Open WebUI runs the search through SearXNG, feeds the results into the model’s context, and I get an answer grounded in today’s web instead of the model’s training cutoff. Appendix: recreating the jail from scratch If you’d like to reproduce this exact setup, here’s the whole thing start to finish, with all my false starts stripped out. It assumes a FreeBSD host with iocage. 1. On the FreeBSD host — create the jail: iocage create -n openwebui -r 13.5-RELEASE \ ip4_addr="em0|openwebui-ip/24" \ defaultrouter="gateway-ip" \ boot=on \ enforce_statfs=1 \ allow_mount=1 \ allow_mount_linprocfs=1 \ allow_mount_linsysfs=1 \ allow_raw_sockets=1 iocage start openwebui 2. Inside the jail — the Linux base and the filesystems it needs: iocage console openwebui # Rocky Linux 9 userland: the linuxulator compatibility layer pkg install -y linux_base-rl9 # Linux binaries expect /proc and /sys; create and mount them mkdir -p /compat/linux/proc /compat/linux/sys mount -t linprocfs linprocfs /compat/linux/proc mount -t linsysfs linsysfs /compat/linux/sys 3. Back on the host — make those mounts survive a reboot: iocage fstab -a openwebui "linprocfs /compat/linux/proc linprocfs rw 0 0" iocage fstab -a openwebui "linsysfs /compat/linux/sys linsysfs rw 0 0" 4. Inside the jail — a self-contained Linux Python, then Open WebUI. No FreeBSD Python, Node, or build toolchain required; a standalone Linux CPython plus prebuilt Linux wheels does it all: fetch -o /tmp/python.tar.gz \ https://github.com/astral-sh/python-build-standalone/releases/download/20240726/cpython-3.11.9+20240726-x86_64-unknown-linux-gnu-install_only.tar.gz mkdir -p /opt/python tar xzf /tmp/python.tar.gz -C /opt/python --strip-components=1 /opt/python/bin/pip3 install open-webui 5. Inside the jail — the rc.d service. Save this as /usr/local/etc/rc.d/openwebui: #!/bin/sh # # PROVIDE: openwebui # REQUIRE: NETWORKING # KEYWORD: shutdown . /etc/rc.subr name="openwebui" rcvar="openwebui_enable" pidfile="/var/run/${name}.pid" logfile="/var/log/${name}.log" procname="/opt/python/bin/python3.11" command="/usr/sbin/daemon" command_args="-f -p ${pidfile} -o ${logfile} /usr/bin/env DATA_DIR=/app/backend/data AIOHTTP_CLIENT_TIMEOUT=600 AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST=10 /opt/python/bin/open-webui serve --host 0.0.0.0 --port 8080" load_rc_config $name : ${openwebui_enable:="NO"} run_rc_command "$1" Then enable and start it: chmod +x /usr/local/etc/rc.d/openwebui sysrc openwebui_enable="YES" service openwebui start Open WebUI is now live on http://openwebui-ip:8080. The last step is in its web admin panel: add an OpenAI-compatible connection pointed at http://inference-server-ip:9292/v1 (llama-swap on ryzen), and set the SearXNG search URL to http://searxng-ip:8888/search?q=