Deep, AI-written investment thesis on 500+ leading US stocks, grounded in filings, earnings, and the numbers.
Looking at buying an existent Crypto Publication Needs to have atleast 500k + monthly views and a good enough amount of years sharing news to the world DM's open submitted by /u/amu4biz [link] [Kommentare]
The Setlur et al result that scaling test time compute without verification or RL is provably suboptimal keeps showing up in my reading and I think it deserves more weight than the "yet another scaling paper" treatment it got. The core claim is that verifier based methods, RL or search guided by a verifier, dominate verifier free methods like distilling successful traces, given a fixed compute budget, and the gap widens as the test time budget grows. What I find underappreciated is how cleanly this maps onto what the deployed systems are now converging on. The single agent ReAct loop is the verifier free extreme, you sample a trace and keep it, maybe with some self reflection that is still the same model grading itself. The multi agent setups that actually move numbers split the verifier off into a separate process. Apodex is the most explicit example I have seen, they train the team behavior in and run a verification team, conflict reviewer, fact checker, draft reviewer, that does not share the reasoning trace, and the reported lift is coming from the verifier not from added parameters. Same trained model, heavy duty mode adds double digits on BrowseComp and FrontierScience-Research. That is exactly the regime the theory predicts, the verifier is where the gain lives. The reason I think this matters beyond benchmark watching is that it reframes where the next chunk of capability comes from. If you believe the VB over VF result, then the path is not just bigger models or longer traces, it is better verifiers that are structurally independent of the generator. The pseudo correctness framing fits here too. The failure mode the verifier has to catch is not the obvious hallucination, it is the answer that passes every self check but is still wrong, and that failure mode is invisible to any verifier that shares context with the generator. What I want to hear from others is the open questions. My list. How much of the verifier gain is transferable to domains without clean reward signals, since the math proof case is the easy one. Whether the independence has to be architectural, separate agents, or whether a sufficiently disciplined prompt separation on one model gets you most of the way. And whether the VB advantage keeps widening or saturates once the verifier itself becomes the bottleneck. The practical version of this for anyone building. If your agent loop has the same model reviewing its own work, you are in the VF regime and the theory says you are leaving capability on the table. The cheapest structural change is to make the verifier a different process with denied context, even if it is the same weights. submitted by /u/Mysterious_Sign_9501 [link] [Kommentare]
Worrying about whether AI can do your job is a blind alley, Cory Doctorow argues. The real danger is AI’s bubble: a speculative fantasy built on convincing bosses to replace workers with systems that can’t actually do what their salesmen promise.
On the 67th edition of the TOP500, LineShine debuts as the new No. 1 system, ending El Capitan's run atop the list and becoming the fifth Exascale system overall. It is the first China-based system to lead the TOP500 since Sunway TaihuLight in 2017. LineShine is installed at the National Supercomputing Centre in Shenzhen (NSCS), China, and was built by the Shenzhen Cloud Computing Center. It submitted a debut measurement of 2.198 Exaflop/s on the HPL benchmark, more than 20% ahead of the No. 2 system, using 13,789,440 cores. The system is based on the custom "LingKun" platform with 304-core LX2 processors running at 1.55 GHz, the proprietary LingQi interconnect, and Kylin OS. LineShine also takes over the No. 1 spot on the HPCG ranking with 22.00 Petaflop/s. On the HPL-MxP benchmark, which measures mixed-precision performance, LineShine debuts in fourth at 7.92 Exaflop/s with a more modest 3.6x speedup, consistent with its CPU-only design. El Capitan, Frontier, Aurora, and JUPITER Booster all remain Exascale-class systems and now occupy No. 2 through No. 5, all still installed at the same sites as last edition. The El Capitan system at the Lawrence Livermore National Laboratory, California, USA, moves to No. 2 on the TOP500. The HPE Cray EX255a system holds at 1.809 Exaflop/s on the HPL benchmark. LLNL's 17.41 Petaflop/s on HPCG now places the system No. 2 on that ranking as well, behind LineShine. El Capitan has 11,340,000 cores and is based on AMD 4th generation EPYC processors with 24 cores at 1.8 GHz and AMD Instinct MI300A accelerators. It uses the Cray Slingshot 11 network for data transfer and achieves an energy efficiency of 60.94 Gigaflops/watt. The Frontier system at the Oak Ridge National Laboratory, Tennessee, USA, is the No. 3 system on the TOP500, holding at an HPL score of 1.353 Exaflop/s. Frontier is based on the HPE Cray EX235a architecture and is equipped with AMD 3rd generation EPYC 64C 2GHz processors. The system has 9,066,176 total cores and also relies on Cray's Slingshot 11 network for data transfer. The Aurora system at the Argonne Leadership Computing Facility, Illinois, USA, holds the No. 4 spot on the TOP500 with 1.012 Exaflop/s on the HPL. Aurora is built by Intel based on the HPE Cray EX - Intel Exascale Compute Blade, which uses Intel Xeon CPU Max Series processors and Intel Data Center GPU Max Series accelerators communicating through Cray's Slingshot-11 interconnect. The JUPITER Booster system at the EuroHPC / Jülich Supercomputing Centre in Germany moves to No. 5, still measured at exactly 1.000 Exaflop/s and remaining the first European Exascale system. JUPITER - JU Pioneer for Innovative and Transformative Exascale Research is located at the Forschungszentrum Jülich campus in Germany and is operated by the Jülich Supercomputing Centre. It is based on Eviden's BullSequana XH3000 direct liquid-cooled architecture, utilizing Grace Hopper Superchips. Rmax and Rpeak values are in PFlop/s. For more details about other fields, check the TOP500 description. Rpeak values are calculated using the advertised clock rate of the CPU. For the efficiency of the systems you should take into account the Turbo CPU clock rate where it applies.
Moody Mush: an Interactive Mushroom Robot: Pat it, poke it, or let it glow softly on your desk. Moody Mush responds with changing moods expressed through light, sound, and motion. It’s an interactive mushroom robot designed to be beginner‑friendly while still aiming for an expressive and cut…
LLMs have broken legibility of effort - our ability to tell, at a glance, whether something took a human real work. What happens next?