Like OP, I've learned at lot in this process. I have versions running in the browser now with a custom WebGPU rendering engine. Still lots of jank and vibes, but it's wild to see what models are capable of with the right tooling. (I've had Claude add extensions into Ghidra for Xbox/Wii specific instruction support)
Wild times we're living in. It's great for software preservation though!
I love stories like this. I used Claude to decompile an old Windows program and port it to Linux, it was extremely challenging, I had a lot of success but ended up just rewriting it as a new app.
I have struggled to get this project working on non-Windows. It just hangs and crashes no matter what I do or try on Linux/Mac. It's a very Windows-oriented project that's slowly losing the shackles right now.
Not gonna lie, some tables require way too much work, every software today wants you to be an engineer with 20+ years of some specific experience, what about just double click and let me play the damn game?
Yeah, I hear you but as I mentioned in my reply to your other post there is a long-term silver-lining to having a bit of onboarding complexity. I think it's a big reason the VPin community is well into its second decade and still full of passionately committed contributors freely sharing awesome stuff. If you want drive-by casual pinball there are reasonably-priced pinball systems on PS5/XBox/Switch and Steam that are quite good. While VPin has gotten much easier in recent years, it's still a hobby that requires active engagement. VPin rewards that effort by enabling unbelievably high-quality, flexibility, customization, community enhancements and an ever-growing library of amazing content that'll take years to explore.
I think the mismatch is when people see all these awesome pinball games "Fer Free!" and assume they're going to click Install and be playing in a couple minutes. I tell my friends to expect at least a half-hour before first play - and that they'll have to read and follow a couple pages of good (but not perfect) instructions to understand and configure a few different tools. If you want things to work reliably:
* Stick to only Visual Pinball (not older emulators like Future Pinball).
* Install it with Pinup Popper and set up your screen mapping and controls based on one of the standard default configs.
* Run tables released or updated relatively recently (3 yrs or so).
* Run tables from well-known release groups and authors (like Visual Pinball Workshop).
* Wait to run newly released tables until they've been out a month, have >200 of downloads and >20 positive reviews.
* Don't run add-ons which mod tables until you're experienced.
And once you're past the install phase and have a bunch of tables fully working with all the bells and whistles you want, there's a new tool called VPin Studio that's great for maintaining your VPin system https://github.com/syd711/vpin-studio.
Re Linux: I've only ever run VPin on Windows. I've seen posts from happy people who run it on Linux so apparently it can work very well but cross-platform is newer so there's less info on it. On Windows getting a full VPin install working is just a little cantankerous but no worse than you'd expect when you realize it's several open source hobby projects which pass data in various ways and aren't usually directly tested together.
It works fairly well for me in Linux on WINE, even with Visual PinMAME. I believe I used the all-in-one installer (vpx7setup.exe, although there's a later version now).
My GPU is an AMD Radeon RX 6600 XT, if that makes a difference.
last time i tried on Debian it just worked... their developer testing app also works flawlessly on Android. Arch Linux has an AUR package with the last git and i updated it yesterday and played a bit before bed
Day -X + 1: Engineer at Alibaba finds the vuln and tells Apache. Patch is pushed to git while new release is coordinated.
Day -X: A black hat sees commits fixing the bug. Attacks start happening.
Day 0: Memes start circulating in Minecraft communities of people crashing servers. Some logs are shared on Twitter, especially in China, of people getting pwned.
Day 0 + ~4 hours: My friend DMs me a meme on Twitter. I look up to find the CVE. Doesn't exist. My friend and I reproduce the exploit and write up a blog post about it. (We name it Log4Shell to differentiate it from a different, older log4j RCE vuln)
Day ~1: Media starts picking it up. Apache is forced to release patches faster in response. CVE is actually published to properly allow security scanners to identify it.
Today: AI makes this happen faster and more consistently. Patches probably should be kept private until a coordinated disclosure happens post-testing and CVE being published?
Hard to say what the right move is, but this is gonna be happening a lot over the next 1-3 years. Lots of companies are going to be getting cooked until AI helps us patch faster than attackers can exploit these fresh 0-days.
I’m with you until that last sentence, which I’ve been thinking about as “… until AI code testing, vulnerability scanning, and developer support tools help to limit the number of 0-days and vulnerabilities making it into production”.
So prevention will be more important than ai-assisted rapid containment or patching, though both of those capabilities will be necessary as part of defense in depth.
And some sort of AI-enabled security analysis across the organization’s architecture that is done as part of testing ahead of new software entering production to ID potential vulnerabilities caused by configuration changes or upgrades that modify how systems interact with each other.
I’ve been trying to guess the timeframe for seeing improved secure development, but I’m hoping it’s a bit closer to 6 months - 1 year given the speed of AI adoption and AI progression. May be closer to 3 years as you stated.
In the meantime, is there more to be done than this (not in order)?
- Patch COTS software
- re-evaluate the scoring for previous vulnerabilities
- set up up containment measures capabilities for systems that can’t be patched / high risk vendors
- use frontier model vuln scanning and patching for home grown systems that may have more 0-days than COTS depending on the organization’s capability
- limit the number of vendors / simplifying the tech stack.
I’d be happy to hear how others are thinking about this.
we simply can't absolve ourselves of responsibility in input and expect a hardened output. It's ABSOLUTELY up to the engineers to have test harnesses and scenarios for testing, vulnerability scanning, etc. Just because we can move faster via prompts doesn't mean we neglect the SDLC.
I think there's opportunity to reinvent the pipeline with AI powered tools to assist but the onus is still on the person to ensure they are deploying something that has been tested.
I have been wondering if 1 Million token context contributes here also. Compaction is much rarer now. How does that influence model performance? For some tasks I do, I feel like performance is worst now after this. Also Plan mode doesn't seem to wipe context anymore?
Also a good fallback if your phone screen cracked 2 hours before. But I can imagine part of the challenge they are facing here are scalpers. TicketMaster app 'rotates' the actual ticket every 30 seconds. Can't rotate paper.
I'd think that having a 2nd factor like presenting ID that matches the ticket would be sufficient there though.
128gb is the max RAM that the current Strix Halo supports with ~250GB/s of bandwidth. The Mac Studio is 256GB max and ~900GB/s of memory bandwidth. They are in different categories of performance, even price-per-dollar is worse. (~$2700 for Framework Desktop vs $7500 for Mac Studio M3 Ultra)
I would honestly guess that this is just a small amount of tweaking on top of the Sonnet 4.x models. It seems like providers are rarely training new 'base' models anymore. We're at a point where the gains are more from modifying the model's architecture and doing a "post" training refinement. That's what we've been seeing for the past 12-18 months, iirc.
> Claude Sonnet 4.6 was trained on a proprietary mix of publicly available information from
the internet up to May 2025, non-public data from third parties, data provided by
data-labeling services and paid contractors, data from Claude users who have opted in to
have their data used for training, and data generated internally at Anthropic. Throughout
the training process we used several data cleaning and filtering methods including
deduplication and classification. ... After the pretraining process, Claude Sonnet 4.6 underwent substantial post-training and fine-tuning, with the intention of making it a helpful, honest, and harmless1 assistant.
Does anybody know when Codex is going to roll out subagent support? That has been an absolute game changer in Claude Code. It lets me run with a single session for so much longer and chip away at much more complex tasks. This was my biggest pain point when I used Codex last week.
Like OP, I've learned at lot in this process. I have versions running in the browser now with a custom WebGPU rendering engine. Still lots of jank and vibes, but it's wild to see what models are capable of with the right tooling. (I've had Claude add extensions into Ghidra for Xbox/Wii specific instruction support)
Wild times we're living in. It's great for software preservation though!