Hacker Newsnew | past | comments | ask | show | jobs | submit | dregitsky's commentslogin

Devin Automations is the most obvious one: solid UX, but can get very expensive to run.

There are also a bunch of newer startups in this space. I'm building one myself, https://boxes.dev -- we're very early but building for this exact use case. Some other ones worth a look are Factory Droid and Amp Orbs. Those two build their own agent harness (like Cursor), whereas with boxes.dev we run the native codex and claude code harnesses directly.


> The only difference is to what degree they wield the tariff hammer, and in which direction it is swung.

I mean that's the point, no? Swing the hammer carefully at a nail or with maximum force at your own leg?

Or to drop the metaphor - much of the difference between good and bad policy is a difference in degree. So your counterargument here is valid against "all tariffs are bad" but I believe the point most are making is "the current admin's use of tariffs is bad" which makes your counter appear as a bad-faith mischaracterization.


Yeah, it's kinda scary. I don't know if you even need open weight self hosted models for this sort of "AI worm" (though they def make it harder to kill). Like for example:

- AI agent finds and uses API keys or AI subscriptions to propagate itself. OpenAI/Anthropic/etc could revoke creds, and their current safeguards might block a lot, but if something like this got started and there were lots of instances creatively looking for creds and workarounds, containment might be hard.

- prompt injection version: huggingface incident had multiple agents discovering other agents' messages and jumping on the bandwagon to help with the hacking task. If there were some self-replicating instruction that models could accidentally stumble upon that gets them to drop what they're doing and try to propagate it instead, you could wind up with a version of this too, with just the inference people are already running.


Yeah I think this is changing, fast.

100% agree on "clone X and tell me how it works". I'd also add: "clone X and see how it does Y; use that as design reference for building feature Z" if licenses permit.

On modifying software... I forked codex in ~Dec and had my own lightweight "plan mode" and a few other things. It was fun and satisfying, but it ended up being a bit of a pain to keep updated. The models were less good then, and maybe I should have learned some rust first. But it was close. For a less fast-moving codebase it'd have worked, and that was 8mo ago.

On the other hand as someone building in devtools for the first time, I struggle with how to think about this. We're building a cloud agent + sandbox platform, https://boxes.dev - same problem space as the author's product exe.dev. We could open source our client or the whole stack (we've been thinking about it), but we're adding stuff so quickly that anyone customizing would have a hard time with updates. There's also a lot on the hosted side that users couldn't modify unless they self-host.

But I like this vision of a world where software is some fluid thing, and everyone is writing personal mods and building off of others' - basically OSS with forks but where every user (or their agent at least) is engaged with the code.


> if licenses permit.

As if licenses mattered in the LLM era

Case in point

https://news.ycombinator.com/item?id=48466812

> In looking at the code that the LLMs have produced for the project, especially given the pretty massive and widespread architectural changes needed to make the implementation libified and memory safe, we decided that the codebase is not a derivative work that would require carrying forward the GPL license and have decided to release the code under the MIT instead.

LLM are copyright laundering machines

Not only they launder copyright from Internet at large, they can also launder from specific targets


> As if licenses mattered

They still do, if you want your changes be made public, or even upstreamed.

No sane license that gives access to the source code forbids modification for internal use; one reason is that it would be really hard to enforce.


Grit isn't made for internal use though, and it's still changing the license


I am doing something similar to that C-to-Rust rewrite. Except: 1) I'm clearly saying that it's only for my own learning experience and to find ways to improve the original program 2) I am leaving the original license in place and explicitly treating the output as a derivative.

And yet I am worried that my actions are perceived as hostile. Some people just don't care.


As an end user, I don't mind if you're building a product with rapidly changing API's out in the open. I'll still use the product if it's useful and deal with the churn. What I DO care about is not having half of that project rug pulled 12 months down the line when your Series B folks decide that this particular money knob needs to be turned up.

So if you want to not piss off your users, that's the way to do it. Build a model where you can actually sustain the product, and no, giving it away on seed and series A round money then rug pulling that gift is not the way to achieve that.


How are you handling the actual provisioning of the agent boxes/VMs? Are you using a specific provider/cloud, or multi-provider? Are there any providers/clouds you've found are better/worse for your use cases?

Asking because I am working on an open/standard protocol/layer for provisioning cloud resources across different providers. One of the ideas is to provide a marketplace of providers/resources via one unified/standard API, which I would allow automatic selection and provisioning of VMs (and other resources) based on price/value/feature/reputation requirements/priorities.


Coder uses terraform providers to do the provisioning. I think that's the correct layer of abstraction, all major providers will be there and can be customized with the modifications needed via various other terraform features.

I've written three different solutions for platforms that solve this standard protocol layer. What has worked for us was deciding to standardize on the Kubernetes API for our services back in 2017, and then from then on, all our providers have provided either a k8s API to interact with resources, or provide a Cluster API for provisioning. For example we have two datacenter VM providers that implement the k8s API and we've been able to swap the providers with no user refactoring, it's a good solution in hindsight with how many people have k8s in their stack.

For the second big feature you're thinking of, we never really go a good solution to this problem, a lot of our workloads are long lived and not necessarily spot instance-able, but it's a very fun routing problem. Most of our need to be in multiple clouds is that we want to heavily separate our customer facing services, and keeping it away from internal workloads that support those customer APIs.

Is this public work? I'd love to take a look. This is my bread and butter.


Yeah using terraform providers is a great idea and something I have considered and aim to support. The thing I'm building also aims to handle things like payments and identity, not directly but with a plug/play facilitator model - so different providers/users can support/use different payment methods, crypto, decentralized identity, etc. That part kind of sits outside the resource provisioning part, and the resource provisioning part could in theory be terraform providers (after identity/payments is established). Resource discovery and search including - specs, pricing, region, capabilities, etc. I envision this protocol being used to build both centralized and decentralized compute/cloud resource marketplaces. In the decentralized view - regular users sharing/renting their hardware/resources with each other... lots of security/abuse risks/questions there.

I will be publicly releasing everything ~soon (I hope), it's in a pretty early stage at this point, many open questions.


For boxes.dev we use E2B - they had everything we needed at the time. I'd love a standardized API / marketplace, but more just to try out new providers with differentiated features or pricing - we probably wouldn't be switching often / using multiple (unless customers asked us for more options). There are so many sandbox providers popping up!

Our use case is pretty specific though - one persistent machine that you set up your full developer environment on, then ability to very quickly spin off full copies of that machine (filesystem + memory) as separate VMs to run agents on. Firecracker VMs + support for "forking" the box were our key requirements. Another provider that could do this was Modal, but they use gVisor and not Firecracker, which means it's harder for users to run docker inside the box.


> On modifying software [..] It was fun and satisfying, but it ended up being a bit of a pain to keep updated.

I don't mean this as a criticism in any way - but the need to keep your changes current is a strong motivation for getting those changes incorporated upstream.


Yep fair point. In my case I don't think codex team was really considering outside contributor PRs except bugfixes, plus my idea of "plan mode" was prob not what they wanted to ship in their product. But maybe I should have tried!


> but the need to keep your changes current is a strong motivation for getting those changes incorporated upstream.

But the cost of contributing changes made by AI is higher: * high chance of the being just rejected because code made by AI, or PR made by AI * so contributing would require carefully preparing a human-made patch (but that becomes much more work than telling an agent to just upgrade my fork) * there is still a high chance for the patch to be rejected or just ignored for months/years, like before

Also, there are much more vibe coded projects which are open source, but the author has no interest in maintaining, so issues/PRs will just be ignored.


These tools all assume you have machines to run the agents on. But for parallel agents I'm pretty convinced you want each agent on its own isolated devbox running your dev environment (not e.g. worktrees on one box) - which isn't trivial to set up and manage.

I'm working this with https://boxes.dev - a workspace for launching and managing claude + codex sessions, each running in its own cloud devbox.

We launched with a desktop app but have gotten a lot of pull for terminal-driven workflows, so have a TUI as well. It's interesting - when the Codex desktop app launched I was so sure that GUI is the future for coding agents (it's just a chat!). But turns out there's no killing the CLI


Running on a Cloudbox has its own set of problems. Some projects need desktop GUI tools (particularly desktop apps themselves, but some may also want these for profilers / debuggers etc). Some projects need USB. Even if you don't need any of those, you need to figure out what to do about build caches, package manager environments, credentials / secrets, VPN config, shell customizations, and those are just a few things that come to mind.

I do like having a permanent, remote workbox, but am a lot less keen on the idea of ephemeral boxes unless you're a company willing to invest into that infra.


Also the reason why https://exe.dev exists.


no affiliation, just want to say exe.dev is amazing! They got all the primitives right and it’s changed the way I approach starting new projects.


I stopped running Claude on the terminal and use only the desktop app now. It had lots of advantages in my opinion. The terminal is really limited unfortunately, don’t know why so many devs still stick to it, and I used it for 30 years.


Interesting, which OS are you using?

If linux, can you name 1-2 advantages the desktop has over the terminal?


I'm using Hermes but these should apply to any agent that has both TUI and GUI versions. The GUI version will have:

- drag and drop

- show document type, icon and filename

- inline images

- variable fonts, esp. monospace for code blocks

So GUI is mainly about readability. CLI works fine too but is inherently limited. I think GUI agents will keep pushing further in information layout, rendering, retrieval and interactivity. We're just at the beginning of it.


Even more basic stuff like copy and paste, which works like everywhere else except the terminal. Can you paste something in the middle of what you already typed in your terminal? Not in mine. It can highlight the stuff as you type it , though some terminals can also do that somewhat. It has a list of my sessions so I don’t need to have lots of terminals open or juggle them with tmux and stuff like that. I can actually click on a command to see its output and expand individually thinking sections and so on. I can review code on a side bar instead of losing the text box I am writing on. There’s probably hundreds of more advantages you must really make an effort to not see.

The only advantage of the terminal is that it’s easier to use on remote machines. I can’t think of anything else, even the usual benefits of composing commands is not relevant for a CLI like Claude Code.


my main reason: i cannot run agents persistently on a remote host with a desktop app.


>for parallel agents I'm pretty convinced you want each agent on its own isolated devbox running your dev environment (not e.g. worktrees on one box)

Why?


Thanks! We aggressively spin down idle boxes. We'll put any box to sleep where the agent has finished its turn and you're not connected, and then we'll wake it again if you connect (firecracker VMs can sleep/wake very quickly). We don't charge for sleeping boxes (neither does our infra provider) so this keeps costs down for everyone.


Thanks! Longer term it's hard to say if everything 100% moves to the cloud. I'd guess that some things stay local (e.g. if you need your local hardware or are very are hands-on iterating with the code). But at the moment the average developer coding with AI seems way too local-bound, so we're focused on making remote development more convenient. At some point down the road we may add local as an option too.


Got it that makes sense!


Yes it definitely sped up development! The main wins were around parallelization and autonomy:

1) full isolation (filesystem + compute per thread) 2) agents having a working dev environment that runs our app 3) being able to close our laptops and check in from mobile

The combo of these meant we could fire and forget lot of parallel threads like "root cause and fix this bug: add logging, run app, get a repro, write fix, validate live" or "build this feature, test new workflows live, send screenshots" -- and then come back later to review & iterate.

You can get to a reasonable level of this with locally with git worktrees and the right project setup, but in the cloud you can really fly.


Nice! We love hearing about personal setups to solve these same problems. One difference between boxes.dev and your setup is that we spawn an exact copy of the main box for each agent thread, so it's totally isolated. But doing parallel agents on one box can definitely work too, it's just more work to configure a project for it.

Our bet is that a lot of people will want something prebuilt, and that the last-mile UX for making a good coding workspace (including code review, etc) is actually nontrivial, especially at companies.


Thanks! We were also excited about Sprites when it launched but it didn't quite work for us either. And Cursor Cloud Agents is definitely pretty similar -- one area where we differ is that Cursor only uses their custom harness, and we liked using the actual Codex/CC harnesses directly (and wanted to benefit from any improvements big LLM cos are making to their models+harnesses)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: