Hacker Newsnew | past | comments | ask | show | jobs | submit | bauerd's commentslogin

They're the exception, not the rule. They get crawled like any other site, but happen to host git repositories. It's not obvious that these are targeted crawls and they likely may just end up in crawling queues a lot generally

By quantifying the result.


What are you talking about?


You want to solve a family of problems using some tool. You measure the relevant solutions performance metrics.

Now, you or someone else vibe-coded a new tool. You created new solutions and measure again.

You have just quantified the result of using the new tool.


How do you measure them if you don't understand what you're doing? A shitty benchmark or small test suite is not how solid software gets made.


You measure results, you benchmark what you care about. It works often enough to be useful


Good, you achieved a 10% speedup for a particular workload that some users said they care about. But how do you find out that was really the feature that should have been built next? How do you prevent adding badly factored code? How to make sure you don't pile on top of existing tech debt in the codebase, that you are solving the most fundamental issues first?


> some users said they care about

But how could they possibly know what they should care about if they don't understand the code?

> how do you find out that was really the feature that should have been built next?

Phew right they don't know. Only the devs understand what software should do.


I actually explained well enough why this requires to a large degree a competent developer to judge.


I think a sufficiently smart non-competent developer can still do this to great effect, but it definitely helps if someone is both a competent developer, smart, and a seasoned user of LLMs.


Just because AI is not yet a god that is better than all humans at creativity and product decisions and design does not mean it is not a huge accelerant right now.


I returned to hand-written notebooks. None of these knowledge management frameworks ever worked for me. Sure my agents have their own semantic memory system but I’m not an agent.


I've taken hand-written, on paper, notes throughout my education, up to high school level. In university, I've switched to hand-written notes on a tablet device. At the end of uni, when I started working on my final thesis & Obsidian was becoming popular, I've switched to it and since then, everything I note is organized into markdown files or at least is somehow referenced through pure text.

Having experienced all of these methods. I've found myself regularly searching through my vault and referencing older material. While the notes hand-written on a tablet, made it much easier to sift through vast amounts of material I was ingesting at the time. I find myself unable to re-use these notes in any meaningful way. They still exist, I can access them, but I don't find myself going back to them. Old paper notebooks? Most of them are destroyed, lost, forgotten to time.

The knowledge management, with easily indexable and searchable contents, has this invaluable perk of being accessible at a moments notice. Keeping potential of being relevant.


I also much prefer handwritten notes. The act of writing something down helps it to stay in my memory more than typing it up. The main for me downside is there's no Cmd+F for my handwritten notes.

I read a post recently about using a local LLM to OCR handwritten notes that I intend to trial soon:

https://andrewdoering.org/blog/2026/remarkable-local-handwri...


The premise of that article is neat, but it looks to be entirely LLM-generated (in addition to the header image) that it’s almost impossible to read.


I did notice that, and the other article I have saved about setting up Ollama on a Macbook Air is even more clearly AI. But I'm assuming/hoping the author had an LLM generate it from their notes/Github project.


I've both a handwritten notebook and an obsidian PKM. When in a meeting, I always have the notebook in front of me. I write plenty of stuff I believe I've to remember. It is really important for my understanding of what is at stake in the meeting. I can't skip it. This live brain dump is required for me. That being said, I almost never come back to what I manually wrote and when I do, I've hard time understanding the logic of my writing... How strange the brain works. It is like I'm using the handwritting to organize the information in my brain. But is done in such a way than I can't no longer understand myself like 24h after the meeting :)

Any thing else goes to the PKM in a structured and often actionable way. It is my notebook, my business development system, my todo list, my brain dump, my list of contacts, etc. What I particularly value is the easy ability to enrich a document with cross-links. This is for me the strongest point of PKM as the one I have in Obsidian.


I think it’s less notebooks or typed notes, and more the issue of cargo cults built around frameworks and systems. Bullet journals are a handwritten PKM, aren’t they? And they were definitely a fad around the fetishized ideal of creating perfect notebooks.


Bullet journals are the opposite of a fad that fetishizes perfect notebooks. Proper bullet journals are very pragmatic.

Some people then add layers of fetishizing to that, because humans gonna human, but the actual idea itself doesn’t deserve to be tarred with the same brush.


Okay, well I don't mean the sense of a perfect notebook that's pristine, so much one that is a panacea for one's knowledge-management needs. Like any productivity trend, it seemed like the bullet journal method inspired a cottage industry of influencers with their bespoke variants and people sharing their journals for the sake of it and heck I'd argue that if there are merchandise- pre-bulleted notebooks sold in stores- then your framework has become a more consumer-driven than pure productivity. But that's no better or worse than GTD, or Zettelkasten for that matter.


Because these productivity increases will be to the detriment of the average person. We won’t get to reap the benefits.


You probably will. It's similar to steam engine and mechanical power loom.


It’s unlikely to stop here. Today is a proprietary 5T parameter model, tomorrow it’s 5 100B parameter models that each specialize to specific applications and you can run them on your phone. The direction of travel is not like you think. The real victim will be medium to large software companies that can no longer rely on the difficulty of reproducing or maintaining or hosting their software as moat.


> We won’t get to reap the benefits.

There is no rational reason to think that. A rising tide lifts all boats.


Not a lot of boat lifting happened in the last few decades. (See "WTF happened in 1971".)


Not a lot of lifting? The standard of living increase for the average person has been substantial. Poverty metrics are falling all over the world.


And inequality is on the rise


>I don’t know how I could work in any of these military industry companies.

You'd sing a different song quite quickly once the threat stops being abstract as you don't get to free-ride on the security a defense industry provides.


The defence industry that would be required to prevent an invasion of the US mainland is at least an order of magnitude smaller than what currently exists to sustain the US empire.


I wouldn’t, but thanks for the reply. I’ve gone through conscription and we are neighbours with Russia. I’ve not lived a day in my life without existential military threat.


You don’t think there are causes worth fighting for? Or that deterrence is immoral? Help me understand.


>What on earth do you do with that many devs on a project like Messenger? I mean, really?

What makes you think it's a simple system to develop at scale?


Cause it's developed. It's been a stable product for a decade. It's like having 10,000 construction workers on a completed building.


10 years ago i wrote a php web chat in 2 hours or so. I pretty much never look at it but the tinestamps suggest it always worked.

I could add more features to it and those will also work.

A friend once worked on an application with a huge team. He often pointed out the window at a large costuction site with a comparable number of people working. He made countless jokes about real work, a real system, real organisation etc Then one day the building was finished and their application kept crashing in production.


It's part of a constantly evolving ecosystem. It's a stable product because reliability engineers make it so and software engineers get the integrations right.


But a messaging program like Signal is fairly similar in scope, but only has like a couple dozen of devs, compared to Messenger's thousands(?).


Signal could definitely use a bunch more devs, if only to fix all the UX bugs I hit on a daily basis.


You haven't used messenger much then, if you're comparing the complexity of the two.


I guess I don't understand which are the features Messenger has that Signal doesn't that requires a billion dollars a year to maintain.


I don't use it. What am I missing?


It still doesn't take a team that large.

When I was at Facebook they decided to re-write Messenger in C. There were people who thought it was a waste of time. There were people who thought it was a great idea. It was a lot of work, took a while, and I wouldn't be suprised if by now it's been re-written to something else.

It's not that hard to make up work, and there's people whose whole job is pretty much just that.


There is no way messenger features or functionality is changing that much year over year.

It’s almost entirely bloat.


> Cause it's developed

You can napkin-math this. How many different team-sized components do you think go into it? If the code were on GitHub, and all they had to do was just update dependencies below them in the stack, and bump the version number for components above them,how many Dependabot PRs would be opened per week for software that's "done"


For a frontend developer? What is there to scale? The app is already done.


>Having said that, some components need to live outside the sandbox (otherwise, who creates the sandbox?).

I run a single-node k3d cluster on each of my MacBooks which uses Agent Sandbox[0] to keep harnesses isolated. Harnesses access models through LiteLLM only. I have aliases for `kubectl exec`ing into whatever harness I need.

[0] https://agent-sandbox.sigs.k8s.io


Fully agree, I only pay the minimum for frontier models to get DeepSeek v4 output reviewed. I don't see this changing either because we have reached a level of good enough at this point.


They can't afford to care about individual customers because enterprise demand exploded and they're short on compute


>On March 4, we changed Claude Code's default reasoning effort from high to medium to reduce the very long latency—enough to make the UI appear frozen—some users were seeing in high mode

Instead of fixing the UI they lowered the default reasoning effort parameter from high to medium? And they "traced this back" because they "take reports about degradation very seriously"? Extremely hard to give them the benefit of doubt here.


Hey, Boris from the team here.

We did both -- we did a number of UI iterations (eg. improving thinking loading states, making it more clear how many tokens are being downloaded, etc.). But we also reduced the default effort level after evals and dogfooding. The latter was not the right decision, so we rolled it back after finding that UX iterations were insufficient (people didn't understand to use /effort to increase intelligence, and often stuck with the default -- we should have anticipated this).


Having a "Recovery Mode"/"Safe Boot" flag to disable our configurations (or progressively enable) to see how claude code responds would be nice. Sometimes I get worried some old flag I set is breaking things. Maybe the flag already exists? I tried Claude doctor but it wasn't quite the solution.

For instance:

Is Haiku supposed to hit a warm system-prompt cache in a default Claude code setup?

I had `DISABLE_TELEMETRY=1` in my env and found the haiku requests would not hit a warm-cached system prompt. E.g. on first request just now w/ most recent version (v2.1.118, but happened on others):

w/ telemetry off - input_tokens:10 cache_read:0 cache_write:28897 out:249

w/ telemetry on - input_tokens:10 cache_read:24344 cache_write:7237 out:243

I used to think having so many users was leading to people hitting a lot of edge cases, 3 million users is 3 million different problems. Everyone can't be on the happy path. But then I started hitting weird edge cases and started thinking the permutations might not be under control.


> people didn't understand to use /effort to increase intelligence, and often stuck with the default -- we should have anticipated this

UI is UI. It is naive to expect that you build some UI but users will "just magically" find out that they should use it as a terminal in the first place.


“after evals and dogfooding” couldn’t have done this before releasing the model? We are paying $200/month to beta test the software for you.


You didn’t anticipate most people stick with defaults?


We anticipated the default would be the best option for most people. We were wrong, so we reverted the default.


It took you a month to revert after multiple complaints. You still blamed users for using the product exactly as you advertised it. And all of your official channels were completely quite for two months, whether it was about new draconian peak hour limits, or about the new defaults, or about exponentially increasing token costs.

People literally started seeing issues immediately as you changed the defaults: https://x.com/levelsio/status/2029307862493618290 And despite a huge amount of reports you still kept it for a whole month.

And then you shipped a completely untested feature with prompt cache misses and literally gaslit users and blamed users for using the product as advertised.

Oh. Remember this https://x.com/bcherny/status/2024152178273989085? "We move fast but test carefully"?

Now untold umber of people have been hit by these changes, so as an apology you reset usage limits three hours before they would reset anyway.

Good job.

Edit. By the way, a very telling sentence from the report:

--- start quote ---

We’ll ensure that a larger share of internal staff use the exact public build of Claude Code (as opposed to the version we use to test new features); and we'll make improvements to our Code Review tool that we use internally

--- end quote ---

Translation: no one is using or even testing the product we ship, and we blindly trust Claude Code to review and find bugs for us. Last one isn't even a translation: https://x.com/bcherny/status/2017742750473720121


Off topic, but I'm hoping you'll maybe see this. There's been an issue with the VS code extension that makes it pretty much impossible to use (PreToolUse can't intercept permission requests anymore, using PermissionRequest hooks always open the diff viewer and steals focus):

https://github.com/anthropics/claude-code/issues/36286 https://github.com/anthropics/claude-code/issues/25018


Yeah, this is so silly.

Anthropic: removes thinking output

Users: see long pauses, complain

Anthropic: better reduce thinking time

Users: wtf

To me it really, really seems like Anthropic is trying to undo the transparency they always had around reasoning chains, and a lot of issues are due to that.

Removing thinking blocks from the convo after 1 hour of being inactive without any notice is just the icing on the cake, whoever thought that was a good idea? How about making “the cache is hot” vs “the cache is cold” a clear visual indicator instead, so you slowly shape user behavior, rather than doing these types of drastic things.


> Instead of fixing the UI they lowered the default reasoning effort parameter from high to medium? And they "traced this back" because they "take reports about degradation very seriously"? Extremely hard to give them the benefit of doubt here.

They had droves of Claude devs vehemently defending and gaslighting users when this started happening


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: