Hacker Newsnew | past | comments | ask | show | jobs | submit | more hectdev's commentslogin

I enjoy molding hyper specific software across my whole eco system of tech through prose with LLMs. Iterating over time to hone the right solutions based on very abstract reactions to the results I get. I would consider this a hobby because it costs me money and I get no value out of it outside of the act.

I also woodwork and I remember this discussion there between hand tools, power tools, and CNCs. It's all the same hobby just different entry points and doing it the least automated way doesn't make your work more or less than anyone else's.


> I would consider this a hobby because it costs me money and I get no value out of it outside of the act.

I LOLd very hard at that one, thanks. Very relatable.

A hobby of mine is also woodworking. I can make jigs, I can buy jigs. I can use hand tools, I can use power tools. I can use screws I can use joints. I can even buy made furniture and just assemble it. Though the latter is more of interior designing. But the point is it is a spectrum and everyone has their own niche. SDLC are not the same and there is different value to be derived. But the process of woodworking makes me handy. The experience allows me to elegantly and cost efficiently fix broken things. It enables me to do interesting things in interesting ways and there is more of a community around that. If I automated away all the woodworking and just bought the furniture... Well there isn't much of a community around that except flexing who has more cash. Idk, I'm trying to have empathy for both those who are for and against it.


> I would consider this a hobby because it costs me money and I get no value out of it outside of the act.

Boy isn't that the truth haha, I feel like my previous phase of dotfile ricing has become "LLM ricing."

Not complaining though, it is genuinely fun!


Maybe if hypothetically all stakeholders express all of their concerns and desires clearly in channels AI can pick up constantly.


Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open


That's a summarized and filtered view of the actual reasoning.

OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.


Older models did show the full unredacted thinking trace, but I don't think Opus has ever shown full CoT.

Here is an archived version of Anthropic's API docs saying that Sonnet 3.7 (only) has unredacted CoT on API: https://web.archive.org/web/20260324051339/https://platform....


That’s cool. I knew o1 hid it since launch, so I assumed Anthropic would also have never shown it.


Got it. Thanks


I've started doing this in my home. All lights are smart LEDs. At 8pm, they shift warm, 80%. From 9pm-10pm, they go from 80% to 0%. At 7am, they start going from 0% to 100% cool (I work from home and it helps with focus).

Having my environment reflect the patterns of nature has really been a nice change.


I do something similar with my exterior lights - at sunset my lights come on at ~1% at 2000K. If they detect motion (in my driveway or back deck) they go up to 100% at 2700K. At sunrise they turn off.

Note that bulbs exist that automatically warm as they dim - so in my case the switches are smart, but the bulbs themselves are all dumb.


What type of system are you using outdoors?


* Automation: Home Assistant

* Smart switches: Lutron Caseta. (With the dimmer for ELV lighting so I can use the middle button for specific brightness/warmth levels.)

* Outdoor bulbs: Philips warm glow bulbs. (2200K-2700K warm dimming, 95 CRI - these ramp from 0% at 2200K to 100% at 2700K)

* Outdoor recessed lights: Globe Electric DuoBright LEDs (2000K-5000K warm dimming, 96 CRI - these ramp from 0% at 2000K to full lumens at 3000K at 80% the dimmer, but then will continue to ramp from 3000K to 5000K at full lumens if the dimmer is turned up.)

* Motion sensors: Philips Hue Outdoor sensors.

Non-smart bulb availability is kinda weird, my specific bulbs may be discontinued or unavailable outside of Canada, but I'd expect local alternatives to be available.


Awesome thank you for sharing. I recently bought a very basic low-voltage system just to have something in the garden but would like to upgrade exterior fixtures at some point and your system seems nice.


The bane of my engineering career is working under engineers like this. It's like we forget we are doing a very analog thing (collaboration and building) under the guise of something digital. We should accept that there will be edge cases, there will be crashes. And unless you're actually in a life-saving industry, that is ok. (I say this with the idea in mind of a 10+ year old code bases spanning many new coding patterns that achieves over 99.5% crash-free)


Agreed. The much better question is: Are we locking ourselves out of a feature, or do we just leave a gap for a rare case?

At work we agreed that some use cases are very niche. These have guards in place to log an alerted-upon marker + return HTTP/500. They have not tripped for years by customers. So, it's fine to not support some rare cases and to deliberately leave known gaps in some contexts. As long as you don't close these paths forward if you need them.

On the other hand, we have contexts like our PostgreSQL instances. Those have a very well defined scope and rooting out all known problems has been the right choice. Most issues we have ignored in that scope have bitten us in the butt sooner rather than later. Very hard in some cases, I may add.

Realizing this about a domain is very important.


As someone with a few unused Teenage Engineering things. The real answer is probably rich tech people who love having things that make people say "I'm not sure who the target audience is".


The TE reference is strong!


Work Louder the company behind this, is TE in mechanical keyboard industry.


Not sure if the mechanical keyboards community would agree...


From my experience, Speech-to-text falls way short of Wispr flow and I would assume the ones that are said to be better than that. It lacks context awareness and formatting


I think they have a better agent personality which pushes back and isn't sycophantic. It has been awhile since I've used the others but that's where it locked me in and I've stuck with it.


> isn't sycophantic

Not sure about that one... But I think the true secret sauce for all these models is how they reason. GPT never outputs how it thinks, which "saves on tokens" but Claude absolutely tells you how it thinks, and there's people who use how it reasons about solving problems to finetune smaller open source models, with surprisingly better output.


From my experience, it has not been sycophantic in the sense that it pushes back and questions my own reasoning in healthy ways. There were moments where I felt I was brushing up against actual AI psychosis, and it pushed back on my questioning of its intentions, that it even had intentions. I'll put it this way: I feel comfortable recommending Claude to people who haven't experienced AI yet. As we've learned from early experiences with other models, leading people down paths of believing they understood math in ways nobody else has and even harming themselves, I put Claude as a safer alternative.


I think Opus can still have sycophantic residue that Fable can point out sometimes. Both models though hold their ground so well.

I have got so use to the Claude personality / style of conversation that I really can't be bothered to try these other models anymore. They need to take a huge jump but that seems to be getting harder and harder because of the jumps Anthropic makes.

This Grok version is a joke if it is not even clearing the bar now. I am just getting use to and using Fable more and more. I am also trying not to forget that this is the highly delayed old Fable model that Grok can't even beat on release. There will be a new version that expands the lead in a week or two.

It all harder and harder to judge too. I just had a prompt/response this morning that Fable finally displayed its intelligence and vowed me. That is partly because anything with even the vaguest reference to biology defaults back to Opus.


I _believe_ the term "AI Psychosis" is a "thought-terminating cliché" that readily puts you in a position to disarm any criticism to your point of view, which, if you're aware of it or not, it's a belief in and of itself. I'm more willing to bet your can't have discussions because you're trying to have debates.

But on your actual point, I don't think AI needs to "surpass current experts abilities in all promised fields" as a marker of its ability. The immediate gains has already shown some remarkable promise and more LLMs should have had safeguards around mental health up front. If I were to put it on a scale, I would say it is net positive long term with a strong negative spike up front which was somewhat preventable. But who knows, maybe is just have "AI Psychosis" and you can easily dismiss me.


My only issue with this was the restriction of "Do not look at any data outside of our working folder" is preventing the tool from doing what it does best. I would have given it access to PubMed to pull the latest research on the subject and validate.

I wouldn't consider Claude itself to be the tool that does a job like this, but the tool that pulls in the best data and gives a supported suggestion. And then go through a number of iterations on where it failed to hone in its assessment.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: