Rendered at 05:52:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
01100011 14 minutes ago [-]
Single shot or with reasoning enabled? My experience is that reasoning dramatically reduces hallucinations and improves output quality. I don't trust models without it.
Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.
simianwords 2 minutes ago [-]
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.
There's no reproducible set either. I'm not gonna trust this report.
sixtyj 33 minutes ago [-]
I would prefer to use agent-assisted python scripts that chatbot.
k7peak 29 minutes ago [-]
Agreed, this works really well for me. Double check the math/python, execute many times without a LLM that can change o
demibabs 28 minutes ago [-]
I like how FT makes me accept cookies from their 46 “technology” (advertising) partners before showing me that the article is behind a paywall anyway.
jb1991 22 minutes ago [-]
You actually like that? I find it kind of annoying.
Wololooo 18 minutes ago [-]
No they do not like it, it is a figure of speech to underline how much they do not like it.
Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.
There's no reproducible set either. I'm not gonna trust this report.