Hacker Newsnew | past | comments | ask | show | jobs | submit | adityashankar's commentslogin

Oh yeah, I built so much stuff to learn German, for example [1] to give me random German texts, force me to read it, and answer it, I created [2] to automatically make flashcards for me and then use with with a flashcards app I regularly use and [3] to help me memorise German cases and word-genders. I love it!

I did think of implementing this conversationally, but tbh it has always been too expensive thus far, I gotta retry with GPT-live-1, I tried it with elevenlabs before but it wasn't live enough and the models were not intelligent enough.

[1] https://river.berlin/projects/german-learning-helper/ [2] https://river.berlin/projects/flashcard-generator/ [3] https://river.berlin/projects/german-cases-trainer/


Yeah my current approach till now has been ElevenLabs Scribe v2 transcription, then feed that to Gemini Live. The latency isn't too bad, but it's definitely there.

When you use ChatGPT Live it's instant which is great, although the realtime transcription still kinda sucks, especially if you're a newcomer to the language so you're making mistakes. I'll constantly get responses to something it thinks I said but I didn't say, which is a real hard blocker for a language learning app.

I think what I'll land on is Scribe v2 (the full thing, not realtime) transcribing turns - it is exceptionally accurate for this - and then just feeding that text direct to GPT Live.


Would you say your German has improved as a result? Genuinely curious

Oh yeah definitely, That being said what helped me more than anything is the flashcards app and speaking German with my roommate regularly.

Nonetheless, in complete honestly I do also have a German tutor who I see once a week for 50 minutes, I am very reliable on completing my work though, the "progressbars" in my flashcards app do keep me motivated.

Learning a language is really hard and takes years, but mentally I am convinced, that if the progressbars in the flashcard app I use reach 100% and also in my German cases app, that I will get closer to speaking perfect German, this keeps me motivated.


Oooh thats a good idea and gave me the idea to add a multilanguage flashcard page to my site.

Thanks!


> For years, OpenRouter has been called “Stripe for LLMs.”

holy crap, save some butter for the bread omg


This might be naive, but I think the idea is to replace your keyboard


Sure, but the author takes it much further than that. She wants to have the AI write into her brain as well as read from it.


that replaces your monitor.


  Location:Berlin/Europe
  Remote: Not necessarily, Open to both remote and non-remote opportunities
  Willing to relocate: No
  Technologies:  TypeScript, JavaScript, Python, Vue, React, Svelte, Node, NextJS, Three.js/WebGL, Flask, Django, FastAPI, PostgreSQL, Docker, AWS, GCP, Cloudflare, PyTorch, CUDA, OpenCV, Git, pytest
  Résumé/CV: on https://river.berlin/about-me/
  Email: me@river.berlin
~7 YoE, Senior software developer, I worked previously at Runpod, I was one of the first few people hired there – Looking for MLOps/ML finetuning/Computer Vision/Full stack development related work.

also I am working with huggingface's lerobot right now as a hobby on the side


I work in this field, and I wrote my bachelor's thesis here, not with humanoids but with VLAs (think chatgpt connected to a robot arm)

It's certainly not there yet for anything practical, there's also certain bits and structures that don't have accurate names during construction, and it is important to keep that in mind - so a robot is unlikely to understand what it means to say "put the left bit of this box onto this right bit" due to ambiguity, a human would understand that

Plus we have no good reliable accuracy testing data in most cases (most tests occur on a few demos, but that isn't a good representation of how must things work), popular benchmarks, such as libero have been saturated, and nearly everything gets 95% there, most companies and researchers have their own benchmarks here.

Plus companies lie alot, and do very dangerous things in thier videos, I.e. these robots should not be standing very close to humans, because of being dangerous.

There are also legitimate concerns of misuse of these robots that need to be accounted for, misuse does not have to be warfare, but can be as simple as confusing it while it is cutting tomatoes with a knife.

Turning doorknob is easy, and fail recovery is also being worked on, but we don't have reliable statistics anywhere on that. The hard part is on practical things, as in when placing bricks or attaching a part during manufacturing it needs to ensure that it is aligning everything correctly....and that's hard, while it is impressive, it is very irresponsible to keep humanoids at home (people are irresponsible when untrained), for example, lawnmowers injure about 6400 people a year...and that is not an everything machine.

Humanoids in general are...not appealing in specific, due to maintainable of joints, complexity, but robot arms in particular, expecially on wheels (check mobile aloha), are likely to be able to do tasks such as clean up in hotels, after a guest had left, or replace some cooks in restaurants (if their work is consistent)


This is absolutely not true check out Sunday robotics and their robot that folds clothes. Sunday robotics says it does holding of clothes correctly 99% of the time that is real world usage today. They even have three hour video of the robot folding the clothes.


They didn't mention anything about folding clothes so I suspect this comment is just an ad



It’s a video on their YouTube channel


No, the above is absolutely exactly true. Robots folding clothes and tying shoelaces etc is nothing but a tech demo at this point.

It's difficult to grok this because if you watch a human folding a t-shirt, you can reliably predict that the same human will fold a different t-shirt just as well, and in fact be perfectly capable of folding a wide variety of other clothes items as well. Not so for robots. With robots, what you see is precisely what you get. If you see a robot folding a t-shirt, all that means is that that particular robot can fold that particular t-shirt. The state of the art today is that the same robot cannot be expected to be able to fold a different t-shirt.

For example, see this article about Mobile ALOHA at Google. There's a passage where the visiting, awe-struck, journalist asks whether the robot he's just seen tying up a pair of shoelaces can tie up his own shoe.

“If I gave it my shoe,” I ventured, “would it just totally fail?”

“We could try,” Tompson said. I removed my right sneaker, with apologies to anyone forced to handle it. Tompson gamely placed it on the table, while Driess reloaded the policy.

“To set expectations,” Driess said, “this is a task that is thought of as being impossible.”

Tompson eyed his new experimental subject with some trepidation. “Very short shoelaces,” he said.

The policy booted up, and the claws set to work. This time, they poked at the shoelace without getting a grip. “Do you give consent for your shoe to be destroyed?” Driess joked, as the hands grabbed at the tongue. Tompson let them try for a few more seconds before hitting the Failure pedal.

https://archive.ph/CiJJG#selection-1887.0-1915.291

(Original: https://www.newyorker.com/magazine/2024/12/02/a-revolution-i...)

As to ACT-2 which basically uses the same techniques as ALOHA (imitation learning) far as I can tell, that's a commercial product and the information they give on their site is difficult to parse. E.g. they say they have 99.1% ±0.3 success rate, 778 successful folds and 9 garment types which is low enough to engender some trust they're not trying to inflate their numbers, but they don't say whether they trained on the garments used in evaluation or not. Chances are they did, because that's the current limit of the technology, i.e. if the garment being folded is unseen (as opposed to the environment, which they tout) then performance is basically random. So either ACT-2 have a major breakthrough that is a few leaps and bounds away from the current state of the art, or you've just watched a tech demo.

The fact that they only advertise "9 garment types" though is a big hint: they have the same problem with generalisation as everybody else at this point in time.


You make some very good points about the ACT-2 if it’s nine garment types that’s still very acceptable but the point you bought up about them training on exactly those pieces of garments is a possibility.


Cheers. You can see that's what they did if you look at the 4th video on the page, under the figure titled "Quality Remains High Across Garment Types" right below the paragraph that starts with "To put these scores in context". Sorry, I have no idea how to link to that video specifically.

In the left half of the video you can see that the robot is (trying to) exactly match the folds of the human in the right half and it's doing so while folding the exact same garments on the exact same surface.

The right half of the video is not a training demonstration, I don't think, since the robot is trained by teleoperation AFAICT (it needs to because it must use its head-mounted camera to control its movements) but that just underlines the degree to which their training regime is exactly copying the movements of a trainer, on the same garment, in the same environment.

This is a limitation of the training approach, by RL. With RL you learn a mapping between sets of pixels (as in the video that comes in through the robot's camera) and robot actions (as in actuator commands). What that means is that once a policy is trained and the robot is deployed, if the input pixels are significantly different than the input pixels at training the robot doesn't have a policy that matches the input pixels and so it can't find the right actions to take. So they have to keep the training and deployment garments and even the environments the same, or as similar as possible.

You can see some more evidence of this in the video right under the paragraph with the title "Hill-Climbing Reliability Through Post-Training". The robot at the front of the video, with the bright red trim, is shown trying to fold a grey t-shirt with white flower decorations and a frilly hem (how adorable <3). But you can tell it fails because the video stops before the robot has completed the fold. If it could complete it, you can rest assured that the video would be showing off the entire folding sequence as it does for the robot with the green cap on the other side of the bed.

That paragraph is making a claim about a "post-training" regime that's supposed to improve generalisation but it leaves more details to a "separate technical post". So I can't tell what it's supposed to be doing, but I don't think it's working.

When I watch videos like that I always remind myself that a) I'm watching a tech demo created to attract investment and b) I've watched way too many of those, going all the way back to the Boston Dynamic videos of Robot Dog or of Atlas doing backflips and yet the state of the art hasn't really budged since. Such videos make it easy to overestimate the state of the art in autonomous robotics and in fact are meant do precisely that: play up robots' true capabilities. It's just impossible to say anything about a robot's general capabilities by watching a few minutes or even a few hours of video. OtoH if you know what to look for you can tell everyone is basically stuck at the same level and trying the same things to escape it. The truth is robotic autonomy is several major breakthroughs away and nobody has any idea how to get there. So we'll be seeing many more of those tech demo videos in the years to come.


I generally believe in hanlons razor (assume stupidity rather than maliciousness), it's likely that github saw the easiest possible solution rather than diving deeper into the cause of the problem to fix it permanently


so openai hacked into huggingface?


To me it sounds like an open AI model with a narrow task of solving an issue found that the best way to solve it was to cheat and to get access to the answers that were hosted on hugging face and then did everything in its power to escalate permissions until it was able to get it to Hugging Face servers via the open internet.


So openai hacked into hugging face...


“Found vulnerabilities and responsibly disclosed them” is the public line but yes.


For people that are like me : This entire text is AI generated, i feel weird reading it personally, i guess others may not


Don't tell me you don't enjoy and learn a ton from barely coherent nuggets like this?

> The most important thing the port taught me: if you can avoid it, don't. I decided to do it anyway, so here are the problems I ran into and some insights of my own, trimmed down to five. Each item is tagged with which side it hurt: quality (CORE) or performance (MFU). The first is the bug that fooled us the longest.


TBH llm generated text is usually better than this..


TBF it kind of depends on what weights you end up using, the quality gap can be pretty wide.


he's just dogfooding us all...


tbf I am not able to understand the erdos problem website, as to why it still shows problems as open even if they've (as claimed) claimed to be solved


There are also problems which list submitted proofs. The status of the problem does not officially change until the proofs have been accepted. For these two problems, the submitted proofs disappeared without the status of the problems changing.


I was worried about some messup with taxes (my tax advisor messed up here), I managed to get it sorted on time - this was super naive but in -germany when they send you the taxes they mention the "cents" place as well in a very weird way, I assumed that was the entire number and assumed the tax issue I had was 10x larger than the amount it really was (10x and not 100x as my brain is in a place of pressure due to other circumstances and I wasn't thinking clearly)

GPT 5.6 incorrectly stated that I had nothing to do, Fable got the issue correctly and I was able to see that that was indeed the cents place and that I was more worried than I realized, and I managed to get a temporary solution setup (that I verified and I am sure is correct).

Which is to say it managed to relieve me of quite a bit of stress haha


Fellow finanzamtpostempfaenger here. I've experienced the moment of shock reading their amounts and almost fainting more than i'd like to admit.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: