The article kind of dismissed Microsoft Soundscape as a failure, but for blind people who want awareness of what is around I still don't think there has been anything that beats it. I work on one of it's successors, Soundscape Community [0]. It is based on the code Microsoft released as open source when they shut down the project.
The problem with all implementations so far is they are slow, verbose, and clunky.
People working on these projects should study pre-digital forms of non-written communication. Commodities exchange pit floor hand signals. Pilot-ATC comms. Military radio. Etc.
If I want Audio AR, I want it to be like a co-pilot flying a jet with me. Crisp, concise, zero fluff.
Make it a slider for people who want to get chatted up vs those that don't.
I was bracing a bit before clicking the link thinking about chances that this is about KU100 and 3Dio Free Space selling to consumers, not exactly like hotcakes but like nicer winter overcoats or more premium lawnmowers. OK, we're not there yet.
It sounds like the title should be "the potential of audio AR", as nothing has really risen yet, and hasn't in the two years since the article was written
Aar is a welcome enhancement (that has its own, but potentially less privacy issues than video capture). There's a whole set of aar enabled glasses that focus more on the user via audio rather than reckless complete video capture like so many "smart" pervert glasses being shilled by large advertising companies.
> They deliver high-quality sound and give developers access to a range of controls and sensor data, from volume to head movement to biometrics. Built-in microphones and simple controls make it easy to trigger assistants and audio apps.
To get good context, you need video data. To get good video data, you need cameras, lots of them, on your head.
To get them on your head you need them to be small and light
To get them to run you need batteries. But the maximum battery size you can get, without wires is about 1-1.5 watt hour.
Then, without making a custom wireless protocol, the maximum bandwith you can reliabily expect to user (with an iphone) is 1megabit.
That means you need to compress the world around you to 1megatbit a second.
Now, with a bit or work, like using eye tracking to segment what you are actually looking at(usually a 64x64 pixel image at 15-20 frames a second), and occasional wide angle view when the scene changed, you can build really good context, even reading a book.
But, you also need good location data, along with room classfication to get context. You can't really use GPS because they aren't accurate enough, and don't work all that well indoors. So you use visual odometry.
Once you have all that, you then now need to teach the machine to understand that 1mbit stream to work out if it knows the answer you need. oh and that has to work in a device that has ~10-15 watthours. or offload to a bit boy machine over a patching network.
i guess my reason for knowing that you are a moronic blowhard that likes to self suck is bc this post is about audio and you went straight to video. even in 2020 the engineers building the AR glasses at Google knew the interface was almost entirely audio bc video overlays are maddening.. if u built "AR glasses" anywhere real instead of in your basement using off the rack components then u would already know the refinement past VR led to an almost entirely monochrome interface that is designed to get out of your way as quickly as possible. You would also know that video is absolutely not required for anything in an AR experience bc everything can be sampled through RF... but u dont know any of this so I'm guessing your AR glasses were a horrible unfinished abortion that u self funded by being a complete idiot aka substandard engineer. Go back to party city where u belong.
At no point did I mention video output. If you _read_ the aria website, you'll see that those do not have a display, because they cost too much money, eat a shit load of power, and require a huge amount of packaging space.
There is a reason why meta's AR glasses system is 6 years late. (Orion was meant to be released as a product in q4 2020, then an SDK, and its pathetic state its in now)
> You would also know that video is absolutely not required for anything in an AR
Ok lets walk through a sample scenario:
I am visually impaired & I walk into a kitchen, I walk towards the electric hob which someone else has only just turned off. This is not my kitchen is I am only vaguely familiar with the layout. My hand goes out to feel for the side of the counter, however its getting dangerously close to hob. Using just RF, how will you get the pose of the hand, identify the hob's location, and identify that the hob is hot?
Oh course you can get an AR experience with just a phone and rough location, but its not that good or compelling. its similar to having notifications read out.
Sure you can have speaker diarisation so you can go full "x said y to z", so you can ask your AR assistant "what did x think about b" but again, you are asking the device, rather that the device knowing enough to provide valuable insight at the right time. The point about AR is that you are being given the right information at the right time without having to ask for it. Again location really helps, as does room classification (ie no recording in the toilet)
One of the best AR demos I have experienced was from the sound team in the sister org. They made a navigation system that used spatial audio and my teams SLAM stack to help you navigate though a forest. Instead of using TTS to say "move x steps forward" it used the sound of a dog. The dog panting was placed on the path, and if you strayed from the path it would get more and more upset until it started whining. You just followed the dog noised and you could walk about with your eyes shut.
Again none of that has a display, but it used cameras to provide precise location.
> refinement past VR led to an almost entirely monochrome interface
The monochrome is interesting, because there are a bunch of reasons why its that way. Most of it is price, colour waveguides are really view dependent, prone to sparkles is another reason. The colour reproduction of the early versions were terrible. Think projector screen in full sun. Finally, they are bigger because they need more power to run and more optics to massage the colours into the right place. Making waveguides is really fucking hard, if you want them curved, even more hard. They also tend to need exotic materials like silicon carbide. Silicon Carbide is a bastard because its almost as hard as diamond and is fairly difficult to etch chemically.
Finally different wavelength LEDs have different efficiency, its much cheaper to have single wavelength microprojectors, because getting all the wavelength at the same output, is expensive.
> I'm guessing your AR glasses were a horrible unfinished
lol if I wanted a display, we'd just use an oculus/pico. It also has a much bigger battery and a massive processor with a boatload of ram. The glasses are real and if you're at the right university, you can use them for your projects.
> that u self funded by being a complete idiot aka substandard engineer.
Haha nice. No, my team was paid to do this, and it made the CTO very happy. The algorithms that we developed are state of the art and in a whole bunch of products.
Also to put a fine point on it "context is hard" is where u lost me which is to say immediately. Context is 0.1% what is in front of u and 99% what personalization data u have curated for yourself and 0.4% location coordinate and 0.5% immediate explicit captured intent. Notice that video is missing from the equation. The actual point of AudioAR is to leave both the camera AND THE PHONE at home bc its irrelevant. Whoever is paying you should drop your ass immediately and find a better meat proxy for their slop canon.
[0] https://github.com/soundscape-community/soundscape/
People working on these projects should study pre-digital forms of non-written communication. Commodities exchange pit floor hand signals. Pilot-ATC comms. Military radio. Etc.
If I want Audio AR, I want it to be like a co-pilot flying a jet with me. Crisp, concise, zero fluff.
Make it a slider for people who want to get chatted up vs those that don't.
https://patents.google.com/patent/US11726740B2/en?oq=US-1172...
(2022)
https://photos.icloud.com/shared/album/0bcjn4GtAKvPziPqWKZqR...
Fear the Greeks when they give you presents.
But i fear it is too late.
To get good context, you need video data. To get good video data, you need cameras, lots of them, on your head.
To get them on your head you need them to be small and light
To get them to run you need batteries. But the maximum battery size you can get, without wires is about 1-1.5 watt hour.
Then, without making a custom wireless protocol, the maximum bandwith you can reliabily expect to user (with an iphone) is 1megabit.
That means you need to compress the world around you to 1megatbit a second.
Now, with a bit or work, like using eye tracking to segment what you are actually looking at(usually a 64x64 pixel image at 15-20 frames a second), and occasional wide angle view when the scene changed, you can build really good context, even reading a book.
But, you also need good location data, along with room classfication to get context. You can't really use GPS because they aren't accurate enough, and don't work all that well indoors. So you use visual odometry.
Once you have all that, you then now need to teach the machine to understand that 1mbit stream to work out if it knows the answer you need. oh and that has to work in a device that has ~10-15 watthours. or offload to a bit boy machine over a patching network.
ie: https://www.projectaria.com/
At no point did I mention video output. If you _read_ the aria website, you'll see that those do not have a display, because they cost too much money, eat a shit load of power, and require a huge amount of packaging space.
There is a reason why meta's AR glasses system is 6 years late. (Orion was meant to be released as a product in q4 2020, then an SDK, and its pathetic state its in now)
> You would also know that video is absolutely not required for anything in an AR
Ok lets walk through a sample scenario:
I am visually impaired & I walk into a kitchen, I walk towards the electric hob which someone else has only just turned off. This is not my kitchen is I am only vaguely familiar with the layout. My hand goes out to feel for the side of the counter, however its getting dangerously close to hob. Using just RF, how will you get the pose of the hand, identify the hob's location, and identify that the hob is hot?
Oh course you can get an AR experience with just a phone and rough location, but its not that good or compelling. its similar to having notifications read out.
Sure you can have speaker diarisation so you can go full "x said y to z", so you can ask your AR assistant "what did x think about b" but again, you are asking the device, rather that the device knowing enough to provide valuable insight at the right time. The point about AR is that you are being given the right information at the right time without having to ask for it. Again location really helps, as does room classification (ie no recording in the toilet)
One of the best AR demos I have experienced was from the sound team in the sister org. They made a navigation system that used spatial audio and my teams SLAM stack to help you navigate though a forest. Instead of using TTS to say "move x steps forward" it used the sound of a dog. The dog panting was placed on the path, and if you strayed from the path it would get more and more upset until it started whining. You just followed the dog noised and you could walk about with your eyes shut.
Again none of that has a display, but it used cameras to provide precise location.
> refinement past VR led to an almost entirely monochrome interface
The monochrome is interesting, because there are a bunch of reasons why its that way. Most of it is price, colour waveguides are really view dependent, prone to sparkles is another reason. The colour reproduction of the early versions were terrible. Think projector screen in full sun. Finally, they are bigger because they need more power to run and more optics to massage the colours into the right place. Making waveguides is really fucking hard, if you want them curved, even more hard. They also tend to need exotic materials like silicon carbide. Silicon Carbide is a bastard because its almost as hard as diamond and is fairly difficult to etch chemically.
Finally different wavelength LEDs have different efficiency, its much cheaper to have single wavelength microprojectors, because getting all the wavelength at the same output, is expensive.
> I'm guessing your AR glasses were a horrible unfinished
lol if I wanted a display, we'd just use an oculus/pico. It also has a much bigger battery and a massive processor with a boatload of ram. The glasses are real and if you're at the right university, you can use them for your projects.
> that u self funded by being a complete idiot aka substandard engineer.
Haha nice. No, my team was paid to do this, and it made the CTO very happy. The algorithms that we developed are state of the art and in a whole bunch of products.