The team started with an existing approach where two neural networks process the images and audio spectrograms, learning to match an audio caption with images containing a given object. However, they modified the image-handling neural network so that it would split the image into a grid of cells, while the audio network cuts up the …
Continue reading “AI can identify objects based on verbal descriptions”
