Posts

Showing posts with the label Bag of Words

[CVML] Random Facts and Advice

Some tips and facts that I took from the summer school. They are pretty random but may be usefull for pactitioners of vision. Many may seem obvious - but I just didn't see it before ... Gist doesn't work on cropped or rotated images. Since it does a kind of whole image template matching, this is pretty clear. And maybe it shouldn't - the scene layout is changed after all. For doing BoW Cordelia Schmid suggests (and I guess uses) 6x6 patches and a pyramid with scale factor 1.2. Scaling is done by Gaussian convolution. Ponce uses 10x10 patches to do sparse coding. If you combine multiple features using MKL or by just adding up kernels (which is the same as concatenating features), normalize each feature by it's variance and then search for a joint gamma. This heuristic get's you out of doing grid search over a huge space! Cordelia Schmid thinks that "clever clusters" don't help much in doing BoW. She thinks it's more important to work ...