Download an interactive system to analyze dietry habits

Transcript
Talking about SenseCams, if we compare adjacent images, we can find two very
different images inside the same event. This is because the images are taken in a low
frequency, compared, for example with the rate of images taken for a video camera, and
a photo could be taken when the user turn himself for a moment, while sitting in front of
the computer. This is the most common reason to trigger false events: slightly change of
position of the user.
Figure 2-9: Examples of false positive boundary [11]
To segment a group of images into events using content-based image analysis, an
adaptation of Heart´s Text Tiling approach is used. [9] With this technique, we take a
reference image, and then compare the block of images previous to it with the block of
images that come afterwards. Each block is represented with the average value of the low
level MPEG-7 visual features of all images present in that block. This way we can solve the
problem of “intruder” images inside a certain event.
New event?
Figure 2-10: The only flower image cannot be considered an event
The MPEG-7 visual features are color structure, color layout, scalable color and edge
histogram. They are calculated making use of the aceToolbox [18].
Color Layout Descriptor (CLD): it is resolution-invariant and it is designed to capture the
spatial distribution of color in an image or an arbitrary-shaped region. To extract the color
layout descriptor, firstly the image is partitioned into 8x8=64 blocks. Each block is
representated by its average color. This result in three 8x8 arrays, one for each color
component (Y: luminance. Cb: Blue-difference chrominance components. Cr: reddifference chrominance components). Then, the DCT transform is applied. We now have
3 matrices of coefficients. The resulting coefficients are zig-zag-scanned and the CLD
descriptor is formed by only 6 coefficients from the Y-DCT-matrix and 3 coefficients from
each DCT matrix of the two chrominance components. The Descriptor is saved as an array
of 12 values. Finally, the remaining coefficients are nonlinearly quantized. [10], [11]
- 11 -