The image is being processed, do not close the tab or the browser until it is finished, thanks.

Step-by-Step Facial Image Acquisition Protocol

This section details the acquisition protocol described in our academic paper, organized into four parts. We start with the essential ethical considerations and IRB approval. The next section introduces the cameras set up, then we cover the capture system’s recording workflow, the standardized photos set up, and finalise with the 3D facial models acquisition device. 

If you are interested in having access to the additional materials needed to replicate this setup at your institution – IRB forms, tripod adapter models, synchronization code, amongst others, – please request access filling the form at the bottom of the page . We can also provide support answering any questions that you might have regarding the acquisition protocol.

In the following sections we will describe the details on how the setup for facial image acquisition has been created and organized.

ETHICS COMMITTEE

This study was reviewed and approved by the Research Ethics Committee at University of Granada, under the certificate No.:4130/CEIH/202. As part of this agreement and to comply with the data protection laws, the identity of the participants the data collected is being pseudo-anonymized using unique ID codes to obscure personally identifiable information.

CAMERAS SETUP

Standardized facial photographs, CCTV recordings, and 3D facial models are collected at several designated locations, including both indoor and outdoor environments, in order to capture variations in lighting conditions. The video recording setup was designed to simulate typical identification scenarios encountered in forensic casework, such as images from public transport, ATMs, shops, or shopping malls. For this purpose, both consumer and commercial-grade cameras are used. In addition, a depth-sensor camera is included in the setup to provide subject-to-camera distance data. Facial photographs are then taken using a Digital Single-Lens Reflex (DSLR) camera positioned at a standardized distance from the participant. Three images are captured for each individual: one frontal and two lateral views (left and right). Finally, three-dimensional facial models are acquired using a specialized photogrammetry system, which captures the face from multiple angles and generates detailed models representing different facial expressions.


Figures 1 and 2 illustrate the configuration of the system at two different locations. Although the same cameras are used, their positioning varies depending on the environment, while ensuring that the depth sensor remains within the range of the other cameras. Table 1 summarizes the specifications for each of the acquisition devices, as well as other relevant information.

Figure 1. Camera set up featuring commercial CCTV cameras and the 3D model acquisition device, located at University of Granada. a) 3D facial scan device: DI4D, b) Tapo C120, c) depth sensor camera Intel Realsense D455, d) Dahua DH-IPC-HDBW2831EP-S-0280B-S.
Figure 2. Camera set up featuring commercial CCTV cameras and the standardized photos camera, located at Panacea Cooperative Research’s office in Ponferrada. a) Dahua DH-IPC-HDBW2831EP-S-0280B-S, b) checkerboard used as a visual clapperboard, c) Tapo C120, d) depth sensor camera Intel Realsense D455, e) Hikvision DS-2CD2046G2-IU, f) NIKON D-5200 camera for standardized photos camera, g) chair for subjects at a fixed distance from the camera.
ReferenceCamera typeSpecificationsResolution 2024 price
Focal lengthFoVCMOS
Dahua DH-IPC-HDBW2831EP-S-0280B-SCommercial surveillance camera2.8 mmH: 105° V: 56°1/2.7”3840 x 2160 px160 €
Tapo C120Consumer camera3.17 mmH: 103° V: 55°1/2.9”2560 × 1440 px40 €
Hikvision DS-2CD2046G2-IUCommercial surveillance camera2.8 mm H: 100.2º V: 54.7º1/3″3840 x 2160 px250 €
Intel Realsense D455Depth Sensor1.93 mmH: 90º V: 65º1/4″ OV 97821280 × 800 px540 €
ReferenceCamera typeFocal lengthISO rangeResolutionPrice range
NIKON D-5200+NIKKOR 18-55mm f/3.5-5.6G II VRDSLR Camera18–55 mm100–64006000 × 4000 px400-600€
DI4D SNAP 6200Photogrammetry systemFocal length and ISO range can be manually altered. Not specified24 mpx30,000 € (2017)
Table 1. Camera specifications, taking into account the type of camera, reference and placement.

VIDEO CAPTURE SYSTEM

The video capture system is composed of three commercial CCTV cameras and a depth-sensor camera. These devices are placed at different heights and angles to reproduce conditions similar to those found in real surveillance environments. All cameras are synchronized so they record at the same time, allowing the subject to be captured from multiple viewpoints. The depth-sensor camera (Intel RealSense D455) acts as the reference point of the system. The other cameras are aligned relative to its central position, which serves as the origin of a shared coordinate system (x, y, z). Once the setup is calibrated, the position of each camera is defined in relation to this reference point. This configuration allows each CCTV camera to estimate the subject-to-camera distance (SCD) based on the subject’s position within the image.

The other cameras, the Hikvision DS-2CD2046G2-IU, Dahua DH-IPC-HDBW2831EP-S-0280B-S and Tapo C120, require a custom adapter to be mounted on a tripod. Any tripod capable of supporting 2–5 kg can be used, and extension tubes may be added to reach heights between 2 and 3 meters. The cameras are positioned at different heights to replicate common surveillance setups. The RealSense depth camera acts as the reference point of the system and should always remain within view of the subject throughout the recording and can be seen in figure 3. This is necessary to calculate the subject-to-camera distance (SCD) for the other cameras.

Figure 3. Shows the location of the CCTV cameras’ placement in relation to the subject, with the blue camera as the depth sensor. Image a) shows the cameras’ focus on the subject that allows for the calculation of the SCD. Image b) shows the synchronization of the cameras.

The placement of the cameras include positioning the Tapo C120 around 1.2–1.5 m (similar to an ATM camera), the Dahua dome camera around 1.5–1.8 m (slightly above eye level), and the bullet CCTV camera above 2 m, simulating standard surveillance in open spaces. Exact positioning may vary depending on the recording location. To ensure consistent data collection, tripod positions should be marked on the floor so cameras can be returned to the same location if moved, an example of this is shown in figure 4. While camera placement should remain consistent, recordings are intentionally captured under different lighting conditions, including both natural and artificial light, to better represent real-world scenarios.

Figure 4. Markings on the floor of the camera’s tripod position.

Participants are asked to walk along a marked path toward the cameras while performing simple actions or expressions (such as answering a phone or speaking). Accessories like sunglasses, hats, or face masks may also be used to partially obscure the face, creating more varied and realistic data. Some examples of the marked pathway and the accessories are seen on figure 5.

Figure 5. a) depicts the accessories used as part of the video recording process, where a person picks up one (e.g cap, mask, eyepatch, or sunglasses) and walks in front of the cameras.
b) depicts the marks on the floor to instruct the subjects where to walk, the blue ones are the path, the red is where the subject is visible to all of the cameras.

All cameras are connected to a computer that controls and stores the recordings. The RealSense depth camera connects directly via USB 3.2, while the surveillance cameras connect through a wireless router. Although recordings are stored on the computer, each camera can also save footage to removable memory as a backup. The cameras can be synchronized using an 8×12 checkerboard pattern that acts as a visual clapperboard. For this to work, the cameras’ fields of view must overlap so the pattern is visible to all devices at the same time. The checkerboard can be displayed on a screen placed in the shared viewing area or briefly revealed by a person holding a printed pattern. The coordinate system is defined relative to the depth camera, meaning its orientation determines the orientation of the entire setup. If the depth camera observes a subject from behind while another camera captures them from the front, the measured depth value corresponds to the back of the head. As a result, the calculated subject-to-camera distance (SCD) will reflect that point rather than the face, which should be considered when positioning the cameras.

STANDARDIZED PHOTOS

The next step in this process is capturing standardized facial photographs. For this purpose we located a tripod with the NIKON D-5200 with a AF-S DX NIKKOR 18-55mm f/3.5-5.6G II VR, 1.50-2m away from the subject, using a lens ranging from 18-55mm, adjusted to a 35mm focal. An adjustable chair was placed separate from the wall, to ensure that the height of the individual can be regulated and the face will always be in the center of the lens. It is recommended that both objects –the tripod and the chair– can be marked on the floor to ensure their consistent location. Figure 6 illustrates the placement of the camera and chair.

Figure 6. Standardized photographs setup, the camera and adjustable chair are both at a fixed position with markings on the floor.

FACIAL 3D MODELS

The last step of the circuit is the acquisition of the 3D facial model of different facial expressions (neutral, smile without teeth, smile with teeth, upset, angry and surprised) . For this purpose, a DI4D photogrammetry System is used. It is composed of six cameras that capture the face from different angles and through its own software a 3D realistic model with coloured texture can be processed. This setup also has two dedicated light sources to ensure uniform illumination, an example of how it is set up can be seen in figure 7, we were able to place it in a windowless room, it is recommended to locate the DI4D system in a place where light can be regulated. Participants should avoid any occluding objects, such as glasses, and should secure hair away from the face to prevent obstructions that could cause errors in model reconstruction. The chair used is also adjustable taking into consideration the height of the subject. Additionally, a uniform, solid-coloured background is required, and participants should avoid wearing clothes similar in colour to the backdrop, as this can interfere with accurate model generation.

Figure 7. DI4D system set up, consisting of 6 cameras mounted on a frame, all of them synchronized with each other and with the 2 flashes. The chair is where the subjects will be sitting and performing the different facial expressions.

ADDITIONAL MATERIAL

If you are interested in receiving additional materials for the implementation of this setup at your institution, you can request it by filling out this form.