PRESS-NEWS.org - Press Release Distribution
PRESS RELEASES DISTRIBUTION

Audio-guided self-supervised learning for disentangled visual speech representations

Audio-guided self-supervised learning for disentangled visual speech representations
2025-01-07
(Press-News.org)

Learning visual speech representations from talking face videos is an important problem for several speech-related tasks, such as lip reading, talking face generation, audio-visual speech separation, and so on. The key difficulty lies in tackling speech-irrelevant factors presented in the videos, such as lighting, resolution, viewpoints, head motion, and so on.

To solve the problems, a research team led by Shuang YANG publishes their new research on 15 December 2024 in Frontiers of Computer Science co-published by Higher Education Press and Springer Nature.

The team proposes to disentangle speech-relevant and speech-irrelevant facial movements from videos in a self-supervised learning manner. The proposed method can learn discriminative disentangled speech representations from videos and can benefit the lip reading task by a straightforward method like knowledge distillation. Both qualitative and quantitative results on the popular visual speech datasets LRW and LRS2-BBC show the effectiveness of their method.

In the research, the researchers observe the speech process and find that speech-relevant and speech-irrelevant facial movements are differences in the frequency of occurrence. Specifically, speech-relevant facial movements always occur at a higher frequency than speech-irrelevant ones. Moreover, the researchers find that the speech-relevant facial movements are consistently synchronized with the accompanying audio speech signal.

Based on the new observations above, the researchers introduce a novel two-branch network to decompose the visual changes between two frames in the same video into speech-relevant and speech-irrelevant components. For speech-relevant branch, they introduce the high-frequency audio signal to guide the learning of speech-relevant cues. For the speech-irrelevant branch, they introduce an information bottleneck to restrict the capacity from acquiring high-frequency and fine-grained speech-relevant information.

Future work can focus on exploring more explicit auxiliary tasks and constraints beyond the reconstruction task to capture speech cues from videos. Meanwhile, it's also a nice try to combine multiple types of knowledge representations to enhance the obtained speech representations.  

DOI: 10.1007/s11704-024-3787-8

END


[Attachments] See images for this press release:
Audio-guided self-supervised learning for disentangled visual speech representations Audio-guided self-supervised learning for disentangled visual speech representations 2 Audio-guided self-supervised learning for disentangled visual speech representations 3

ELSE PRESS RELEASES FROM THIS DATE:

From logs to security: How process analysis is transforming access control

From logs to security: How process analysis is transforming access control
2025-01-07
Researchers at the University of Electro-Communications have developed a groundbreaking framework for improving system security by analyzing business process logs. This framework focuses on ensuring that role-based access control (RBAC) rules-critical to managing who can access specific system resources-are correctly implemented. Noncompliance with these rules, whether due to error or malicious activity, can result in unauthorized access and pose significant risks to organizations.   RBAC is a widely used access control model that relies on predefined roles assigned to users. However, as business processes become more complex, ensuring ...

Dronedarone inhibits the proliferation of esophageal squamous cell carcinoma through the CDK4/CDK6-RB1 axis in vitro and in vivo

Dronedarone inhibits the proliferation of esophageal squamous cell carcinoma through the CDK4/CDK6-RB1 axis in vitro and in vivo
2025-01-07
Esophageal squamous cell carcinoma (ESCC) is a severe health threat, being a predominant subtype of esophageal cancer and contributing significantly to cancer-related mortality globally. Despite advancements in combination therapies, patient prognosis remains poor, highlighting an urgent need for novel treatment strategies. In this context, a study explores the potential of dronedarone, an FDA-approved drug, in inhibiting ESCC proliferation through the CDK4/CDK6-RB1 axis, both in vitro and in vivo. The research reveals that dronedarone, ...

Photonic nanojet-regulated soft microalga-robot

Photonic nanojet-regulated soft microalga-robot
2025-01-07
Micro/nanorobots hold exciting prospects for executing different tasks in complex microenvironments due to their small size, high flexibility, controllability, and environmental adaptability. However, traditional rigid micro/nanorobots are still difficult to perform different biomedical tasks in complex and unstructured narrow microenvironments due to their limited flexibility and insufficient deformability. To address this problem, in a new paper published in PhotoniX, a team of scientists led by Professor Hongbao Xin from Institute of Nanophotonics, Jinan University, China, has developed a new soft microalga robot (saBOT). They innovatively used microalga, ...

How do directional connections shape complex dynamics in neuronal networks?

How do directional connections shape complex dynamics in neuronal networks?
2025-01-07
Uncovering the relationship between structure (connectivity) and function (neuronal activity) is a fundamental question across many areas of biology. However, investigating this directly in animal brains is challenging because of the immense complexity of their neural connections and the invasive surgeries that are typically needed. Lab-grown neurons with artificially-controlled connections have the possibility of becoming a useful alternative to animal testing, particularly as we learn how to accurately characterize their behaviour. A research team at Tohoku University used microfluidic devices to reveal how directional connections shape the complex dynamics ...

Drug-resistant hookworms put pets and people at risk

Drug-resistant hookworms put pets and people at risk
2025-01-07
Canine hookworms are becoming increasingly resistant to drugs across Australia, according to new research. Scientists at The University of Queensland and The University of Sydney have identified widespread resistance to benzimidazole-based dewormers which are commonly used to treat gastrointestinal parasites in dogs. Dr Swaid Abdullah from UQ’s School of Veterinary Science said almost 70 per cent of the hookworm samples studied showed genetic mutations that can cause drug resistance. “This is a big problem, as hookworm infections ...

New strontium isotope map of Sub-Saharan Africa is a powerful tool for archaeology, forensics, and wildlife conservation

2025-01-07
A team of researchers led by UC Santa Cruz recently released a sophisticated new map that reveals, for the first time, the unique “geologic fingerprints” for most of the African continent.  The map will help archaeologists, conservation scientists, and forensics experts match artifacts and plant, animal, and human remains found at locations around the world back to their most likely region of origin within Africa, offering new insights on issues ranging from the history of the transatlantic slave trade to modern wildlife trafficking and human migration patterns.  The research team’s ...

‘Sandwich carers’ experience decline in mental and physical health

2025-01-07
People who care for both their children and older family members – also known as ‘sandwich carers’ – suffer from deterioration in both their mental and physical health over time, finds a new study by UCL researchers. The research, published in Public Health, analysed data from around 2,000 sandwich carers and 2,000 non-sandwich carers from the UK Household Longitudinal Study between 2009 and 2020. Sandwich carers juggle the responsibilities of caring for ageing parents or older relatives while raising dependent children ...

A new way to determine whether a species will successfully invade an ecosystem

2025-01-06
CAMBRIDGE, MA -- When a new species is introduced into an ecosystem, it may succeed in establishing itself, or it may fail to gain a foothold and die out. Physicists at MIT have now devised a formula that can predict which of those outcomes is most likely. The researchers created their formula based on analysis of hundreds of different scenarios that they modeled using populations of soil bacteria grown in their laboratory. They now plan to test their formula in larger-scale ecosystems, including forests. This approach could also be helpful in predicting whether probiotics or fecal microbiota treatments (FMT) would successfully combat infections of the human GI tract. “People ...

A change in the weather in the U.S. Corn Belt

2025-01-06
A change in the weather in the U.S. Corn Belt Intensive farming and shallow groundwater affect precipitation patterns The sweeping land use changes and irrigation of the U.S. Corn Belt, along with the influence of the area’s shallow groundwater, have significantly altered precipitation patterns in that vital agricultural region, new research shows. The study, published in the Proceedings of the National Academy of Sciences, focuses on “precipitation recycling” — a process in which the moisture released to the atmosphere by plants, soils, lakes, and other features of the landscape returns to the same area in the form of rain. By using advanced ...

How we classify flood risk may give developers, home buyers a false sense of security

2025-01-06
Common methods of communicating flood risk may create a false sense of security, leading to increased development in areas threatened by flooding. This phenomenon, called the “safe development paradox,” is described in a new paper from North Carolina State University. Lead author Georgina Sanchez, a research scholar in NC State’s Center for Geospatial Analytics, said this may be an unintended byproduct of how the Federal Emergency Management Agency classifies areas based on their probability of dangerous flooding. Known as flood mapping, this classification system describes areas in terms of their likelihood of being flooded each year. These ...

LAST 30 PRESS RELEASES:

Tracing the quick synthesis of an industrially important catalyst

New software sheds light on cancer’s hidden genetic networks

UT Health San Antonio awarded $3 million in CPRIT grants to bolster cancer research and prevention efforts in South Texas

Third symposium spotlights global challenge of new contaminants in China’s fight against pollution

From straw to soil harmony: International team reveals how biochar supercharges carbon-smart farming

Myeloma: How AI is redrawing the map of cancer care

Manhattan E. Charurat, Ph.D., MHS invested as the Homer and Martha Gudelsky Distinguished Professor in Medicine at the University of Maryland School of Medicine

Insilico Medicine’s Pharma.AI Q4 Winter Launch Recap: Revolutionizing drug discovery with cutting-edge AI innovations, accelerating the path to pharmaceutical superintelligence

Nanoplastics have diet-dependent impacts on digestive system health

Brain neuron death occurs throughout life and increases with age, a natural human protein drug may halt neuron death in Alzheimer’s disease

SPIE and CLP announce the recipients of the 2025 Advanced Photonics Young Innovator Award

Lessons from the Caldor Fire’s Christmas Valley ‘Miracle’

Ant societies rose by trading individual protection for collective power

Research reveals how ancient viral DNA shapes early embryonic development

A molecular gatekeeper that controls protein synthesis

New ‘cloaking device’ concept to shield sensitive tech from magnetic fields

Researchers show impact of mountain building and climate change on alpine biodiversity

Study models the transition from Neanderthals to modern humans in Europe

University of Phoenix College of Doctoral Studies releases white paper on AI-driven skilling to reduce burnout and restore worker autonomy

AIs fail at the game of visual “telephone”

The levers for a sustainable food system

Potential changes in US homelessness by ending federal support for housing first programs

Vulnerability of large language models to prompt injection when providing medical advice

Researchers develop new system for high-energy-density, long-life, multi-electron transfer bromine-based flow batteries

Ending federal support for housing first programs could increase U.S. homelessness by 5% in one year, new JAMA study finds

New research uncovers molecular ‘safety switch’ shielding cancers from immune attack

Bacteria resisting viral infection can still sink carbon to ocean floor

Younger biological age may increase depression risk in older women during COVID-19

Bharat Innovates 2026 National Basecamp Showcases India’s Most Promising Deep-Tech Ventures

Here’s what determines whether your income level rises or falls

[Press-News.org] Audio-guided self-supervised learning for disentangled visual speech representations