Prior-Conditioned Dense Prediction: Completion and Matching of Auxiliary Evidence in 2D and 3D Semantic Segmentation

Monday, August 17, 2026 - 01:00 pm
2277 (M. Bert Storey Engineering and Innovation Center)

DISSERTATION DEFENSE
 

Author :  Ziyu Zhao
Advisors: Dr. Song Wang
Date: Aug 17, 2026
Time: 01:00 pm
Location:  2277 (M. Bert Storey Engineering and Innovation Center)
Link:  https://teams.microsoft.com/meet/240422703068254?p=EsBfT7Q4WmKzfnWa8K


Abstract
While the accuracy of semantic segmentation has improved rapidly, extending a trained model to new categories remains expensive: each addition still requires collecting
dense annotations and training again, which limits the use of segmentation where the categories of interest are not known in advance. This dissertation studies how external prior information can assist semantic segmentation in open-vocabulary and few-shot settings, and refers to this as Prior-Conditioned Dense \mbox{Prediction}. In these settings the target category is specified by an input, consumed the same way in training and at test time, rather than learned into the model's parameters, so new categories require no retraining. This is where assistance is needed: the inputs themselves are difficult to use directly, for two different reasons. A category name identifies what to segment but provides no appearance evidence with which to localize the category in a particular image. A labeled example identifies what to segment by showing it, but in 3D segmentation such examples conventionally take the form of point clouds in which the annotation is given point by point, and producing even a few of them is laborious. The dissertation therefore constructs auxiliary prior evidence, generated from the names, transferred from a cheaper modality, or synthesized from what is given, and converts it into a form directly comparable with the scene. Dense matching between the constructed evidence and the scene then yields the labels at every pixel or point. Where the constructed evidence is unreliable, because it crosses modalities or contains generated content, the model additionally adapts it to the scene or weights it by estimated quality.

Three studies instantiate this view, ordered by how much of the evidence must be constructed. In the first study, for open-vocabulary 2D segmentation, category names are paired with diffusion-generated exemplars that supply the appearance the names lack, combined into a dual-prompt cost volume refined at inference. In the second study, for label-efficient 3D segmentation, the point-wise labeled 3D examples are replaced by labeled 2D images, lifted into pseudo RGB-D point clouds by monocular depth estimation and matched to the scene through prototypes in a shared embedding space. In the third study, where a single labeled view leaves the geometry incomplete, further views are synthesized by RGB-D inpainting; since the newly revealed regions are generated rather than observed, learned view and point weights discount them when prototypes are formed. On standard 2D and 3D benchmarks, each study improves over methods that consume the supplied inputs in their original form.

Together, the studies show that auxiliary prior evidence can strengthen dense prediction well beyond what the task's own inputs support, once it is constructed into a comparable
form and consumed through localized matching. They further indicate that adaptation and weighted aggregation are best introduced according to the semantic, cross-modal, or visibility mismatch present in each setting, rather than adopted uniformly

Toward Adaptive Blockchain Design: Enabling Redactability, Referability, and Priority-Aware Processing

Friday, August 14, 2026 - 04:00 pm
online

DISSERTATION DEFENSE

Author :  Matthew Sharp
Advisors: Dr. Chin-Tser Huang
Date: Aug 14, 2026
Time: 04:00 pm
Location: Online
Link:  https://teams.microsoft.com/meet/211658530586897?p=AAPp2DUFHhla2JclD4

Abstract

Traditional blockchain systems typically process transactions using uniform ordering policies and fixed validation procedures, regardless of urgency, workload type, or contextual importance. While this design supports consistency and security, it limits the effectiveness of blockchain systems in applications where time-critical or trust-sensitive operations must be handled quickly without compromising fairness, accountability, or auditability. This dissertation addresses these limitations by introducing adaptive blockchain mechanisms that combine reputation signals, ledgernative structural controls, and priority-aware scheduling policies to support more responsive transaction processing.

In this dissertation, we propose priority-aware and adaptive blockchain mechanisms for time-sensitive and trust-sensitive environments. The motivation is to improve the responsiveness and fairness of blockchain systems when processing heterogeneous workloads that may differ in urgency, trust level, service requirements, or application context. The goal is to provide flexible blockchain mechanisms that can dynamically distinguish between urgent and ordinary work while preserving the trust
guarantees expected from distributed ledger systems. To achieve this goal, the dissertation builds on a smart-marker-based blockchain design that encodes control semantics directly within the ledger. These markers support operations such as branching, merging, referencing, and priority-aware processing while maintaining auditability and consistency. The proposed approach is grounded in distributed consensus, trustbased reasoning, and fairness-aware scheduling. Rather than treating blockchain
processing as a uniform sequence of transactions, this work frames blockchain opii eration as a decision-making problem shaped by dynamic workload characteristics, reputation signals, and policy constraints. By combining reputation-aware evaluation, ledger-native control structures, and bounded scheduling policies, the proposed system balances responsiveness, fairness, and security. This enables structured adaptation without abandoning the accountability and verification properties that make blockchain systems trustworthy. As a result, the dissertation supports the practical deployment of adaptive blockchain designs across distributed applications where timing, trust, and workload differentiation are critical.

The main research questions in this dissertation are as follows:

How can blockchain mechanisms support multiple concurrent computational outcomes without sacrificing ledger integrity?
Our first contribution proposes a smart-marker-based reputational probabilistic blockchain architecture that enables branching and merging within the ledger.
This design allows multiple agents or algorithms to produce concurrent outcomes, which are preserved as branchchains rather than discarded through premature consensus. Each branchchain is associated with a probabilistic reputation score that reflects historical performance and reliability. This approach enables the system to maintain alternative results while guiding decision-making through trust-aware evaluation. Experimental results demonstrate that the architecture supports multi-agent collaboration and improves interpretability in decision-oriented applications.

How can adaptive blockchain mechanisms be applied in real-world, resource-constrained environments?
Our second contribution evaluates the proposed mechanism through an application study in blockchain-based electronic voting. A smart-marker-based branching architecture is used to execute multiple vote-counting algorithms in parallel, improving robustness and decision reliability. The system is implemented on a resource-constrained platform to assess feasibility under limited computational resources. Experimental results demonstrate that the architecture supports scalable processing, low power consumption, and transparent auditability in practical deployment scenarios.

3. How can blockchain systems process heterogeneous workloads in a fair and efficient manner?
Our third contribution introduces a priority-aware scheduling mechanism for blockchain systems. Transactions are classified based on urgency, service tier, and workload characteristics. A scheduling mechanism is developed to allocate processing resources while enforcing fairness constraints such as bounded delay and anti-starvation guarantees. The system balances faster service for highpriority transactions with continued progress for lower-priority workloads. Evaluation metrics include latency, throughput, and contract satisfaction, demonstrating that the proposed scheduling approach improves responsiveness while maintaining equitable access.

4. How can blockchain systems reduce operational latency for timesensitive applications without weakening consensus guarantees?
Our fourth contribution proposes an early admission mechanism that allows privileged blocks to become partially validated and visible before full validation is complete. Admission decisions are governed by reputation thresholds and partial validation criteria. Smart markers are used to record provisional states and ensure auditability. The mechanism reduces time-to-inclusion while preserving the security guarantees of eventual consensus. A game-theoretic model is planned be developed to analyze adversarial behavior and determine safe operating parameters for early admission policies.

This dissertation advances blockchain design from static, uniform transaction processing toward adaptive, fairness-aware, and application-aware operation. By integrating reputation, structural control, and scheduling mechanisms, the proposed
mechanisms enable blockchain systems to support both ordinary and time-sensitiveworkloads within a unified and trustworthy infrastructure.

Medical Image Segmentation: From Perceptual Enhancement to Robustness under Incomplete Data

Tuesday, August 11, 2026 - 02:00 pm
Meeting room 2267

DISSERTATION DEFENSE

Author :  Hongpeng Yang
Advisors: Dr. Yan Tong
Date: Aug 11, 2026
Time: 02:00 pm
Location: Meeting room 2267
Link:  https://teams.microsoft.com/meet/216585525604914?p=nTgWeDedV9AAIGAdFn

Abstract

Deep learning models for medical image segmentation are typically trained under assumptions that clinical data do not always satisfy: distinguishable anatomical boundaries, complete volumetric observations, and access to every required imaging modality. In practice, lesion contrast may be weak, slices may be missing along the depth axis, or MRI sequences may be unavailable. Each condition removes a different form of evidence and calls for tailored compensation. This dissertation develops three task-oriented frameworks for such structured input imperfections, that is, imperfections whose location or identity is known before the prediction is made. Although addressing distinct problems, all three adapt internal representations to the affected information and keep compensation tied to segmentation.

The first framework, Dark Vision Network (DVNet), addresses weak boundary and texture cues in U-shaped networks, by analogy with perception under low illumination. It decomposes skip-connection features into frequency subbands and uses Mamba-based, contrast-oriented fusion to enhance them before decoding. Its plug-and-play design leaves the backbone unchanged and applies to CNN-, Transformer-, and Mamba-based architectures. Across these backbone families, the module lowers boundary distance error in every reported comparison and improves region overlap in most of them, and the same design transfers from volumetric brain MRI to abdominal MRI and two-dimensional microscopy.

The second framework, InterFrameNet, addresses spatial incompleteness when only the endpoint slices of a local window are observed. Rather than reconstructing missing images, it predicts intermediate segmentation features from endpoint features and relative slice positions. An auxiliary delta-aware objective regularises structural changes between adjacent predictions. Experiments under increasing slice sparsity show that segmentation-directed cross-frame prediction outperforms copy- and mean-based filling, and that its advantage widens as the gap between observed slices grows.

The third framework FAR-Seg, addresses missing modalities in brain tumour segmentation. It estimates a target-specific feature for each absent MRI sequence while preserving the acquired-modality features. A predicted class-wise failure map then guides residual correction of the initial segmentation. Evaluation over every non-empty modality combination shows that separating feature completion from output correction improves segmentation accuracy, with the clearest gains where the absent sequence carries evidence specific to one tumour region, such as enhancing tumour when contrast-enhanced T1 is unavailable.

Together, these frameworks support a common design principle for medical image segmentation under imperfect inputs: compensation should reflect the degraded or missing evidence and operate at the stage most directly connected to segmentation. This principle links feature enhancement, cross-frame prediction, and feature completion with output correction within a single task-oriented perspective.
 

Artificial Intelligence-Driven Materials Informatics: From Property Prediction to Crystal Structure Discovery

Monday, August 10, 2026 - 04:00 pm
2265 (M. Bert Storey Engineering and Innovation Center)

DISSERTATION DEFENSE
 

Author :  Sadman Sadeed Omee
Advisors: Dr. Jianjun Hu
Date: Aug 10, 2026
Time: 04:00 pm
Location:  2265 (M. Bert Storey Engineering and Innovation Center)
Link:   https://sc-edu.zoom.us/j/4997546955


Abstract
The discovery of novel materials is essential for advances in energy storage, electronics, catalysis, superconductivity, and many other technologies. However, the vast chemical design space and the high cost of experimental and first-principles methods such as density functional theory (DFT) make large-scale materials discovery challenging. This dissertation develops artificial intelligence (AI) and machine learning (ML) methods to accelerate materials discovery through advances in property prediction, model generalization, crystal structure prediction (CSP), and generative modeling.
In the first topic, we introduce DeeperGATGNN, a scalable graph attention neural network that overcomes over-smoothing through architectural innovations, enabling deeper networks and achieving state-of-the-art performance for materials property prediction. In the second topic, we present the first comprehensive benchmark of out-of-distribution (OOD) generalization in materials property prediction, providing a systematic evaluation of leading models under challenging OOD settings and offering insights into their ability to generalize beyond the training distribution in realistic materials discovery scenarios. In the third topic, we present ParetoCSP, a hybrid AI-evolutionary framework that combines ML interatomic potentials with age-fitness Pareto optimization to improve the efficiency of CSP. In the fourth topic, we extend this framework with ParetoCSP2, which enhances polymorphism prediction through an adaptive space group diversity control mechanism, improved structural initialization, and iterative relaxation, leading to more reliable recovery of multiple stable crystal structures. In the fifth topic, we introduce Diffhedron, an E(3)-equivariant diffusion model that generates crystal coordination polyhedra from chemical compositions, providing a physically grounded intermediate representation for future CSP and inverse materials design problems.

Overall, this dissertation contributes new predictive models, benchmark studies, search algorithms, and generative frameworks that collectively advance AI-driven materials informatics. These developments move the field closer to the long-term vision of autonomous discovery systems capable of reliably linking composition, structure, and properties to accelerate the design of next-generation materials.

Neurosymbolic Methods for Retrieval, Reasoning and Planning

Wednesday, August 5, 2026 - 08:00 am
STB 529, AI Institute

DISSERTATION DEFENSE

Author :  Vedant Khandelwal
Advisors: Dr. Amit Sheth
Date: Aug 5, 2026
Time: 08:00 am
Location: STB 529, AI Institute 
Link:   https://sc-edu.zoom.us/j/82296571999

Abstract
Neural and symbolic methods fail in complementary ways. A neural model can answer almost any question but cannot distinguish correct from plausible. A symbolic model says what correct means but covers only what its authors anticipated. Many ways to combine them exist, and no principle says which one a task needs; this dissertation supplies one.

The symbolic part fixes the space of allowed options, and the neural part ranks within it. What varies is where along the pipeline, from training data to returned output, that space gets fixed. A check decides whether a candidate satisfies a stated criterion without constructing it, whether a coloring is valid, whether a plan is executable, and whether an answer is correct. Placement follows from what a task's check settles, what it costs, and what it leaves open, not from the application. Retrieval admits no check at query time, since relevance is a human judgment, so the space is fixed at the input. Reasoning admits one that settles an answer outright, so it is fixed after generation. Planning admits an equally cheap check but has too few training domains, so it is fixed before learning, and again at generation, where an exact search keeps every step legal while the heuristic ranks.

Applied systems meet several signals at once and take no new rule, since each subproblem gets the placement its own check predicts. Removing each in turn shows what it contributes. The result is a decision procedure from how a task can be checked to where symbolic structure belongs, what it gains over a neural baseline, and when it stops paying.
 

Time Series Segmentation and Dense Human Activity Recognition for Puff Detection

Friday, July 10, 2026 - 10:00 am
online

 THESIS DEFENSE

Author :  Jakub Jerzmanowski
Advisors: Dr. Homayoun Valafar
Date: July 10, 2026
Time: 10:00 Am
Location:  Room 2265, Storey Innovation building

Link:  https://teams.microsoft.com/l/meetup-join/19%3ameeting_NWQzZDZhOTItMjNi…

Abstract
Cigarette smoking remains the leading cause of preventable death in the United States, claiming approximately 480,000 lives annually. Clinicians who tailor cessation interventions are limited by their data: knowing that a person smoked is far less useful than knowing when each puff began and ended, information from which count, duration, and inter-puff interval follow. Wrist-worn inertial measurement unit (IMU) systems can sense smoking unobtrusively, but existing methods classify coarse fixed-length windows rather than the puffs themselves, and so cannot recover this structure. We frame puff detection as dense, per-timestep segmentation, labeling every IMU sample and evaluating at the event level. To our knowledge, this is the first treatment of smoking-puff detection as a segmentation task. Using leave-one-subject-out cross validation, we show that window classifiers, run densely, collapse under strict overlap (event F1 at IoU 0.75 near 0.05), placing predictions in the right neighborhood but the wrong shape, whereas a 1D U-Net adapted to the time domain reaches 0.714 on our 1,500-hour, six-participant in-situ dataset. Decomposing the residual error shows that localization is essentially solved (matched-puff IoU of 0.93 to 0.99); the entire remainder is missed detections, concentrated in one hard participant who drives a pooled miss rate of 0.101. Because the bottleneck is per-person recall, we close it with personalization: warm restarting on as few as ten of a participant’s own puffs cuts the hard subject’s miss rate from 0.38 to 0.05 while boundary quality holds, reframing the problem as one of data and personalization rather than architecture.

CAA-MFA: Context Aware And Adaptive Multifactor Authentication

Thursday, July 9, 2026 - 09:00 am
online

DISSERTATION DEFENSE

Author :  Jonathan Sharp
Advisors: Dr. Csilla Farkas
Date: July 09, 2026
Time: 09:00 Am
Location:  Virtual
Link:  https://teams.microsoft.com/meet/22071693515848?p=LRAW7BhvdNDjot6moR


Abstract
In this dissertation, we studied how to improve multi-factor authentication using context-aware and adaptive authentication methods. We developed the ContextAware Adaptive Multi-factor Authentication (CAA-MFA) framework to enhance the usability and security of authentication systems in dynamic environments. Traditional multi-factor authentication (MFA) systems often rely on static combinations of factors regardless of contextual risk. This limits their effectiveness against evolving threats such as phishing, social engineering, credential compromise, and MFA
fatigue attacks (1). CAA-MFA addresses these challenges by adjusting authentication requirements based on real-time context, trust, policy constraints, and access risk.
The proposed framework treats authentication as a decision-making problem shaped by dynamic risk and policy compliance. It models user and environmental context semantically, evaluates the trustworthiness of available authentication factors, selects
authentication factors through constraint solving, and quantifies access risk using a Risk Level Assessment (RLA) model. By modeling context through ontologies and enforcing factor-selection constraints formally, CAA-MFA supports structured adaptation and scalable policy enforcement across heterogeneous environments (2). The framework also builds on trust-based reasoning for adaptive authentication by assigning trust scores to authentication factor-source pairs and selecting factors that satisfy constraints related to trustworthiness, privacy, usability, and required security level
(3).
The framework was evaluated using 12,000 labeled login attempts, including
iii
10,000 legitimate login attempts and 2,000 attacker attempts. The evaluation measured authentication performance, computational overhead, runtime cost, and the contribution of SAT-based factor selection and RLA. The strongest configuration,using both SAT-based selection and RLA with an SVM classifier, achieved an Equal Error Rate (EER) of 0.0102, Area Under the Curve (AUC) of 0.9965, and F1-score of 0.9852. These results show that CAA-MFA can distinguish legitimate and adversarial login attempts with strong performance while providing a structured method
for adjusting authentication strength according to contextual risk. The main research questions in this dissertation are as follows:


1. How can we model user and environmental context to support adaptive multi-factor authentication?
 

We proposed a semantic context model that captures relevant features from users, devices, behavior, history, and the surrounding environment. These features include user roles, device attributes, network conditions, location, time, behavioral indicators, and privacy requirements. The model uses ontologies to support structured reasoning about contextual conditions, enabling the authentication system to interpret contextual changes and supply meaningful inputs to downstream trust evaluation and factor selection.


2. How can we dynamically select authentication factors using constraintsolving based on contextual trust?
 

We proposed a formal mechanism for selecting authentication factors using trust evaluation and constraint satisfaction. The framework supports both passive and active authentication factors. Passive factors include device identifiers, location data, application history, and other contextual signals, while active factors include biometric input, typing behavior, user-entered PINs, and other explicit verification methods. Each factor-source pair is assigned a trust score iv reflecting reliability, source integrity, and contextual relevance. The selection of factors is modeled as a constraint satisfaction problem, and a SAT solver is used to enforce policy requirements such as minimum trust thresholds, usability constraints, and privacy constraints. This approach enables dynamic adaptation to changing contexts without relying on static authentication workflows.
 

3. How can we quantify the cost versus benefit of utilizing CAA-MFA over traditional MFA?
 

We evaluated the practical trade-offs introduced by deploying CAA-MFA compared to static MFA systems. The evaluation considers usability, computational overhead, implementation complexity, and security effectiveness. The results show that CAA-MFA introduces additional operational cost through context modeling, SAT-based factor selection, RLA computation, and classifier evaluation. However, these costs are justified in environments where authentication risk varies across users, devices, networks, and resources because the adaptive model provides measurable improvement in authentication performance and policy-aware factor selection.
 

4. How can we quantify access risk and use it to adjust authentication strength in real time?


We developed an access-risk model that incorporates contextual factors, authentication strength, historical user behavior, user clearance, and resource sensitivity. The model generates a continuous Risk Level Assessment (RLA) score that supports real-time adjustment of authentication strength. This score helps determine when to escalate verification requirements in high-risk contexts and when to reduce unnecessary authentication burden in low-risk contexts. The RLA model is integrated with the context-aware factor selection framework and evaluated through the authentication performance and ablation studies.

Elevating the Usability of Contactless Perception Systems: From Millimeter-wave to Multimodal Platforms

Thursday, July 2, 2026 - 09:45 am
Room 2277, Storey Innovation building

 DISSERTATION DEFENSE

Author :  Moh. Sabbir Saadat
Advisors: Dr. Sanjib Sur
Date: July 02, 2026
Time: 09:45 Am
Location:  Room 2277, Storey Innovation building

Link:  https://teams.microsoft.com/meet/29819819216139?p=qjSmXRE3Pm9C9PUJ2P

Abstract
Contactless perception systems are an attractive paradigm for applications in security or health monitoring due to their non-intrusiveness and ease of instrumentation compared to wearable or contact-based sensors. This dissertation advances the usability of contactless perception along two connected directions: millimeter-wave (mmWave)-based sensing and imaging, and multimodal digital health assessments.

MmWave signals are high-frequency (24.0 – 300.0 GHz) radio signals that can operate in low visibility, sense through some occlusions, avoid direct camera-like visual capture, and can sense micro-motions that are not possible with conventional vision systems (monocular camera, stereocamera, lidar, IR sensors). Moreover, mmWave signals constitute an important enabling layer in 5Gand-beyond networking paradigms, particularly in indoor networking scenarios. This presents us with an opportunity to bring mmWave sensing for applications where conventional vision systems fail – privacy concerns, low-lighting conditions, requirement of detecting micro-motions, etc. This dissertation addresses a set of fundamental and system-level challenges that hinder the wide adoption of mmWave signals in perception systems: increased and rapid temperature increase due to higher power consumption, information loss resulting from specular reflections, networkingsensing performance trade-offs in integrated system, and motion errors and sparse measurements in hand-held operation.

In the final work of this dissertation, we address the limitation of mmWave system, or any singlemodality perception system, to address perception in more complex digital health applications such as the post-stroke recovery assessment. While mmWave signals capture the fine-grained limb motion abnormality and pose asymmetry in stroke survivors, post-stroke symptoms span across a wider range in terms of scale and functionality: from broad body pose and motor dysfunction to fine-grained limb coordination and facial weakness, to speech and conversational impairments. This dissertation presents a first step towards a multimodal approach to automating post-stroke recovery assessment using vision and audio while retaining mmWave signals as a complementary integration. This work is an interdisciplinary collaboration between our lab and researchers from the University of South Carolina, School of Medicine, paving the way towards research in multimodal digital health sensing on a wider scale.

Understanding Group Emotion by Multi-Stream Deep Neural Networks

Monday, June 22, 2026 - 02:00 pm
online

DISSERTATION DEFENSE
 

Author : Ahmed Shehab Khan
Advisors: Dr. Yan Tong
Date: June 22, 2026
Time: 02:00 pm
Place: Virtual (Zoom)
Link:  https://sc-edu.zoom.us/j/86147086296

Abstract
Group Emotion Recognition (GER) is the task of inferring the collective emotional state of a group of individuals from a single image. The task has several inherent challenges. First, relevant evidence is spread across faces, body language, objects, and the surrounding scene, and no single cue is sufficient. Second, individuals and regions within a group contribute unequally to the perceived emotion, with the relative importance shifting from image to image. Third, existing methods that integrate these signals rely on multi-stream pipelines with separate detectors and networks, making inference computationally expensive. This dissertation develops three deep learning frameworks for GER, each building on the previous to address these challenges.


First, we proposed a four-stream hybrid network that combines features from individual faces, the scene, and the spatial arrangement of faces within the image. A face-location aware stream captures the relationship between faces and scene through an attention heatmap; a multi-scale face stream handles the high variance in face size found in images collected in the wild; and a global blurred stream learns scene-only features by suppressing face appearance. These four streams are combined with hand-engineered fusion weights.

Second, we proposed Regional Attention Networks with Context-aware Fusion. Building on the multi-stream approach, this work addressed two limitations: how to determine the importance of individual persons and objects within a group, and how the relative weight of different streams should depend on the image. A regional attention mechanism estimates the importance of each person or object from the image, and a context-aware fusion module replaces the fixed stream weights of the prior framework with values derived from the image content itself. To reduce computational cost, feature extraction is consolidated onto a single shared backbone.

Third, we proposed LG-GER, a language-guided distillation framework that addresses the inference cost of detector-driven multi-stream pipelines. To overcome the lack of spatial supervision in existing GER datasets, a multimodal large language model serves as an offline annotator that generates dense, spatially grounded emotion evidence (bounding boxes with emotion signals and confidence scores) for the training images. This structured evidence is distilled into a single vision-language backbone through four complementary losses: classification, region-text grounding, spatial emotion, and spatial confidence regression. At inference, the framework requires no detectors, no MLLMs, and no multi-stream fusion, making it the first detector-free framework for group emotion recognition.

All three frameworks were evaluated on publicly available GER benchmarks. Visualization and case studies illustrate how the attention and fusion components identify the most informative regions for group emotion recognition.

Learning and Exploiting Causal Structure for Robust and Transferable Configuration Optimization in Cyber-Physical Systems

Friday, May 22, 2026 - 10:30 am
online

DISSERTATION DEFENSE

Author : Md Abir Hossen
Advisors: Dr. Pooyan Jamshidi
Date: May 22, 2026
Time: 10:30 am
Place: Virtual (Zoom)
Link:  https://sc-edu.zoom.us/j/88165893836?pwd=wVzSgMF42SJFqBHkE7QtSQO1Tsemp4…

Abstract


Cyber-physical robotic systems expose a combinatorially large configuration space comprising interacting hardware and software parameters. Incorrect configurations can lead to functional faults that are difficult to diagnose due to the intricate and often hidden dependencies between system settings and performance. This dissertation addresses these challenges by learning and exploiting causal structure to enable robust, sample-efficient, and transferable configuration optimization.


The first part of this work introduces CaRE (Causal Robotics DEbugging), a causal diagnosis framework that identifies the root causes of observed functional faults. By learning causal relationships between configuration parameters and performance indicators from observational data, CaRE enables precise fault localization and validation through targeted interventions across both simulation and physical robot platforms.

Building on this causal foundation, the next stage introduces CURE (Causal Understanding and Remediation for Enhancing Robot Performance), a configuration optimization method that identifies causally relevant parameters and restricts optimization to a reduced subspace. CURE improves convergence efficiency and supports transfer across environments by leveraging causal knowledge obtained in low-cost simulations and applying it to real-world robot deployments.

The final part of this dissertation introduces RESCUE (REducing  Sampling cost with Causal Understanding and Estimation), which extends the optimization setting from a single-fidelity target environment to multi-fidelity settings where multiple information sources with different costs and accuracies are available. RESCUE uses causal structure to construct a causal prior and guide configuration-fidelity selection, reducing costly high-fidelity evaluations while preserving optimization quality. Empirical evaluation across synthetic and real-world problems shows improved sample efficiency, more effective fidelity allocation, and lower constraint violation rates than competing methods.

Collectively, this dissertation establishes a causal foundation for reliable, efficient, and transferable configuration debugging and optimization in cyber-physical systems, validated through both synthetic benchmarks and real-world robotic applications.