Nilgün Şengöz
AI Researcher working across Explainable AI, LLM-based optimisation, computer vision, and quantum machine learning.

Research Across High-Stakes Domains
I am an Assistant Professor (Dr. Öğr. Üyesi) in the Department of Information Systems and Technologies at Burdur Mehmet Akif Ersoy University's Gölhisar School of Applied Sciences, currently on leave for postdoctoral research abroad. My research spans AI applications across healthcare, computer vision, defense systems, and combinatorial optimisation.
I am conducting postdoctoral research at the University of Nottingham from February 2026 to February 2027, supported by the TÜBİTAK 2219 International Postdoctoral Research Fellowship, within the Computational Optimisation and Learning Lab.
My research is built on a consistent principle: algorithms that perform well in one domain tend to transfer to others. A visual anomaly detector developed for medical histopathology can be adapted for industrial inspection or defense surveillance. This transferability, combined with interpretability requirements, defines the common thread across my published work.
Ph.D. in Computer Engineering
Specialization in AI & Image Processing
TÜBİTAK 2219 Fellow
International Postdoctoral Research Fellowship
CRC Press Co-Editor
Explainable Artificial Intelligence (XAI) in Healthcare, 2024
Assistant Professor
Burdur Mehmet Akif Ersoy University
Research Areas
My work spans four interconnected areas, united by a common question: how do we build AI systems that are accurate, interpretable, and transferable across high-stakes domains?
Explainable AI (XAI)
Post-hoc and intrinsic interpretability for deep learning systems in regulated environments. Methods include Grad-CAM, SHAP, LIME, and attention visualisation. Published work spans clinical diagnostics, defense AI, and cross-sector compliance. Co-editor of Explainable Artificial Intelligence (XAI) in Healthcare (CRC Press, 2024).
LLM-Based Combinatorial Optimisation
Using large language models as an offline knowledge source for algorithm selection. Developed LLM-STAR, a generation hyper-heuristic that distils GPT-4 reasoning into a deterministic seven-rule engine for examination timetabling, presented at PATAT 2026. The deployed engine runs without the LLM.
Computer Vision for High-Stakes Environments
Attention-based anomaly detection in complex visual data where standard detectors fail. Published research on military camouflaged object detection (ECJSE, 2026). Methods transfer across medical imaging, industrial inspection, and defense surveillance contexts.
Quantum Machine Learning
Trainability analysis for parameterised quantum circuits in the NISQ era, including gradient covariance structure and spectral criteria for barren plateau behaviour. An emerging strand of my research programme alongside classical machine learning.
Contributions to the Scientific Community
Explainable Artificial Intelligence (XAI) in Healthcare
This volume addresses how explainable AI can raise the trustworthiness, performance and sustainability of AI systems deployed in healthcare. Across twelve contributed chapters, it covers XAI techniques, frameworks and evaluation metrics, with applications spanning disease diagnosis, medical image processing, drug discovery, precision medicine and digital twins. The book is written for graduate students, researchers, industry practitioners and clinicians working on high-stakes decision systems.
For a complete list, visit my Google Scholar profile.
Experience & Education
On Leave 2026–27
Talks & Presentations
Conference presentations and invited lectures. Slide decks are available for download where the work has already been presented publicly.
Large Language Model-Based Explainable Hyper-heuristics for Examination Timetabling
University of Nottingham, Jubilee Campus · 25–28 August 2026
Introduces LLM-STAR (Structure-Aware Selector), a generation hyper-heuristic that uses GPT-4 offline to distil selected literature on examination timetabling and graph colouring into a seven-rule engine. The rules map five structural features of a conflict graph to a candidate set of constructive heuristics; once generated, the engine runs without the LLM. Matches pool-best on all 13 Toronto benchmark instances and all 18 generated validation instances. Joint work with Prof. Dr. Ender Özcan and Dr. Jeremie Clos at the Computational Optimisation and Learning Lab.
From Black Box to Glass Box: Explainable AI as the Precondition of Trust in High-Stakes Decisions
London Metropolitan University, London · 13 June 2026
Invited keynote examining why interpretability is becoming a deployment requirement rather than an optional feature, across healthcare, defence and industrial AI systems.
Towards Trustworthy Scholarly Writing with Large Language Models: A Practical Verification Checklist for Reducing AI Hallucinations
Introduces the Scholarly LLM Verification Checklist (SLVC), a protocol for reducing hallucination risk when large language models are used in academic manuscript preparation. The framework sets out verification steps for citations, quantitative claims and attributed statements before a manuscript is submitted.
Recognition & Milestones





Click any photograph to view it full size. Use the arrow keys or on-screen arrows to browse.
🏆 NATCOR Winning Team
NATCOR @ University of Nottingham · EPSRC
Winning Team of the Practical Challenges Competition in "Heuristic Optimisation and Learning", solving problems on complexity theory, heuristics, meta-heuristics, hyper-heuristics and large-scale data analytics. (13–17 April 2026)
🎉 100+ Google Scholar Citations
Google Scholar
Reached the milestone of 100+ citations across publications in deep learning, explainable AI, and medical image processing.
View Profile📄 PATAT 2026 — Paper Presented
International Conference on the Practice and Theory of Automated Timetabling
Paper presented at PATAT 2026, the 15th Conference on the Practice and Theory of Automated Timetabling, held at the University of Nottingham. The work distils large language model reasoning into a deterministic rule engine for heuristic selection in examination timetabling. The extended abstract and slides are available in the Talks section.
TÜBİTAK 2219 Fellowship
TÜBİTAK · The Scientific and Technological Research Council of Turkey
International Postdoctoral Research Fellowship supporting research at the University of Nottingham, UK.
Thoughts, Travels & Discoveries
Sharing my journey through AI research, academic life abroad, and the places I explore along the way.
Binlerce Kilometre Uzakta, Kendi Tarihimle Yüzleşmek — British Museum
Osmanlı eserlerini Londra'da görmenin verdiği o tuhaf his: bir yanda heyecan, öte yanda derin bir hüzün. Çiçekler dahi ait olduğu toprakta büyür...
Londra'da British Museum'a gittiğimde, içimde bambaşka bir his uyandı. Özellikle Osmanlı tarihi bölümüne doğru nerdeyse koşar adımlarla gittim. Sanki mâverâdan bir ses beni müzenin o kısmına çağırıyordu. Cam vitrinlerin arkasında duran her eser; bir çini parçası, bir tuğra, işlemeli bir kaftan, bir ferman... Hepsi benim tarihimden, benim köklerimden bir parçaydı.
İnsan kendi medeniyetinin izlerini binlerce kilometre uzakta, yabancı bir toprağın göbeğinde, yabancı bir dilin rehberliğinde görünce garip bir his yaşıyor. Bir yanda coşku var — "Evet, atalarım böyle eserler bıraktı, bunlar var oldu, hayatta kaldı" diye bir gurur. Öte yanda ise sessiz, ağır bir hüzün oturdu gönlüme.
Çünkü o eserler orada olmamalıydı. Her şey ait olduğu yerde anlam kazanır. Bir çini, İznik'te; bir kaftan, Topkapı'da; bir ferman, yazıldığı toprakta soluduğunda gerçek sesini verir. Çiçekler dahi ait olduğu toprakta büyür, gelişir, kök salar. Söküldüğünde yine çiçektir belki, ama kokusu eksiktir, toprağından uzaktır.
Müzeden çıkarken şunu düşündüm: Bir bilim insanı olarak ben de şu an kendi toprağımdan uzaktayım. TÜBİTAK bursuyla Nottingham'dayım, araştırıyorum, öğreniyorum, kendimi geliştiriyorum. Ama fark şu ki ben geri döneceğim. Topraklarıma, öğrencilerime, bilgimi taşıyacağım ülkeme. O eserler ise ait olmadığı yerde kalacaklar.
British Museum dünya coğrafyasındaki tüm eserleri görebileceğiniz harika bir yer ama atalarımın mirasını olması gerektiği gibi kendi vatanımda yer almaması bizden sonraki nesillere anlatacak hikayelerimizin hep yarım kalmasına sebep olacaktır ne yazık ki...
Explainable AI: Why Transparency Matters in Healthcare
As AI systems become more prevalent in medical diagnosis, the need for transparency and interpretability grows exponentially...
As AI systems become more prevalent in medical diagnosis, the need for transparency and interpretability grows exponentially. In my research, I focus on Explainable AI (XAI) methods that help clinicians understand how algorithms reach their decisions.
Gradient-weighted Class Activation Mapping (Grad-CAM) is one of the key techniques I use to visualize which parts of a histopathological image the model focuses on when making a diagnosis. This visual feedback is crucial for building trust between AI systems and healthcare professionals.
The challenge is not just building accurate models. Doctors must understand them, trust, and ultimately use to improve patient outcomes. This is where XAI bridges the gap between algorithmic power and clinical practice.
In our recent study on paratuberculosis diagnosis, we demonstrated that Grad-CAM heatmaps closely aligned with the regions pathologists identified as diagnostically relevant, validating the model's reasoning process.
LLM-STAR: The Model Reads the Literature and Writes the Rules
Presented at PATAT 2026 in Nottingham and published in the proceedings as an extended abstract. The paper, the slides and a three-minute excerpt of the talk, collected in one place.
Research context: This work was conducted at the University of Nottingham under a TÜBİTAK 2219 international postdoctoral fellowship, within the Computational Optimisation and Learning (COL) Lab. The paper is co-authored with Prof. Dr. Ender Özcan, who supervises the fellowship, and Dr. Jeremie Clos.
LLM-STAR was presented at PATAT 2026, the 15th Conference on the Practice and Theory of Automated Timetabling, held at the University of Nottingham on 25–28 August 2026, and is published in the conference proceedings as an extended abstract. This page collects the paper, the slides and a three-minute excerpt of the talk.


What the paper does
Examination timetabling reduces, in its simplest form, to graph colouring: exams are vertices, shared students are edges, and the fewest clash-free time-slots is the chromatic number. Constructive heuristics such as Welsh-Powell, DSATUR, Smallest-Last, BFS and DFS solve this quickly, but no single one of them dominates across the structures that real conflict graphs take. On the Toronto instance lse-f-91, Smallest-Last needs 18 time-slots where DSATUR needs 19. Choosing the heuristic from the structure of the instance is an algorithm selection problem, and in scheduling practice the selection has to be defensible, so human-readable rules are preferable to an opaque learned model.
LLM-STAR (Structure-Aware Selector) uses GPT-4 once, offline, as a knowledge source. The model is given selected literature on examination timetabling and graph colouring together with the structural features and observed heuristic performance of the 13 Toronto benchmark instances. Its reasoning is filtered by a four-of-five consensus rule and converted, through maximum-margin thresholding, into seven priority-ordered rules over five structural features: edge density, global clustering coefficient, degree variance, maximum degree and instance size. Once the rules exist, the deployed engine contains no language model. Every selection names the rule that fired and the rationale attached to it.
Results and stated limits
The engine matches the pool-best heuristic on all 13 Toronto instances and on all 18 independently generated validation instances, executing at most two of the five heuristics per instance. Against a static Smallest-Last baseline, a two-tailed Wilcoxon signed-rank test gives p = 0.002 with Cohen's d = −1.05. The paper states its limits openly: the rules target the structural subspace of realistic examination conflict graphs and claim no generality beyond it; one rule never activates on either benchmark; and the heavy-tailed rule is calibrated on a single instance, which the paper flags as an overfitting risk.
Paper, slides and talk
The same three-minute excerpt of the talk is available on LinkedIn and on X, so either link works.
Academic correspondence on offline LLM distillation, algorithm selection and explainable hyper-heuristics is welcome.
Military Camouflage Detection: What Attention Mechanisms See When Human Eyes Fail
Standard object detectors fail near chance level on well-camouflaged targets. Attention-augmented architectures change the question from "is there an object?" to "is there a statistical anomaly?"
Research context: This work has been published in El-Cezeri Journal of Science and Engineering (ECJSE), Vol. 13, No. 2, pp. 146–160, 2026. Co-authors: G. Karaman, M. S. Çeliker, N. Y. Çan. DOI: 10.31202/ecjse.1747013
Why Standard Detectors Fail
Camouflage is one of the oldest and most persistent challenges in visual perception. Its effectiveness relies on a straightforward principle: match target appearance to background statistics so thoroughly that the human visual system finds no foothold. Standard object detection architectures are optimised for their dominant training data: objects that are visually distinct from backgrounds, with clear edges, consistent texture, and reasonable contrast. Military camouflage specifically engineers against all three of these properties.
The result: naive application of standard models produces detection rates near chance level on well-camouflaged targets. The problem is not model capacity. It is misaligned inductive bias.
The Attention-Based Approach
Our published work integrates attention mechanisms into established segmentation and classification backbones, specifically Attention U-Net and ResNet-50, to re-frame the detection task. Rather than asking whether there is an object matching a given class, the attention-augmented model is guided toward regions whose properties are inconsistent with the surrounding background.
This re-framing is the core insight. A camouflaged target cannot match the background perfectly; there will always be statistical residuals: micro-scale shadow inconsistencies, subtle texture regularity, edge artefacts at boundary regions.
Attention gates allow the network to suppress irrelevant background activation and concentrate representational capacity on regions where camouflage artefacts are strongest. Full architectural details, dataset description and quantitative results are reported in the published paper.
Attention-Guided Detection · Conceptual Flow
Conceptual illustration of how attention gating balances fine detail against scene context. Simplified for readability; the published paper reports the full architecture.
The Same Problem in Four Other Domains
Explainability
Given my broader research programme in Explainable AI, I paid particular attention to whether the model's attention maps were interpretable by domain experts. Grad-CAM visualisations showed consistent focus on boundary regions and texture discontinuities that human experts also identified as informative. This alignment between model attention and expert judgement is a necessary, though not sufficient, condition for operational trust.
Transferability
The attention-based approach is not domain-specific. The model learns to find statistical inconsistencies in visual data. The same problem appears in medical imaging (subtle lesions in complex tissue), industrial inspection (surface defects that blend with material texture), remote sensing (concealed installations), and infrastructure security (foreign object detection).
The same attention-based approach applies to each of these problems with adaptation. The domain changes what "anomaly" means. The algorithm for finding anomalies remains the same.
On Algorithm Architecture: Why Domain-Agnostic AI Is the More Durable Skill
Looking across research spanning several distinct domains, a pattern emerges: the most transferable contribution is not domain knowledge. It is the ability to design algorithms that generalise.
Looking back across my published work, I notice a pattern I did not plan but can now articulate clearly. My publications span histopathological image classification, military camouflaged object detection, LLM-based combinatorial optimisation, hybrid neural architectures, and agricultural image analysis. On the surface, these seem unrelated. On closer inspection, they share a common structure: each involved designing an algorithm to find structure in data that is not obvious, then explaining why that structure exists.
The Distinction That Matters
Consider two researchers. The first is a "medical imaging AI specialist." Invaluable in that context, but when the clinical application changes, their algorithmic toolkit must adapt to new domain knowledge. The second researcher asks first: what is the structure of this problem at the level of the algorithm? They recognise that detecting a tumour in histopathological tissue and detecting a concealed object in a complex visual background are, algorithmically, the same problem, and apply the same attention-based architecture to both.
The domain changes what "anomaly" means. The algorithm for finding anomalies transfers.
Evidence from My Own Research
CLAHE contrast enhancement, developed for veterinary histopathology, transfers directly to military vision, because both involve low-contrast targets in complex backgrounds. XGBoost hybrid ensembles, developed for rotten fruit detection, apply to medical image classification, because both combine structural and appearance features under class imbalance. Grad-CAM attention visualisation, developed for clinical XAI, transfers to defense AI, because both require operator-interpretable evidence maps for high-stakes decisions.
Each transfer emerged organically, not from planning, but from recognising that two problems with different application contexts were structurally identical at the algorithmic level.
Research Arc · How the Work Developed
Research Profile · Indicative Figures
What Algorithm Architecture Means in Practice
Problem abstraction before domain engagement. Before reading the domain literature, ask: what is the formal structure of this problem? Classification with class imbalance? Detection with low signal-to-noise? Combinatorial problem with hard and soft constraints? The formal structure determines which algorithmic family is relevant. The domain determines the data characteristics and validation criteria.
Method transfer as a first hypothesis. When facing a new problem, the first hypothesis is always: has this algorithmic structure been solved elsewhere? If so, does the solution transfer, and if not, where does it break?
Generalisability as a research contribution. A paper reporting "97% accuracy on dataset X in disease Y" has made a domain contribution. A paper showing "architecture Z outperforms baseline B on low-salience detection across three domains" has made an algorithmic contribution. The second type has a longer citation half-life.
A clarification: this is not an argument against domain expertise. It is an argument that the most productive AI researchers tend to be those whose primary fluency is in the underlying algorithms, engaging domain knowledge as a constraint and validation mechanism rather than a primary lens.
XAI Beyond Healthcare: Explainability as a Cross-Sector Deployment Requirement
As the EU AI Act enters into force and the UAE National AI Strategy 2031 prioritises trustworthy AI, explainability is transitioning from research interest to operational necessity across sectors.
My research in Explainable AI originated in a clinical context: making deep learning diagnostics trustworthy to pathologists. That context remains important; the CRC Press volume I co-edited in 2024 addresses it directly. Over the course of this research, I have become increasingly convinced that the XAI problem is not primarily a healthcare problem. It is a consequence of deploying powerful AI systems in any context where human accountability is required, and as regulatory frameworks mature, that context is becoming essentially everywhere.
The Regulatory Shift
The EU AI Act, which entered into force in 2024, establishes a risk-based framework imposing transparency requirements on AI systems in proportion to potential harm. High-risk applications, including medical diagnosis, credit scoring, employment screening and critical infrastructure, are subject to requirements including human oversight, model behaviour documentation, and the ability to explain individual decisions.
Globally, similar frameworks are emerging. The UAE National AI Strategy 2031 explicitly prioritises trustworthy and ethical AI. NATO AI principles require meaningful human control over autonomous systems. The FDA has issued guidance on AI medical devices. The direction of travel is consistent: AI systems making consequential decisions must be explainable.
Where XAI Requirements Are Most Acute
Healthcare & Medical Devices: EU MDR and FDA guidance require AI transparency in clinical decision support. Diagnosis, triage, and treatment recommendation systems must provide auditable reasoning to clinicians and regulators.
Defence & Autonomous Systems: NATO AI principles and international humanitarian law require meaningful human control over autonomous weapons. Explainable decision trails are a legal prerequisite.
Financial Services: GDPR Article 22 and regional governance frameworks require right-to-explanation for automated credit, insurance, and fraud decisions.
Smart City & Infrastructure: AI systems managing urban resources such as traffic, utilities and emergency response require auditable reasoning for public accountability and operator override capability.
The Methodological Toolkit
Grad-CAM produces spatial heatmaps for image classification tasks, answering where the model was looking. SHAP provides feature attribution for tabular data, grounded in game-theoretic foundations. LIME offers instance-level explanation of any black-box model when global interpretability is impossible. Attention visualisation reveals what transformer-based architectures attended to at each step.
Choosing the right method depends on three factors: the model architecture, the nature of the explanation required (spatial vs. feature vs. instance), and the audience (clinician vs. regulator vs. operator). These methods are not interchangeable.
Global XAI Regulation · Key Milestones
XAI Method Selection · When to Use Which
An Honest Assessment
The most widely used post-hoc explanation methods are approximations. They reveal something true about model behaviour, but they do not provide a complete or guaranteed-faithful account. There is active and legitimate debate in the research community about whether these approximations are reliable enough for high-stakes operational deployment.
This debate is pushing the field toward intrinsically interpretable architectures: models designed for transparency from the outset of the design process. Within five years, XAI will likely be a prerequisite for deployment approval in regulated environments rather than an optional add-on.
I served as a co-editor of Explainable Artificial Intelligence (XAI) in Healthcare (CRC Press, Taylor & Francis, 2024). The volume brings together twelve contributed chapters on deploying XAI methods in clinical and other high-stakes settings, covering disease diagnosis, medical imaging, drug discovery and precision medicine. Academic correspondence on XAI methodology, cross-sector applications, or the regulation-explainability interface is welcome.
Let's Collaborate
I welcome correspondence on research collaboration, joint publications, and academic partnerships in explainable AI, computer vision, and combinatorial optimisation.
🇬🇧 Currently in Nottingham
I'm conducting postdoctoral research at the University of Nottingham as a TÜBİTAK 2219 Fellow until February 2027.