Thanks, this is really helpful. I have created a separate GitHub repository with an initial structure focused specifically on CDSS evaluation, including folders set up for clinical scenario templates, OSCE-style scoring, SpaceRxQA-style medication/countermeasure reasoning, safety and never-event testing, DAG-based causal reasoning evaluation, and required expertise/data sources. I’ll look at compiling a small pilot set before scaling towards a larger validated scenario bank. The repository is here
Hi! I would love to contribute.
Hi,
I just came across this paper which proposes a shift from treating perception as passive input processing toward viewing perception itself as an active reasoning process. Rather than simply consuming multimodal data, the system learns to determine what additional information should be gathered, where attention should be directed, and how perception and reasoning can operate in a closed loop.
Potential relevance to NASA AI/ML efforts:
• Autonomous science agents and decision-support systems
• Human-AI teaming in uncertain or information-constrained environments
• Multi-sensor fusion for Earth and space observations
• Long-duration mission planning where information acquisition is itself a decision variable
One interesting question for the community: How might active perception architectures improve scientific discovery workflows compared to current retrieval-augmented or tool-augmented agent approaches?
Paper: Native Active Perception as Reasoning for Omni-Modal Understanding
Hi all
Pelase join the meeting ![]()
Hi all , @AWGall
we talked about this framework to make CDSS , please read it here if you missed today’s meeting :
CDSS Group: Proposed Research-to-Product Plan
Main objective
Our primary focus is developing a useful CDSS product. Research is the pathway that allows us to move from an evidence-based concept to validated software:
Research → Validation → Development → Product
Why focus on SANS?
SANS is a practical starting point because:
-
We already have experience working with the RR9 dataset.
-
We have strong ophthalmologists as clinical partners.
-
We have relevant clinical knowledge, data experience and an existing DAG foundation.
First research paper
The project will have two main pillars:
-
Expert-based ground-truth labels
-
CDSS software development and evaluation
We currently have access to two datasets, but they are primarily multi-omics datasets intended for research. The NASA Twins Study and Inspiration4 provide useful scientific and longitudinal inspiration, but they are less directly suitable for evaluating a clinical decision-support system.
Therefore, we need to create a dedicated, expert-labelled SANS-CDSS benchmark.
Proposed CDSS workflow
1. Define the scope
-
Main clinical concerns
-
Intended users and setting
-
Questions the CDSS should answer
-
Required inputs and expected outputs
2. Create longitudinal astronaut profiles
Proposed assessment points:
-
L−30: primary baseline
-
L−15: preflight follow-up
-
L: launch
-
L+45: inflight assessment
-
L+90: later inflight assessment
These profiles will initially be created and managed in Excel.
Potential variables include:
-
Age and sex
-
Relevant past medical history
-
OCT measurements, including RNFL and choroidal thickness
-
Fundus findings, including vascularity index and optic-disc measurements
-
Visual symptoms and functional findings
-
Heart rate and other relevant physiological or mission variables
-
Measurement quality and missing-data status
3. Establish the expert reference standard
For each profile and timepoint, experts should determine:
-
Is the available data reliable and of sufficient quality for interpretation?
-
Are there any missing or inconsistent data that limit assessment?
-
Is the finding clinically concerning in the context of SANS?
-
Is the condition stable, progressing, improving or uncertain?
-
Is additional examination, imaging or measurement required?
-
What is the recommended next step or management action?
-
Is escalation needed (e.g., notify the Crew Medical Officer)?
-
Is immediate or emergency action required?
-
Should the system abstain due to insufficient or low-quality data?
-
What is the level of confidence in this assessment?
The expert framework should therefore cover data quality, clinical interpretation, recommended action, escalation, uncertainty and confidence.
4. Develop four comparison systems
-
Rule-based CDSS
-
Bayesian-only CDSS using the DAG
-
LLM/RAG-based CDSS
-
Hybrid SANS-CDSS combining rules, Bayesian inference and evidence retrieval
5. Evaluate the systems
Profile and timepoint data → System answer → Comparison with expert ground truth
Every system should receive the same information and produce the same structured output format.
6. Evaluate the outcomes
-
Correct versus incorrect decisions (with clear definitions of what constitutes a correct clinical action)
-
Certainty and uncertainty (including calibration of confidence scores)
-
Appropriate abstention (when the system correctly identifies insufficient data)
-
Robustness to missing or poor-quality data (including sensitivity analyses)
-
Detection delay (time to identify clinically relevant changes)
-
False-alert rate (frequency of unnecessary alerts)
-
Unsafe-action rate (instances where recommendations could lead to harm)
-
Agreement with expert recommendations (including inter-expert variability where applicable)
The first paper will demonstrate whether the hybrid SANS-CDSS provides safer, more accurate and more explainable decisions than the three individual approaches.
Hello, is it still too late to join and contribute?
Hi @smaransure
It’s not too late to join the team at all! We develop our plan and make it better in every meeting, and there’s a lot more engagement needed.
Sounds good AliReza. I’d be glad to take ownership over any tasks. I can send over a resume.
Sure, please message me and we can find suitable project tasks for you ![]()
Sent it over!
Hi @AliReza-H . Sorry to be a nuisance. Are there minutes of the meeting I could look at. I get spammed with multiple copies of the meetings from otter and also read ai but I am adverse to opening any of these. Kind regards wayne
Hi, this looks really exciting to work on. I would like to contribute.
Dear @Wilester2025 , thank you so much for asking this , i will make sure always sending notes here for all reasons you mentioned .
CDSS Subgroup Meeting Summary — 17 July 2026
1. Research-to-product strategy
Alireza proposed that the group’s primary objective should be development of a useful and validated CDSS product, with research serving as the pathway:
Research → Validation → Development → Product
SANS was proposed as the initial clinical use case because the group already has experience with the RR9 dataset, ophthalmology collaborators, relevant clinical and data expertise, and an existing DAG foundation.
The first project would combine two pillars:
-
Expert-defined ground-truth labels.
-
Development and comparative evaluation of CDSS approaches.
Existing resources such as the NASA Twins Study and Inspiration4 provide useful longitudinal and multi-omics context but are not sufficient for direct CDSS evaluation. A dedicated expert-labelled SANS-CDSS benchmark is therefore needed.
2. Proposed SANS-CDSS benchmark
The group should first define the intended users, clinical setting, principal SANS concerns, required inputs, expected outputs, and questions the system must answer.
Longitudinal astronaut profiles would then be created at representative timepoints:
-
L−30: primary baseline
-
L−15: preflight follow-up
-
L: launch
-
L+45: inflight assessment
-
L+90: later inflight assessment
Initial profiles could be structured in Excel and include demographics, medical history, visual symptoms, OCT measurements such as RNFL and choroidal thickness, fundus biomarkers, optic-disc measurements, physiological variables, and indicators of missing or poor-quality data.
Clinical experts would establish the reference standard by assessing:
-
Data quality, completeness, and reliability.
-
Whether findings are concerning for SANS.
-
Stability, progression, improvement, or uncertainty.
-
Need for further testing.
-
Recommended management or escalation to the Crew Medical Officer.
-
Need for immediate or emergency action.
-
Whether the system should abstain because of insufficient information.
-
Confidence in each assessment.
Four approaches were proposed for comparison:
-
Rule-based CDSS.
-
Bayesian CDSS using the DAG.
-
LLM/RAG-based CDSS.
-
Hybrid CDSS integrating rules, Bayesian inference, and evidence retrieval.
All systems would receive identical profile data and return the same structured outputs. Their recommendations would be compared with expert ground truth using accuracy, confidence calibration, appropriate abstention, robustness to missing or low-quality data, detection delay, false-alert rate, unsafe-action rate, and agreement with experts.
The planned first paper would test whether the hybrid approach produces safer, more accurate, and more explainable decisions than the individual approaches.
3. Psychological and cognitive components
The group identified astronaut psychological health as an important but insufficiently addressed component of the CDSS. Existing work has focused more heavily on cognitive changes during and after flight, while psychological assessment and decision support remain less developed.
Nic expressed interest in contributing to this area and has been reviewing astronaut-health materials and work from the Brain Group. A focused project or discussion on integrating psychological and cognitive risks into the CDSS was proposed.
4. Sam’s framework and potential collaboration
Jian introduced Sam’s GitHub repository on conversational AI and longitudinal CDSS evaluation. Its focus appears to include evidence-grounded, mission-appropriate decision-making over time, DAG-based reasoning, and synthetic-data generation.
Alireza clarified that Sam had shared the repository in response to the subgroup’s multilayer CDSS framework and is already connected to the project, although not yet actively involved. Ritika noted that Sam had previously participated in a UI/UX meeting and expressed interest in CDSS work.
The group proposed inviting Sam to present her framework, clarify its evaluation strategy, and determine how it could be integrated with the SANS-CDSS plan.
5. Multi-agent simulation
Jian proposed extending the CDSS into a multi-agent simulation in which different agents represent crew members with distinct:
-
Mission roles and professional skills.
-
Medical and psychological histories.
-
Physical and behavioral characteristics.
-
Domain-specific knowledge and reasoning capabilities.
The discussion emphasized that an agent’s specialized skills, prompts, knowledge sources, reasoning structures, DAGs, and retrieval systems may be more important than the underlying language model alone.
This environment could generate longitudinal synthetic scenarios that are scarce or absent from existing mission datasets, including radiation illness, limited medical resources, missing data, psychological deterioration, interpersonal conflict, resource scarcity, and ethical emergencies.
Such scenarios could complement the structured SANS benchmark by testing CDSS behavior under rare, extreme, and evolving conditions. Their value would be in evaluating system robustness, safety, escalation decisions, uncertainty management, and performance over time.
6. Meeting communication
Because attendance was limited, the group discussed improving meeting communication through:
-
A dedicated CDSS email address.
-
Reminders approximately one week and two to three days before meetings.
-
Invitations containing a short recap, agenda, and requested preparation.
-
Investigation of emails being filtered outside primary inboxes.
Ritika will investigate the dedicated email and reminder process. Alireza offered to help improve the invitation format.
Proposed next steps
-
Finalize the scope and structured outputs of the SANS-CDSS.
-
Create the initial longitudinal-profile template.
-
Develop the expert-labelling and consensus framework.
-
Define the four comparison systems and common evaluation protocol.
-
Invite Sam to present and align her framework with the project.
-
Develop the psychological and cognitive component with Nic.
-
Explore multi-agent scenarios as a complementary source of synthetic evaluation cases.
-
Improve the meeting invitation and reminder process.
Wonderful !
Thank you so much AliReza. Would a proposed SANS CDSS be able to proactively trend across all crew members as well at the same time , perhaps in the case of something like very slowly increasing co2 levels over time
Hi all — I was sorry to miss the meeting while traveling, but I’ve caught up on the proposed Research → Validation → Development → Product framework and the SANS focus. I’m very aligned with this direction, particularly the expert-grounded benchmark and comparative evaluation of rule-based, Bayesian, LLM/RAG, and hybrid approaches.
As a reminder, I also set up the initial GitHub repo for the CDSS work to give us a shared space for architecture, governance, RAG/reasoning, safety, and documentation: GitHub - aka79/NASA-CDSS-for-Long-Duration-Spaceflight: Clinical Decision Support System (CDSS) for Long-Duration Spaceflight · GitHub . Happy to align it with the current SANS plan and contribute where most useful.
Sure @ayse that would be great!
Thanks for sharing this, @ayse. I really like the structured Research → Validation → Development → Product roadmap.
One thing that came to mind while reading the benchmark section is that it might be useful to define a standardized output schema for expert annotations from the beginning (for example: structured fields for data quality, clinical assessment, recommended action, confidence, uncertainty, and rationale). Having a consistent annotation format should make it easier to compare the rule-based, Bayesian, LLM/RAG, and hybrid systems fairly during evaluation.
It may also be useful to log which evidence each system relied on for every recommendation (rules fired, Bayesian path, retrieved evidence, or LLM reasoning summary) so explainability can be evaluated alongside accuracy.
Looking forward to contributing to this. Here’s my mail hitaeshi25sehgal@gmail.com
Thanks for providing meeting summaries here, it’s really helpful. I will be able to attend the CDSS meeting this Friday and can discuss the github/CDSS evaluation then.
Hi, @Wilester2025
I think we are talking about prediction based on the current and previous state/trends, rather than detection. I also think that every variable measured through sensors and entered into the system as a digital input can potentially be used by the CDSS.
I’ll make sure to include this variable when creating the possible scenarios. Thanks for the suggestion. ![]()