INTOSYN AI Lab — AI Research Division

Towards Trustworthy and Controllable AGI

We investigate how intelligent systems make decisions, when they can be trusted, and how they can be tested and corrected. Five active projects turn these questions into testable hypotheses and experiments.

5
Active projects
Theory × Experiments
From questions to testable hypotheses
GPU
Compute for research
In dialogue with these research communities
  • Nature
  • Science
  • NeurIPS
  • ICML
  • ICLR
  • CVPR
  • ICCV
  • ECCV
  • ACL
  • EMNLP
  • AAAI
  • IJCAI
OUR APPROACH

Clear questions. Rigorous research.

Start with testable hypotheses

Break trust and control into specific questions, with explicit assumptions and evidence that could challenge the hypothesis.

Test against meaningful alternatives

Design evaluations around baselines, controlled comparisons and failure cases. Distinguish observations from hypotheses and their limits.

Connect theory to working systems

Use agents, financial research, sequence generation, cryo-EM and CAD as experimental settings for testing methods on concrete tasks.

ONGOING RESEARCH

Five projects. Five open questions.

From model switching to scientific validation, one question connects our work: how can we understand intelligent systems and intervene on the basis of evidence?

CommitGate

Agent / LLMIn progress
State AbstractionCounterfactual EvaluationModel Routing

Does progress survive a model switch?

Switching models can change an agent’s next actions without necessarily changing the task outcome. CommitGate investigates which aspects of task state determine this difference, and whether a full rerun is needed to predict it.

Approach & evaluation
Research approach

Study state abstractions beyond raw action logs and compare continuations under controlled model substitutions to characterize outcome equivalence between trajectories.

Evaluation focus

Compare predictions with full reruns across tasks, switching points and model pairs; identify where predicted equivalence holds and where it fails.

BITBUDGET

Adaptive InferenceIn progress
Adaptive Data AnalysisGeneralizationInformation Budget

How much validation information does search consume?

Adaptive search repeatedly uses feedback from the same data to refine hypotheses. Trial counts alone may not capture that dependence. BITBUDGET studies how information drawn from validation data relates to out-of-sample performance.

Approach & evaluation
Research approach

Characterize validation-specific information absorbed by the final hypothesis. Use financial research as a testbed to compare information budgets across feedback mechanisms and search policies.

Evaluation focus

Test whether budget estimates explain generalization differences between search policies and predict performance decay across periods and markets.

RONDO

AlignmentIn progress
Credit AssignmentWeak SupervisionCounterfactual Intervention

Can overall preferences reveal local contributions?

An overall score does not directly reveal which part of a long sequence affected its quality. RONDO studies temporal credit assignment through preference comparisons and local interventions, without dense segment-level labels.

Approach & evaluation
Research approach

Compare original sequences with targeted edits. Begin with musical structure and study the assumptions under which local contributions become identifiable.

Evaluation focus

Check localization against independent structural information, test whether local repairs improve overall preference, and examine conditions for transfer to other long-sequence tasks.

CryoMender

AI for ScienceIn progress
Surrogate ModelingCalibrationCryo-EM

When can a low-cost evaluator be trusted?

CryoMender studies local repairs to atomic models built from cryo-EM data. It explores whether surrogate models can screen candidates, reduce expensive validation calls and identify cases that need closer verification.

Approach & evaluation
Research approach

Combine candidate ranking, uncertainty estimates and calibration to allocate a limited physical and geometric validation budget to uncertain repairs.

Evaluation focus

Evaluate validation cost, repair quality and false acceptance risk together, with particular attention to resolution changes and distribution shift.

ReproWorld

World ModelIn progress
Causal LocalizationExecutable ProgramsMinimal Repair

Can a failed build reveal its first faulty step?

Early errors in a parametric CAD build can propagate through later dependencies. ReproWorld studies how to trace downstream symptoms to causal errors and restore execution by changing as few operations as possible.

Approach & evaluation
Research approach

Treat construction histories as executable programs. Use step-level interventions to investigate error propagation, causal localization and dependency-aware minimal repair.

Evaluation focus

Compare with full regeneration using execution cost, edit scope, restored correctness and preservation of constraints and editability.

Infrastructure

Compute for reproducible experiments

Support training, simulation and evaluation in consistent computing environments, so configurations, runs and results can be compared and reviewed.

GPU compute

Compute resources for model training, inference and research evaluation.

Distributed experiments

Interconnected nodes for parallel workloads and multi-node training.

Data and experiment artifacts

Storage for datasets, checkpoints and run records.

Research tooling

PyTorch, JAX and CUDA tooling, with environments organized around each project.

Illustrative research computing workflow

Start with a clear question. Build the evidence together.

Choose a project that interests you and share relevant experience or independent work. Let’s discuss the research question, possible approaches and the time you can commit.