Research practice / Version 1.0

How we separate development results from independent evaluation.

A useful experiment is a step in development. To understand what it proves, we also need to know how its inputs were chosen, what changed and what remains untested.

What did an experiment actually establish?

A model can look convincing on a familiar sequence and behave differently when the scene, camera or background changes. A result is easier to interpret when its evaluation conditions are recorded with it.

Our research practice separates the work used to develop a system from the evidence reserved to assess it. This note describes that distinction. It reports no new performance result.

Keep the roles of data explicit.

Training data helps fit a model. Development data helps the team investigate behaviour and make engineering decisions. Independent evaluation data is held apart so that the same examples do not quietly become both the design target and the final assessment.

These roles apply to complete source sequences where nearby frames are closely related. Calling one frame “development” and an almost identical neighbouring frame “test” can exaggerate how independent the test really is.

Once a sequence has informed a change, it remains useful development material. It should not subsequently be described as untouched evidence of generalisation.

A completed training run is not an accepted improvement.

Finishing training means that a process completed. Acceptance is a separate engineering decision. A candidate needs to meet the declared criteria across the evaluation conditions that matter to its intended role.

Improvement in one condition may accompany regression in another. For that reason, a single attractive aggregate score can hide a material failure. We preserve reference systems, candidate records and the reasons for accepting or rejecting a change.

A failed candidate still teaches us something. The distinction is between learning from it and promoting it into the system.

Keep the evidence categories distinct.

Simulation gives control over a defined model of a scene. Recorded video gives repeatable visual inputs. A physical-camera bench exposes real optics and hardware behaviour. Integrated flight introduces the response of the moving platform.

These categories support each other, but success in one does not establish success in the next. A recorded replay, for example, can expose a loss of visual continuity. Its images cannot change in response to a guidance output because those images were already recorded.

Describe a result with the evidence category that produced it. A bench result remains a bench result until the next test supplies its own evidence.

Make the research record interpretable.

A useful record makes it possible to understand which inputs and system version were used, what question was asked and how the conclusion was reached. We keep source provenance, evaluation roles, experiment configurations and output records alongside the work.

When reporting a measurement, the denominator and operating conditions belong with the number. Timing also needs its context: the computing platform, the workload and the distinction between startup and continued operation.

Corrections belong in the record too. If a scoring rule or interpretation changes, the revised conclusion should explain what changed rather than quietly replacing the earlier account.

Scope and limitations.

This is a description of research practice, not an independent audit, benchmark release or claim of operational reliability. Kodanda’s current work spans simulation, recorded development inputs and physical bench experiments. Integrated autonomous flight validation remains ahead.

The recorded example on our research page uses an external dataset and a borrowed detector baseline. It illustrates a replay format; it is not a representative score for an operational task.

Our proposed Cue-to-Track protocol is still in preparation. Public evaluation materials and measured results require their own complete descriptions and release review.

Related material

Inspect the recorded development example ↗

Read the first-product scope ↗

Dataset source and media credits ↗

Version 1.0 / 14 September 2026. Original methodology note by Kodanda AI. No new benchmark result is reported in this article.

Partnerships

Build the next step
with us.

We welcome conversations with platform builders, test partners and researchers in onboard autonomy.

Discuss a partnership