Read-only review fixture

← Review index

Manuscript

RAG grounding & fintech adoption — Can retrieval-grounded drafting reduce unsupported claims in empirical manuscripts?

Abstract

Scholarly drafting assistants increasingly rely on retrieval-augmented generation (RAG), yet authors still ship unsupported claims when passage-level audits are skipped (Gao et al., 2024; Min et al., 2023). We evaluate a verify-first project workflow that keeps literature synthesis, empirical readouts, and manuscript sections on one traceable spine, using Southeast Asian fintech adoption (N = 420) as a substantive illustration. Trust (β = 0.34, p < .001) and perceived usefulness (β = 0.21, p = .003) predict behavioral intention; perceived ease of use is positive but marginal (β = 0.11, p = .068). The model explains 41.8% of variance (adjusted R2 = .409). The contribution is methodological: grounded drafting should be judged by whether claims remain tethered to attached sources and empirical outputs—not by fluency alone.

Introduction

Large language models compress heterogeneous sources into fluent prose faster than most scholars can type, which makes the drafting bottleneck feel solved long before the evidence bottleneck is (Gao et al., 2024). Retrieval-augmented generation (RAG) improves coverage by injecting passages at generation time, but coverage is not the same as accountability: a draft can cite the right paper while misstating what the paper actually shows (Min et al., 2023). Hallucination research therefore frames the realistic target as mitigation rather than elimination—systems should surface uncertainty, preserve trace, and invite author verification before claims harden into manuscript sections (Ji et al., 2023). That framing matters for empirical work in particular, where a single misaligned coefficient description can invalidate an entire results paragraph. Existing tools often split the workflow. Literature lives in one surface, regressions in another, and the manuscript in a third. The author becomes the integration layer, manually copying numbers forward and hoping the introduction still matches the results section after the third rewrite. We argue that pre-submission review quality depends on…

Literature Review

Retrieval-augmented drafting RAG systems combine parametric model weights with non-parametric retrieval over a document index (Lewis et al., 2020). Surveys emphasize retrieval quality, chunking, and faithfulness metrics, noting that downstream authoring still requires author judgment about which retrieved passages license which sentences (Gao et al., 2024). Atomic evaluation frameworks such as FActScore decompose long outputs into checkable facts and score precision against trusted sources (Min et al., 2023). Hallucination and mitigation Ji et al. (2023) catalog hallucination modes in natural language generation and show that retrieval and verification reduce but do not remove unsupported statements—especially when models synthesize across domains. For drafting assistants, the implication is that UI copy should avoid promising elimination and instead train authors to expect audit loops. Project-spine integration Few systems connect empirical estimation outputs to manuscript sections with stable identifiers. Authors routinely bridge results by hand, which introduces transcription risk (wrong sign, wrong decimal, stale R2). A project spine stores analysis configuration, result JSON, …

Methods

Sample and procedure We analyze the fintech_sea_survey.csv dataset attached to this project (N = 420). Respondents completed multi-item scales for behavioral intention (BI), trust, perceived usefulness (PU), and perceived ease of use (PEOU), along with demographic fields (age, gender). Participation was voluntary and responses were anonymized before analysis. Measures and estimation All constructs were measured with Likert-type items aggregated to scale means. We estimate ordinary least squares (OLS) regression with BI as the dependent variable and trust, PU, and PEOU as simultaneous predictors. Coefficients are reported with heteroskedasticity-robust standard errors (HC1). Assumption checks Residuals were inspected for severe non-normality and heteroskedasticity patterns. No imputation was applied; listwise deletion retained the full N = 420. While OLS on Likert means is common in IS adoption research, we note the limitation that treating ordinal items as continuous can understate uncertainty. Workflow instrumentation Separately from the statistical model, we log project decision traces whenever empirical outputs are bridged into manuscript sections—capturing section id, before/af…

Results

Table 1 summarizes the primary OLS model predicting behavioral intention. Trust is the strongest predictor (β = 0.34, SE = 0.08, t = 4.25, p < .001), followed by perceived usefulness (β = 0.21, SE = 0.07, t = 3.00, p = .003). Perceived ease of use is positive but does not reach conventional significance at α = .05 (β = 0.11, SE = 0.06, t = 1.83, p = .068). PredictorβSEtp (Intercept)0.820.312.64.009 Trust0.340.084.25<.001 Perceived usefulness (PU)0.210.073.00.003 Perceived ease of use (PEOU)0.110.061.83.068 Note. OLS with HC1 robust standard errors; N = 420. Model fit: R2 = .418, adjusted R2 = .409, F(3, 416) = 28.6, p < .001. Interpretation for drafting Authors should describe ease-of-use as directionally positive but not definitive in this sample. Trust and usefulness support stronger language (conditional on the cross-sectional design). Any manuscript bridge from the empirics lane should preserve the p-value for PEOU honestly rather than rounding up to “significant.”

Discussion

The empirical illustration reproduces a familiar adoption pattern: trust and perceived usefulness dominate behavioral intention, while ease of use plays a secondary role (Fonseca, 2013). For the methodological contribution, the important point is not novelty of the coefficients but whether the manuscript text stays aligned with them after iterative drafting—an alignment problem RAG tools exacerbate when they paraphrase results too aggressively (Gao et al., 2024). A verify-first workflow changes the author’s task from “generate prose” to adjudicate claims. When regression output and section text share a project id, bridging operations become diffable: reviewers can see whether β = 0.11 became “strongly drives” in the discussion. Design implications Bind empirical JSON to manuscript sections. Expose open comment and suggestion queues before export. Default to mitigation language in AI-generated summaries. Limitations The survey is cross-sectional; coefficients are associational, not causal. The sample is regional; generalization requires new data, not richer paraphrase. Pre-submission checklist Before export, authors should resolve all open passage comments, reconcile results and dis…

Conclusion

Retrieval-augmented drafting accelerates manuscript assembly but does not absolve authors of evidentiary duty. FActScore-style atomic checks and project-native empirical bridges point toward the same practice: keep claims small, tied to passages, and revisitable (Min et al., 2023). In the fintech adoption illustration, trust and perceived usefulness predict behavioral intention with moderate model fit; ease of use remains suggestive rather than definitive. Those substantive findings are secondary to the workflow thesis—submission-quality drafts emerge when literature, empirics, and prose share one inspectable spine. We invite tool builders to optimize for pre-submission review, not first-draft fluency. The cost of an extra audit loop is smaller than the cost of desk rejection for unsupported empirical language.

Read-only. No forms, no mutation controls, no write APIs on this mirror.

Explore a demo workspace·Demo data — not your account