2026/08/19 更新

イマガワ タカヒサ
今川 孝久
IMAGAWA Takahisa
Scopus 論文情報
総論文数: 10 総Citation: 117 h-index: 5

棒グラフ及び折れ線グラフは最大で直近20年分が表示されます。

所属
大学院情報工学研究院 知的システム工学研究系
職名
助教
外部リンク

研究分野

  • 情報通信 / 知能情報学  / 人工知能,強化学習,プランニング

取得学位

  • 東京大学  -  博士(学術)   2018年03月

  • 東京大学  -  修士(学術)   2015年03月

  • 東京大学  -  学士(教養)   2013年03月

学内職務経歴

  • 2024年02月 - 現在   九州工業大学   大学院情報工学研究院   知的システム工学研究系     助教

論文

  • Cost-Aware Embedding Dimension Selection via Accuracy and Training-Cost Surrogate Models 査読有り

    Imagawa, Takahisa and Enokida, Shuichi

    18th International Joint Conference IJCCI 2026   2026年10月

     詳細を見る

    担当区分:筆頭著者, 責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

  • 物体検出器の知識蒸留におけるアノテーションコストを考慮した蒸留手法に関する研究

    横石和真, 今川孝久, 榎田修一

    SSII2026   2026年06月

     詳細を見る

    担当区分:責任著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

  • Exploring Pre-Service Teachers’ Reflection on Nonverbal Behavior in Microteaching Through Three-Point Comparison Feedback 査読有り

    Shirasaka S., Imagawa T., Enokida S.

    Education Sciences   16 ( 5 )   2026年05月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)

    Discrepancies among feedback sources are typically treated as measurement errors, yet they may serve as catalysts for deeper professional reflection. This exploratory, single-group mixed-methods study, conducted in one Japanese teacher education context, examined how three-point comparison feedback (3PCF)—the simultaneous presentation of automated video-based evaluation, peer evaluation, and self-evaluation—relates to pre-service teachers’ reflection on nonverbal teaching behavior in microteaching. Drawing on Hattie and Timperley’s feedback model and the concept of cognitive conflict, 27 participants received 3PCF on multiple nonverbal behaviors and completed written reflections analyzed using an ordinal coding scheme, keyword detection, and text mining. Quantitative analysis revealed that agreement between automated and peer evaluation was strongly item-dependent (e.g., voice volume: r = 0.853; facial expression: r = 0.164). Qualitative analysis showed that discrepancies were associated with multi-layered reflection; as exploratory, prompt-sensitive indicators, keyword detection suggested that 67% of participants recognized gaps between self-perception and external evaluations, 41% reasoned about why sources diverged, and 70% formulated specific behavioral improvement plans. Text mining further identified distinct reflection patterns, suggesting multiple cognitive pathways. These findings, based on a single cohort, suggest that structured comparison across feedback sources can reframe evaluation discrepancies as educational resources associated with reflective and actionable responses in teacher education.

    DOI: 10.3390/educsci16050760

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=105040046374&origin=inward

  • MOTIP における入力時系列長に対する ID マッチング精度の予測と その活用による再学習コスト削減手法

    川上舞, 今川孝久, 榎田修一

    CVIM2026年1月研究会   2026年01月

     詳細を見る

    担当区分:責任著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

  • The double disconnect: why automated facial expression analysis does not predict peer evaluation in microteaching 査読有り

    Shirasaka S., Imagawa T., Enokida S.

    Frontiers in Education   11   2026年01月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)

    Introduction – When automated facial expression recognition (FER) is applied to teaching, two structural barriers—a double disconnect—may prevent its outputs from corresponding to human evaluations: AI-side construct mismatch between discrete emotion categories trained on posed expressions and the multimodal constructs human observers use, and human–side evaluation constraints (halo effects, evaluability, forced distribution rating) that may prevent evaluators from assessing facial expressiveness as a distinct dimension. This study tested whether these disconnects jointly account for the absence of FER–evaluation correspondence in pre-service teacher microteaching. Methods – Twenty-seven Japanese pre-service teachers conducted 3-min microteaching sessions analyzed using DeepFace; all participants evaluated all other presentations across nine dimensions under a forced distribution rating procedure (N = 702 ratings, self-evaluations excluded). Analyses were organized at three evidence levels: a theory-constrained focal test (Level 1), a primary multiplicity-controlled 63-test matrix (Level 2), and post-hoc cluster patterns (Level 3). Results – At Level 1, the primary DeepFace analysis yielded a near-zero focal estimate (r = .002, p = .99). A matched Py-Feat reanalysis under the same 63-test BH-FDR family likewise yielded no FDR-surviving effects (0/63), although its uncorrected focal estimate was positive (β = +0.283, p = .010, p_FDR = .302). At Level 2, no effect survived Benjamini-Hochberg correction. At Level 3, five uncorrected-significant effects formed a vocal cluster (cross-feature corroborated under dynamic-feature reanalysis) and a peak-aggregation-dependent tension cluster (not robust under distributional reformulation). Principal component analysis confirmed strong halo structure (PC1 = 70.8%, Cronbach's α = .94). A target-level rank-aggregated regression showed that singular fits induced by the forced distribution design did not distort the substantive null when conclusions were derived under an estimand aligned with that design (β correlation r = 1.000 with the primary multilevel model across the 63-test matrix). Discussion – The findings characterize informative non-correspondence under this design rather than high-powered exclusion of moderate effects, and weaken several prominent design-based rival explanations for the null. The contribution is to reframe the practical role of FER in teacher education: not as a replacement evaluator, but as an interpretive contrast that exposes what automated and human channels each can and cannot capture.

    DOI: 10.3389/feduc.2026.1835461

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=105046123444&origin=inward

  • Cross Bisimulationに基づく暗黙的模倣学習によるサンプル効率的な強化学習 査読有り

    今川 孝久, 榎田 修一

    人工知能学会全国大会論文集 ( 一般社団法人 人工知能学会 )   JSAI2026 ( 0 )   5M1GS2b02 - 5M1GS2b02   2026年01月

     詳細を見る

    担当区分:筆頭著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    <p>強化学習は様々な応用事例がある有用な手法であるが,一般に大量のデータが必要で,このことは改善すべき課題の一つである.本研究では,必ずしも質の高くない少量の模倣対象データを活用し,強化学習のサンプル効率を改善する手法Cross Bisimulation based Implicit Imitation Learning (CBI2L) を提案する.本研究では,模倣対象となるmentorと強化学習の実行者であるobserverのエージェントそれぞれのマルコフ決定過程に対し,その間での累積報酬差に関する擬距離cross bisimulation metricを定義する.そして,不動点定理によるcross bisimulation metricの一意存在性,mentorとobserverの累積報酬期待値の差とcross bisimulation metricの関係性について理論的に分析し,observerの学習にmentorの累積報酬を用いるCBI2Lの妥当性を示す.さらにはCBI2Lをsoft actor-critic (SAC)に組み込み,PointMaze環境でSACと比べてサンプル効率が改善することを実際に示す.</p>

    DOI: 10.11517/pjsai.jsai2026.0_5m1gs2b02

    CiNii Research

  • Biased Exploration Q-Learning: A Simple Method for Embedding Knowledge into Reinforcement Learning 査読有り

    Imagawa T., Enokida S.

    International Conference on Agents and Artificial Intelligence   2   1242 - 1253   2026年01月

     詳細を見る

    担当区分:筆頭著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

    The cumulative regret of Q-learning, one of the representative methods in reinforcement learning, has been analyzed; however, most of them assume that the agent learns from scratch. If the agent has knowledge about the learning domain, its learning will be more efficient, as suggested in transfer learning research. Therefore, we modify an existing Q-learning method and propose Biased Exploration Q-learning (BEQ), which assumes that the agent can acquire domain knowledge in advance. We analyze BEQ theoretically and clarify the conditions under which pruning of suboptimal actions is possible and show that the upper bound of the regret of BEQ is significantly improved compared to the Q-learning. Furthermore, our experiments show that BEQ outperforms Potential-Based Reward Shaping (PBRS), a common method for introducing domain knowledge. Moreover, we apply BEQ in deep reinforcement learning and demonstrate its effectiveness in Atari games with simple domain knowledge.

    DOI: 10.5220/0014246600004052

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=105041701330&origin=inward

  • Implicit Imitation Learningの応用による強化学習の効率化

    今川孝久, 榎田修一

    IBIS2025   2025年11月

     詳細を見る

    担当区分:筆頭著者, 責任著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

  • Optimization of Masking Selection Ratio in Diffusion Action Segmentation 査読有り

    Ooyama, Yuta and Imagawa, Takahisa and Enokida, Shuichi

    Proceedings of 8th International Symposium on Future Active Safety Technology towards Zero-Traffic Accidents   2025年09月

     詳細を見る

    担当区分:責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

  • Pose Estimation Using Mask Processing for Robustness Against Occlusions in Overlapping People 査読有り

    Ide, Yuki and Imagawa, Takahisa and Enokida, Shuichi

    Proceedings of 8th International Symposium on Future Active Safety Technology towards Zero-Traffic Accidents   2025年09月

     詳細を見る

    担当区分:責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

  • Efficient Embedding Dimension Number Selection Method for Deep Learning Focusing on Loss Variation in the Early Learning Phase 査読有り

    Uchida, Aoto and Imagawa, Takahisa and Enokida Shuichi

    Proceedings of 8th International Symposium on Future Active Safety Technology towards Zero-Traffic Accidents   2025年09月

     詳細を見る

    担当区分:責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

  • 境界重視型損失関数を追加したHQ-SAM の提案とダクトホース内部検査への応用

    久保田愛梨, 相馬康宏, 相馬貴之, 今川孝久, 榎田修一

    FIT2025   2025年09月

     詳細を見る

    記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

  • GroupPose を用いた複数人同時姿勢推定におけるアテンションのマスキングに関する研究

    笹岡亮太, 今川孝久, 榎田修一

    MIRU2025 Extended Abstract集   2025年08月

     詳細を見る

    担当区分:責任著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

  • ハンドクラフト特徴量による深層学習ベースの三次元物体検出器の性能向上

    平川絢士, 今川孝久, 榎田修一

    DIA2025講演論文集   2025年03月

     詳細を見る

    記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    Kyutacar

  • 事前知識を用いたQ学習

    今川孝久,榎田修一

    IBISML研究会論文集   2024年12月

     詳細を見る

    担当区分:筆頭著者, 責任著者   記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    Kyutacar

  • 深層学習における学習序盤の損失変化に注目した効率的な次元数選択手法(共著)

    内田青澄,今川孝久,榎田修一

    FIT2024 第23回情報科学技術フォーラム講演論文集   2024年09月

     詳細を見る

    記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    Kyutacar

  • 人物の重なりに対する頑健性向上のためのマスク処理を用いた姿勢推定

    井手宥希,今川孝久,榎田修一

    MIRU2024 Extended Abstract集   2024年08月

     詳細を見る

    記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    Kyutacar

  • Diffusion Action Segmentationにおけるフレーム特徴量へのマスキングに関する研究

    大山佑太,今川孝久,榎田修一

    MIRU2024 Extended Abstract集   2024年08月

     詳細を見る

    記述言語:日本語   掲載種別:研究論文(研究会,シンポジウム資料等)

    Kyutacar

  • DROPOUT Q-FUNCTIONS FOR DOUBLY EFFICIENT REINFORCEMENT LEARNING 査読有り 国際誌

    Hiraoka T., Imagawa T., Hashimoto T., Onishi T., Tsuruoka Y.

    ICLR 2022 - 10th International Conference on Learning Representations   2022年01月

     詳細を見る

    担当区分:最終著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

    Randomized ensembled double Q-learning (REDQ) (Chen et al., 2021b) has recently achieved state-of-the-art sample efficiency on continuous-action reinforcement learning benchmarks. This superior sample efficiency is made possible by using a large Q-function ensemble. However, REDQ is much less computationally efficient than non-ensemble counterparts such as Soft Actor-Critic (SAC) (Haarnoja et al., 2018a). To make REDQ more computationally efficient, we propose a method of improving computational efficiency called DroQ, which is a variant of REDQ that uses a small ensemble of dropout Q-functions. Our dropout Q-functions are simple Q-functions equipped with dropout connection and layer normalization. Despite its simplicity of implementation, our experimental results indicate that DroQ is doubly (sample and computationally) efficient. It achieved comparable sample efficiency with REDQ, much better computational efficiency than REDQ, and comparable computational efficiency with that of SAC.

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=85146871335&origin=inward

  • Off-Policy Meta-Reinforcement Learning with Belief-Based Task Inference 査読有り 国際誌

    Imagawa T., Hiraoka T., Tsuruoka Y.

    IEEE Access   10   49494 - 49507   2022年01月

     詳細を見る

    担当区分:筆頭著者   記述言語:英語   掲載種別:研究論文(学術雑誌)

    Meta-reinforcement learning (RL) addresses the problem of sample inefficiency in deep RL by using experience obtained in past tasks for solving a new task. However, most existing meta-RL methods require partially or fully on-policy data, which hinders the improvement of sample efficiency. To alleviate this problem, we propose a novel off-policy meta-RL method, embedding learning and uncertainty evaluation (ELUE). An ELUE agent is characterized by the learning of what we call a task embedding space, an embedding space for representing the features of tasks. The agent learns a belief model over the task embedding space and trains a belief-conditional policy and Q-function. The belief model is designed to be agnostic to the order in which task information is obtained, thereby reducing the difficulty of task embedding learning. For a new task, the ELUE agent collects data by the pretrained policy, and updates its belief on the basis of the belief model. Thanks to the belief update, the performance of the agent improves with a small amount of data. In addition, the agent updates the parameters of its policy and Q-function so that it can adjust the pretrained relationships when there are enough data. We demonstrate that ELUE outperforms state-of-the-art meta RL methods through experiments on meta-RL benchmarks.

    DOI: 10.1109/ACCESS.2022.3170582

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=85129665822&origin=inward

  • Meta-Model-Based Meta-Policy Optimization 査読有り 国際誌

    Hiraoka T., Imagawa T., Tangkaratt V., Osa T., Onishi T., Tsuruoka Y.

    Proceedings of Machine Learning Research   157   129 - 144   2021年01月

     詳細を見る

    担当区分:責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

    Model-based meta-reinforcement learning (RL) methods have recently been shown to be a promising approach to improving the sample efficiency of RL in multi-task settings. However, the theoretical understanding of those methods is yet to be established, and there is currently no theoretical guarantee of their performance in a real-world environment. In this paper, we analyze the performance guarantee of model-based meta-RL methods by extending the theorems proposed by Janner et al. (2019). On the basis of our theoretical results, we propose Meta-Model-Based Meta-Policy Optimization (M3PO), a model-based meta-RL method with a performance guarantee. We demonstrate that M3PO outperforms existing meta-RL methods in continuous-control benchmarks.

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=85137101456&origin=inward

  • Learning robust options by conditional value at risk optimization 査読有り 国際誌

    Hiraoka T., Imagawa T., Mori T., Onishi T., Tsuruoka Y.

    Advances in Neural Information Processing Systems   32   2019年01月

     詳細を見る

    担当区分:責任著者   記述言語:英語   掲載種別:研究論文(国際会議プロシーディングス)

    Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters, these methods only consider either the worst case or the average (ordinary) case for learning options. This limited consideration of the cases often produces options that do not work well in the unconsidered case. In this paper, we propose a conditional value at risk (CVaR)-based method to learn options that work well in both the average and worst cases. We extend the CVaR-based policy gradient method proposed by Chow and Ghavamzadeh (2014) to deal with robust Markov decision processes and then apply the extended method to learning robust options. We conduct experiments to evaluate our method in multi-joint robot control tasks (HopperIceBlock, Half-Cheetah, and Walker2D). Experimental results show that our method produces options that 1) give better worst-case performance than the options learned only to minimize the average-case loss, and 2) give better average-case performance than the options learned only to minimize the worst-case loss.

    Scopus

    その他リンク: https://www.scopus.com/inward/record.uri?partnerID=HzOxMe3b&scp=85090171999&origin=inward

▼全件表示

口頭発表・ポスター発表等

  • 事前知識を用いたQ学習

    今川孝久

    第55回 IBISML研究会  2024年12月  電子情報通信学会

     詳細を見る

    開催期間: 2025年12月20日 - 2025年12月21日   記述言語:日本語   開催地:北海道大学   国名:日本国  

科研費獲得実績

  • 他者データを学習のバイアスとして活用することによる強化学習の効率化

    研究課題番号:25K21153   2025年04月 - 現在   若手研究

     詳細を見る

    強化学習は学習者が試行錯誤を繰り返しその結果の良し悪しをもとに学習する方法で,広い応用先を持つ.
    しかし,多くの場合,無数の試行錯誤が必要となり,このことが実応用を妨げる一因となっている.
    本研究では,強化学習をする学習者の他に,他者が存在する状況を想定する.
    その場合,他者が学習者にとって有益な行動をとれば,学習者の学習が大きく進展する可能性がある.
    本研究では他者からの学習を実現するための基礎となる手法を考案する.

担当授業科目(学内)

  • 2025年度   知的システム工学実験演習Ⅱ

  • 2025年度   システム制御基礎

  • 2025年度   ロボティクス基礎

  • 2025年度   知的システム工学実験演習Ⅰ

  • 2024年度   知的システム工学実験演習Ⅱ

  • 2024年度   知的システム工学実験演習Ⅰ

▼全件表示

担当経験のある授業科目(学外)

  • 深層強化学習総合演習

    2025年02月 - 現在   機関名:深層強化学習総合演習

     詳細を見る

    科目区分:その他  国名:日本国

    KyutechARISEにて社会人向けに強化学習の講義,演習の授業を提供中.

FD活動への参加

  • 2024年12月   新任FD研修2024 12月

  • 2024年11月   新任FD研修2024 11月

  • 2024年10月   新任FD研修2024 10月

  • 2024年07月   新任FD研修2024 7月

  • 2024年06月   新任FD研修2024 6月

  • 2024年05月   新任FD研修2024 5月

  • 2024年04月   新任FD研修2024 4月

▼全件表示