1. 83df861 Internal change by Google AI Edge · 4 hours ago upstream/main
  2. b319bde Clean up `runner_test_suite.h` by Quentin Khan · 7 hours ago
  3. 31c0eca Allow the YNNPACK delegate to rebuild its subgraph outside of delegation without invoke GetNodeAndRegistration directly by Google AI Edge · 8 hours ago
  4. fe837be Support Gemma 4 26B INT4 MoE convert-path model on Metal GPU delegate. by Fengwu Yao · 8 hours ago
  5. bbb1312 Rename `LiteRtStaticLinkedAcceleratorCpuDef` to `LiteRtStaticLinkedAcceleratorXnnpackDef` and remove `__attribute__((weak))` from `RegisterCpuAccelerator`. by Gerardo Carranza · 11 hours ago
  6. 4c60cf4 Update CODEOWNERS file to include a url link pointing to github team by Google AI Edge · 11 hours ago
  7. 4018aed Merge pull request #10448 from graham0824:dev/hungjuiw/fix-addn-test2 by Copybara-Service · 11 hours ago
  8. 1052937 Internal change by Maria Lyubimtseva · 11 hours ago
  9. b0e378a Add a LiteRT cpu_backend build flag and a minimal build guide by Terry Heo · 11 hours ago
  10. dd07838 Implement WebGPU shape/indexing, spatial conv/pool, and LSTM operations and enable GPU NumericalTestSuite. by Ping Yu · 12 hours ago
  11. 97fe163 Merge pull request #10270 from gunes-arm:pr/naming-changes by Copybara-Service · 12 hours ago
  12. 0dedf30 Expose delegation metrics through LiteRT C and C++ compiled model APIs. by Gerardo Carranza · 13 hours ago
  13. 9b04b48 Remove obsolete highwayhash patching from Windows wheel build script by Terry Heo · 14 hours ago
  14. 42c8fe8 Fix -Wpass-failed build break in the XNNPACK MoE kernel under sanitizers. by Fengwu Yao · 14 hours ago
  15. eb610ee Use hardware simd_sum and single-phase shared-memory reduction in gated_delta_update shader. by Google AI Edge · 14 hours ago
  16. e52511b Make the OpenVINO android_x86_64 TAP project actually compile src_gen. by Matt Kreileder CA · 14 hours ago
  17. 5b39c03 PR #9968: Qualcomm AI Engine Direct - Switch from coarse-grained to fine-grained ops. by Jiun Kai Yang · 15 hours ago
  18. d703c55 Simplify GQA graph construction in YNNPACK and tighten SDPA decode1 threshold. by Volodymyr Kysenko · 15 hours ago
  19. 261401c Add a migration guide from TFLite Support to LiteRT Support. by Jun Jiang · 15 hours ago
  20. 2db38fd Optimize XNNPACK MoE expert weight dequantization, single-token GEMV, and scratch memory. by Fengwu Yao · 16 hours ago
  21. 9de33b5 Avoid linking XNNPACK in BuiltinOpResolverWithoutDefaultDelegates by Terry Heo · 17 hours ago
  22. 2f17ea5 Add YNNPACK CPU backend support to LiteRT ATS. by Gerardo Carranza · 17 hours ago
  23. ecdc41c Legalize REDUCE_ANY in the OpenVINO compiler plugin. by Tommy Chiang · 17 hours ago
  24. 724a247 Decouple YNNPACK CPU accelerator registration from compile-time preprocessor macros. by Gerardo Carranza · 17 hours ago
  25. 2134610 Add filegroup target for LiteRT C++ SDK sources by Catalin Termure · 18 hours ago
  26. 06705e9 Add mixed-precision and FP32 recurrent state support to LiteRT GPU delegate and custom ops. by Google AI Edge · 19 hours ago
  27. 541ebbb Move MakeConvWithBatchIds to common MoE utils and reuse it in LiteRT composite kernel. by Raman Sarokin · 19 hours ago
  28. 471a9fd Update README.md for LiteRT Swift user guide. by Jun Jiang · 27 hours ago
  29. cde1321 Enable BOOL mask pruning for all Flash-Decode head dims in sdpa_transposed. by Fengwu Yao · 30 hours ago
  30. dcfd9b7 Do not clamp Flash-Decode active tokens to param[0] for single-query causal SDPA. by Fengwu Yao · 31 hours ago
  31. 0d4a8cb Qualcomm AI Engine Direct - Fix addn_test.cc issue due to namespace by hungjuiw · 31 hours ago
  32. eed62af Support the tanh-approximated GELU activation (GeGLU) in the odml.swiglu GPU kernel. by Fengwu Yao · 31 hours ago
  33. b5de9a0 Support V RMSNorm, Q-only mode, partial RoPE proportion, and head_dim up to 512 in qkv_norm_rope. by Fengwu Yao · 32 hours ago
  34. 6c6af01 Remove the unused BUILD_CONVERTER option from the LiteRT wheel build by Terry Heo · 32 hours ago
  35. 7116bc9 Fix weight sharing, non-sharing compilation now produces separate context bin. by Andrew Zhang · 32 hours ago
  36. b50f4ba Update copyright header for colabs by Maria Lyubimtseva · 32 hours ago
  37. ef7ea74 Add the documentation and tests of nvidia's benchmark script by Google AI Edge · 33 hours ago
  38. 7970321 The NVIDIA dispatch library copied a CUDA tensor buffer to the device every by Google AI Edge · 33 hours ago
  39. 948486f Overlap the softmax and the score products of the tiled attention kernel by Google AI Edge · 33 hours ago
  40. 9a1d7f2 Run Gemma 4 prefill fully connected layers on the tensor-core GEMM by Google AI Edge · 33 hours ago
  41. 4122e7f Add a tensor-core GEMM for BF16 activations and INT4 weights by Google AI Edge · 33 hours ago
  42. 71684ba Accumulate the scores of BF16 queries in FP16 on the tiled attention kernel by Google AI Edge · 33 hours ago
  43. cffedb5 Run Gemma 4 prefill attention on a tiled tensor-core kernel by Google AI Edge · 33 hours ago
  44. d6ae6a4 Fuse Gemma 4 global attention into tensor-core CUDA kernels by Google AI Edge · 33 hours ago
  45. 519999c Implement WebGPU binary, comparison, logical, and ternary (Select/SelectV2) operations for LiteRT Tensor API. by Ping Yu · 34 hours ago
  46. 7278fe8 Add LiteRT Quantizer (litert_quantizer) to the LiteRT repository. by Maria Lyubimtseva · 2 days ago
  47. d408614 Remove org_tensorflow and tf_workspace dependencies from LiteRT OSS by Terry Heo · 2 days ago chromium/8083
  48. b460105 Rewrite QuantizationTests.swift in XCTest and drop the swift-testing dep due to OSS rules_swift version mismatch. by Jun Jiang · 2 days ago
  49. 2588243 Internal change by Tenghui Zhu · 2 days ago
  50. b263ebb Require cl_arm_import_memory_android_hardware_buffer for AHWB <-> OpenCL interop. by Fengwu Yao · 2 days ago
  51. 769e75b Add LITERT_WITH_TENSORFLOW to test TensorFlow targets in CI by Terry Heo · 2 days ago
  52. 3445317 Add DMA-BUF and peak memory reporting to LiteRT tools and runner script for embedding models. by Andrew Zhang · 2 days ago
  53. f60efba Prevent redundant TFLite ArenaPlanner allocations for subgraph I/O by Terry Heo · 2 days ago
  54. c0575cf Add INT4 FullyConnected test coverage to LiteRT ATS. by Gerardo Carranza · 2 days ago
  55. 91b2c64 Implement WebGPU unary, activation, cast, and shape-aliasing operations for LiteRT Tensor API. by Ping Yu · 2 days ago
  56. b585783 Extract functions that can be reused to run with YNNPACK. by Quentin Khan · 2 days ago
  57. 5ee5bda Automated Code Change by Google AI Edge · 2 days ago
  58. c4d5081 Remove source_location shim now that absl::SourceLocation is open source. by Quentin Khan · 2 days ago
  59. 0b67fbc Add GQA head ratio (H_v vs H_k) support to CPU and GPU gated_delta_update kernels. by Google AI Edge · 2 days ago
  60. 321b969 Add a test in attention when the KV cache grows. by Quentin Khan · 2 days ago
  61. 3291071 Prepare tests to run over multiple backends. by Quentin Khan · 2 days ago
  62. 68a99d8 Extract buffer mapping logic that can be shared between XNNPACK and YNNPACK runners. by Quentin Khan · 2 days ago
  63. 165f86d Support broadcast-based grouped-query attention (GQA) during prefill in YNNPACK. by Volodymyr Kysenko · 2 days ago
  64. a0c165e Unify Flash-Decode SDPA across Metal and OpenCL with UCL wave-SIMD support. by Fengwu Yao · 2 days ago
  65. 75d54fd Add TensorFlow's third-party repositories to the LiteRT OSS shim by Terry Heo · 2 days ago
  66. 826c40d Protect XNNPACK runtime creation and destruction with workspace_mutex_. by Google AI Edge · 2 days ago
  67. e272cd5 Merge pull request #10032 from graham0824:dev/hungjuiw/fix-addn-test by Copybara-Service · 2 days ago
  68. 8a9cc74 Merge pull request #9665 from graham0824:dev/hungjuiw/qcom-target-backend by Copybara-Service · 2 days ago
  69. c00c771 [LiteRT][Qualcomm] Set HTP file read memory budget and release mmap pages after QNN context deserialization. by Weiyi Wang · 2 days ago
  70. b015746 Add `cint2_fp32_int4_e8m0_drq` fusion pass and reference `FullyConnected` kernel. by Majid Dadashi · 2 days ago
  71. 59097e4 Add `cint2_fp32_int4_e8m0_drq` fusion pass and reference `FullyConnected` kernel. by Majid Dadashi · 2 days ago
  72. e868789 Merge pull request #9480 from graham0824:dev/mingxiup/expand_dims_op by Copybara-Service · 2 days ago
  73. 3233ef5 Merge pull request #10188 from graham0824:dev/hungjuiw/transformation-in-compile by Copybara-Service · 2 days ago
  74. 950d251 Fix SdpaTransposed reference evaluation for bool masks and 32-aligned active KV length by Gerardo Carranza · 2 days ago
  75. 0159226 Migrate remaining ML Drift enum references in //third_party/odml to kPascalCase. by Juhyun Lee · 2 days ago
  76. acbb3a0 Add stand-in XLA and LLVM repositories for LiteRT OSS by Terry Heo · 2 days ago
  77. 3884c94 Add a stand-in TensorFlow repository for LiteRT OSS by Terry Heo · 3 days ago
  78. 4fe2408 Bump pinned XNNPACK commit to pick up qs8_qc4w neoni8mm GEMM kernels. by Changming Sun CA · 3 days ago
  79. dbc8e13 Package TensorFlow Lite CoreML and Metal delegates as standalone xcframeworks and expose CLiteRT_static in SwiftPM. by Jun Jiang · 3 days ago
  80. 76fe528 Make ml_drift delegate C++17 compliant. by Google AI Edge · 3 days ago
  81. c112443 Update readme of LiteRT github repo by Google AI Edge · 3 days ago
  82. b6a9c77 Optimize MoE GPU delegate graph for decode and top-k reduction. by Fengwu Yao · 3 days ago
  83. 143d165 This is an internal change. by Chunlei Niu · 3 days ago
  84. 34d6836 Migrate ML Drift enum references in //third_party/odml/litert/ml_drift to kPascalCase. by Juhyun Lee · 3 days ago
  85. 391c38f Fix heap-use-after-free in moe_experts_parser_test under ASAN. by Fengwu Yao · 3 days ago
  86. c7bc981 Add an internal only Kotlin API to control fallback to CPU in tests. by Chunlei Niu · 3 days ago
  87. 516e119 Migrate ML Drift enum references in //third_party/odml/litert/tensor to kPascalCase. by Juhyun Lee · 3 days ago
  88. 7f8b9a7 Automated Code Change by Google AI Edge · 3 days ago
  89. ee5819f Consolidate MoE builder utilities. by Raman Sarokin · 3 days ago
  90. 0750b18 Optimize SELECT_V2 inner loops by hoisting scalar broadcast values. by Dillon Sharlet · 3 days ago
  91. 3971f72 Add grouped-query SDPA kernel tests for the decomposed graph. by Fengwu Yao · 3 days ago
  92. 732dea9 Automated Code Change by Google AI Edge · 3 days ago
  93. 77a7a2c Harden SdpaTransposed ATS generator and fix MLDrift SDPA mask/softcap handling. by Gerardo Carranza · 3 days ago
  94. 6d59736 Fix SdpaTransposed ATS generator Q tensor layout and GPU GQA fallback. by Gerardo Carranza · 3 days ago
  95. efb191a Support blockwise int8 MoE expert scales in the XNNPACK delegate. by Fengwu Yao · 3 days ago
  96. 85d78cd Include QAIRT version in Qualcomm SDK package description. by Andrew Zhang · 3 days ago
  97. 58d4927 Use ucl::Init<float4> in the short_conv_step kernel. by Fengwu Yao · 3 days ago
  98. ec4de24 [LiteRT][MediaTek] Fix dangling pointer in GetDimensions, enum sanitizer traps, and SoC case sensitivity by Andrew Zhang · 3 days ago
  99. 873d8e5 Remove TensorFlow deps from strip_buffers and analyzer_wrapper by Terry Heo · 3 days ago
  100. adce58e Support blockwise-quantized MoE expert weights on the GPU delegate. by Fengwu Yao · 3 days ago